Showing posts with label transcription. Show all posts
Showing posts with label transcription. Show all posts

Tuesday, 2 April 2013

Crowdsourcing text correction and transcription of digitised historic newspapers: a list of sites


Last month two new websites were launched giving the public access to digitised historic newspapers.  The release of a new ‘old’ digitised newspaper site is becoming a regular monthly occurrence now, with a library somewhere in the world completing a newspaper digitisation project with astonishing regularity, after what seems like such a long wait. 

The two new sites this month were the Welsh Newspapers online and the Louiseville Leader.

The Welsh site has been several years in progress and seriously considered using the National Library of Australia software for text correction, before putting it in the ‘too hard basket’. The National Library of Wales is to be commended on making the Welsh Newspapers service free (unlike the English newspapers which are still in a subscription model from the British Library).
 
The Louiseville site delivers all the issues of a key African American community newspaper covering local, national, and international news published in Louisville, Kentucky from 1917-1950. Unfortunately the building which housed original copies of the paper was badly damaged by a fire. The remaining issues, loaned by Kentucky State University and the widow of the publisher, were microfilmed by the University of Louisville, with the digital files created from that microfilm. The long and winding road the texts have taken toward digital representation has made them less than ideal candidates for optical character recognition (OCR), which has difficulty transcribing faded, torn, or misaligned texts, even when they are readable to the human eye. For this reason the site has enabled public transcription to help improve the accuracy and searchability of the newspaper content.
 
It’s great to see both of these new sites and I fully understand the difficult process many libraries have gone through to get to this point, having been there and managed a newspaper digitisation project myself. I still have a particular interest in those newspaper sites which involve the public in text correction, which is another step perhaps just too challenging for many libraries to take.  After the worldwide library applaud of the Australian Newspapers/Trove text correction beta five years ago, now an internationally hailed success, and the stated intent of many libraries to follow suit with public text correction the question arises “how many actual did?”
 
There are many libraries internationally that now offer websites to search across digitised historic newspapers and I’m not going to list all of them, just the handful that give their users the text correction or transcription ability. With Australian text correctors, now addicted to text correction of newspapers and looking elsewhere to sate their ample appetites I thought it was time to compile a list specifically of text correction websites for historic newspapers. To the best of my knowledge there are 9 sites now.  Who will be the 10th?? If I have inadvertently missed a site perhaps let me know in the comments. Most of the sites are for English language content but it is interesting to see a few coming through for other languages.  As a note of interest there were several foreign language historic newspapers published in Australia (Chinese, Greek, Hebrew, German) but these were put in the ‘too hard basket’ for the first stage of Australian Newspapers/Trove and sadly did not make it into the second stage either.  They give a very interesting perspective on sub communities within a wider community.
 
Congratulations to all the libraries listed below who took the first difficult step to digitise and then the more challenging step to crowdsource. Happy text correcting to all the amazing people that volunteer their valuable time to help libraries make old newspapers more accessible, I hope you enjoy the list. The sites are all slightly different but work on the general basis of showing a digitised page and asking for public correction/transcription of the OCR text created from that page. If the OCR text is improved then keyword searching of the newspapers is improved.  It particularly helps to correct people’s names, especially in family notices, births and deaths, since these are often the first thing that users search on.
 
List of historic/old digitised newspaper sites that offer public text correction/transcription: March 2013
US Newspapers
 
Australian Newspapers
Finnish Newspapers
Vietnamese Newspapers
Russian
 
Useful Resource:
Frederick Zarndt’s recent PowerPoint on crowdsourcing in libraries with a particular focus on newspapers:

Saturday, 10 November 2012

National Archives of Australia embraces crowdsourcing and releases ‘The Hive’.


 
The National Archives of Australia (NAA) has made a bold step into the cultural heritage crowdsourcing arena with ‘The Hive’ which was released two weeks ago. The brand makes a clever play on the word ‘Archive’ combined with the idea of a hive of working bees (the public).  The site encourages the public to transcribe archive records.

Early this year when David Fricker became Director General of the NAA he was quick to encourage staff to think innovatively, embrace change, and to harness opportunities such as crowdsourcing to improve access to our collections. He publicly spoke in favour of  crowdsourcing and a changing business model for archives at the International Council of Archives Congress in August:

“Another key development in expanding access is crowdsourcing. As many of us are now seeing, by allowing the public to contribute to the description of archival resources we are enhancing the ability of future generations to discover and learn from our archives. I also think it is a wonderful opportunity for the public to be more engaged with us as archives and to share in the work we do – preserving the memories of our nations. There is still some work to do here, in order to maximise the value of contributions and to maintain the integrity of our archives as authentic and accurate. However, I do not believe these problems are insurmountable, and indeed I believe these systems can to some extent be self-correcting.

This is a type of the co-design, citizen first activity… drawing on the interest and enthusiasm of the community to bring more of our archives into view – discoverable and retrievable…Access will be online and everywhere, improved by rich new data visualisation techniques and expanded descriptive contributions from an engaged citizenry”.

The Hive is the Archives pilot and experimentation into the potential of large scale transcription crowdsourcing to improve access to records.  Staff have looked closely at other crowdsourcing sites on offer and attempted to build on their knowledge and techniques, to provide a site that could be used as a large scale platform for a variety of transcription crowdsourcing projects.

At present the site offers just over 800 lists for the public to transcribe. Some of these are typed and some handwritten.  They are rated in difficulty as easy, medium or hard.  Part of the difficulty with this project is that the public need to have some understanding of how archives receive and describe their records to make sense of what they are being asked to do.  In simple terms archives receive vast amounts of records (referred to as consignments).  Each consignment comes with a list of the items in it.  However because of the large volume of records being received it is usual that only the consignment record is entered into the catalogue e.g. ‘100 boxes of plans and drawings’ from x government agency, rather than all the individual items on the consignment list being described in the catalogue.  The ideal scenario for users of the archives is that every item e.g. plan and drawing is described on the catalogue so that it can be found.  Without this a lot of guess work goes into finding relevant things, or alternatively personal visits are required to view the hard copy consignment lists.
 
The project that the archives is undertaking is to digitise consignment lists and then make them available for transcription by the public. Once transcribed they become searchable and the items within them can be found more easily.  Because so many of the lists are old and handwritten it is virtually impossible to get good OCR on them.  That’s where the public come in who can read them with the human eye. Also the time of the public is needed to speed up the access. Projections on the time it would take archives staff to describe the lists without public help currently stand at 210 years.  It is anticipated that a member of the public could with relative ease describe several hundred items per hour with the Hive tool, which would make a big difference, especially if there was a swarm.

The consignment lists in the pilot are those that have proved most popular with researchers and contain items in the ‘open period’, that is older than 30 years and now open to the public.  The top interest is lists of architectural drawings and historic buildings. This is closely followed by PNG patrol officer records, maritime incidents, personal records from the war office, prisoners of war, meteorology and cyclones, WW1 intelligence, and oil drilling on the Great Barrier Reef.

In the first 2 weeks 300 records have been transcribed of the 800. There is a definite preference for the lists rated hard (handwritten) and ones that involve names.

The site is well presented and gives volunteer transcribers things we know they want such as progress chart, recent activity, points scoring system, rewards, optional login using Open ID e.g. their Google ID, ability to search and choose items, or just take the next one served up, to pick easy or difficult items, to add a marker for where they got to if they are interrupted, and to favourite records.  The only slight drawback is the placing of the transcription window at the bottom of the screen rather than right or left, which often means it is hard to see the transcription window and the content you are transcribing at the same time. Also the OCR text in the transcription window and the cursor is not hooked directly to the text in the image so it is easy to get lost whilst transcribing sometimes.  This is largely because most of the lists are in tables, and the table rows and columns have not been retained in the OCR, so the OCR is somewhat muddled.  Further development of the site will largely depend on feedback given by the public users, and the ability of the archives to keep up a steady supply of new, interesting digitised consignment lists to the Hive.  The Archives is still considering how it may be able to integrate the public content back into its main catalogue RecordSearch, or integrate the Hive into RecordSearch. In the meantime the list content will remain searchable in the Hive.

There is obviously an expectation from the Archives that by making its content more discoverable it will lead to more access requests.  This is why at point of transcription there is a button which enables the user to request a copy of the item.  These requests are being met by digitising the item, and then uploading them into the main catalogue ‘RecordSearch’ with the full item description.

I congratulate the National Archives of Australia Access Team on the development of this exciting new site, which holds so much potential to improve access to records and engage with our citizens in new ways.

The screenshots below show the site in action:



 Easy level transcription- Archived drawings

Medium Level Transcription - ABC Drama Scripts
 


Difficult level transcription - Plans
 


Saturday, 25 August 2012

Crowdsourcing and Social Media at US National Archives (NARA). The Citizen Archivist Dashboard


Last week I attended the International Congress of Archives (ICA 2012) which was held in Brisbane. Over 1,000 Archivists from 93 countries attended.

The much anticipated opening keynote on the first day was given by David Ferriero
head of US National Archives.  He is the first librarian to become a National Archivist, previously being in charge of New York Public Library and known for promoting use of social media and relationships with Google and Wikipedia.  His talk was called ‘A world of social media’. I was looking forward to hearing what the US National Archives are doing with social media and crowdsourcing.  People were generally of the opinion that this organisation will/is leading by example in this field.

David Ferriero took to the stage and took us by surprise.  He only used 20 minutes of his 40 minute slot, gave no presentation, instead reading from his notes at breakneck speed and bombarding us with statistics that were largely out of context. At the end he took no questions and dashed off the stage.  He left a surprised and bewildered audience behind.  I for one was immensely disappointed not to see and hear more about some of the exciting US Archives activities. He of course may have had mitigating circumstances that I am totally unaware of.  He did however give small tasters of what his organisation is doing. There was brief mention of large scale crowdsourcing on unspecified projects, a citizen archivists dashboard, and a relationship with Wikipedia which peaked my interest.

So I decided to follow up online and find out for myself what may be happening at NARA. I took me quite some time to search the internet and blogs and get the information I had hoped David would give in his keynote, but it was worth it. Here is what I found:

1. Citizen Archivist Dashboard Webpage http://www.archives.gov/citizen-archivist/

In January 2012 the US National Archives launched the Citizen Archivist Dashboard. This is a great webpage bringing all the online and physical social engagement and crowdsourcing activities together.  It is easy for someone to see what options they may have to help the US National Archives. It is very clearly designed and I like it a lot.

 
2. Transcription Projects

There are two transcription projects going on for handwritten records. Firstly the National Archives Transcription Pilot Project. It appears still to be in ‘pilot’ mode (started in January 2012) since only 300 documents (about 1,000 pages) are available for transcription. They have been very carefully selected from a collection of billions of pages and graded by colour codes according to how difficult the handwriting is to read. This pre-selection must have taken very valuable staff time. You can browse or search by difficulty of transcription, year, and the status of transcription: “Not Yet Started,” “Partially Transcribed,” and “Completed.” You then choose a page to work on and then that page is blocked to other users, so it’s not being edited by multiple users at the same time.  The interface is very simple, much like the Australian Newspapers. In a free text box beside the image you can transcribe what you see. No login is required, though you do have to complete a captcha. 

The missing part is that I can’t see how many people have transcribed what.  It’s not clear if the documents disappear from here when fully transcribed, and how and where they become full text searchable in the collection.  It also seems to be a time consuming process for NARA staff to do the pre-selection and difficulty rating of the documents. This is of course a very small pilot and hopefully lessons will be learnt and the site will be developed further to reach it’s full potential. Also it would be good if more documents became available for transcription. This is one of the easiest handwritten transcription tools I have seen.  I could not find any information about who developed the tool and if it is available open source.

Interestingly David Ferriero says that many US school children are no longer taught cursive handwriting and therefore cannot read handwriting. He says ‘Help us transcribe records and guarantee that school children can make use of our documents’. I’m not quite clear if he thinks this is a potential crowdsourcing exercise for school children to learn handwriting and become better educated, or if adults are supposed to do it so that school children can just read the finished text.

The National Archives have developed a relationship with the Wikipedia Community and currently have a Wikipedian in residence. As part of that program they have shared some primary handwritten national documents into ‘Wikisource’ for transcription via the Wikisource Tool. These documents are mostly at the beginner level in terms of difficulty. I’m not clear if they are the same ones in being used in the Archives own pilot, or different documents. I’m also not clear why they are piloting two different methods for transcription, or what the initial results are compared to each other. Wikisource offers more than transcription however, Wikipedians (if they can get access to original documents or copies) can also scan documents and OCR them.

3. Scanning Projects

  • Scanathons
For reasons I don’t understand the US National Archives has only digitised 750,000 of its 40 million images. This is a very low figure for an organisation like this. They seem to be focusing quite a lot of effort on getting physical volunteers to come in person to the Archives to digitise/scan images for them at ‘Scanathons’. This started in 2011. In January 2012 there was a 4 day Wikipedia ExtravaSCANza. Over the 4 days a group of Wikipedians met in the Still Pictures Research Room and scanned 500 images on desktop scanners. Each day there was a theme: NASA, women’s history, Chile, and battleships.

NARA encourages readers to take their own photos of records in the reading rooms and upload them to a special group in Flickr.  The important thing here is that they should also be described with title, series, and record group if possible so they can be found. So far only 20 people have joined the group and 133 photos have been uploaded (most of these by the same person). I’m not clear how NARA intends to link these digital images back to the item descriptions in their collections but this is a great idea to tackle large scale digitisation of images.

 

The tagging facility, unlike the other pilots seems to me to be unlikely to succeed in its objectives. This is perhaps because of the tight controls that have been placed around it and the isolation of the activity from normal search and browse behaviour. Whilst anyone can easily transcribe a record without needing to login the process for tagging is difficult.

The activity is focused on Tuesdays and themed around a topic.  Records for the topic are pre-selected by the Archives and available in an online group e.g. Elvis, Titanic.  Volunteers must register and follow a set of guidelines; Tags will be reviewed by NARA staff before being accepted and going live on the database. I looked at the topics and it was unclear to me why if the Archives had already identified the items as being about Elvis they couldn’t simply generate an automatic tag for ‘Elvis’. In my opinion tagging is not actually a crowdsourcing activity because individuals are motivated to add tags to help themselves find things, it is a by product of search. Research shows it is rare for users to have concensus on tag terms and use. Crowdsourcing activities achieve a big clear goal that could not be achieved by individuals alone, and everyone in the crowd should be aware of how they are helping the ultimate goal.  

5. Indexing the 1940 Census

On April 2, 2012, NARA released the digital images of the 1940 United States Federal Census after a 72 year embargo. The census images will be uploaded and made available on Archives.com, FindMyPast.com, National Archives, ProQuest, and FamilySearch.org. The entire 1940 census data will be indexed by a community of volunteers and made available for free. The free index of the census records and corresponding images will be available to the public for perpetuity.

6. Useful Links

I found a recent presentation given this year by Pamela Wright – Chief Digital Access Strategist at NARA which gives screenshots of what I have talked about above. ‘From access to engagement’

7. Social Media

NARA are active users of social media channels and they have started to monitor their activity. The Social media statistics from NARA May 2012 may be interesting reading for some.

I would be interested in reading more presentations or articles about the citizen archivist pilot projects from NARA and finding out what they have achieved and learnt so far. I hope this information is made available to the archives and library community soon.  Please reply in comment if you have any more information on the pilot activities.

Sunday, 17 June 2012

If only they would crowdsource! – Diamond Jubilee - Royal Archives at Windsor Castle


Many years ago I worked for a software company installing the first archive management systems into large UK archives such as the London Metropolitan Archives, Cumbria Archives at Carlisle Castle and the Royal Archives at Windsor Castle.  It was a challenging time for archives going from paper systems to computer systems, in fact very similar to the challenges archives now face transitioning from managing paper records to born digital records.  Ironically I have just returned again to the archives sector and am now working at the National Archives of Australia on the second challenge.

When the first computer systems were installed in archives it often came as a shock to archivists to discover that when the system was installed it would be ‘empty’ and their records would not somehow miraculously appear in the system. This was the first piece of news I usually had to convey in training before showing an online process for acquisitions. I particularly remember that at the Royal Archives they estimated with their current staff of 4 it would take them 700 years to record their archive collection into their new system, and they were somewhat despondent to say the least. Nevertheless the Queen was pleased with the install of the first computerised system at Windsor Castle and awarded the software company I worked for the Royal Warrant, which meant we could use the Royal Coat of Arms on our letterhead.  The warrant is more often seen on pots of jam and pickle than on software. The celebration of implementation party at Windsor Castle with members of the Royal Household and staff was one to remember. 

The Round Tower at Windsor Castle contained every hand written record every monarch and members of their household had ever created. Queen Victoria’s collection was particularly large.  The Royal Archives could only be contacted by letter and each year less than 10 well vetted members of the public were allowed to access a very restricted and pre-agreed part of the collection under strict supervision.  Because it was largely uncatalogued, described or known there was a terrible fear of what a member of the public might find in the archives. This was understandable since household records such as the cost of banquets were intermingled with personal letters and diaries.  From the public's point of view the archive is that of our Kings and Queens and we would like to access it, but from HRH's view it is her private family archive. Although it is now more acceptable to expose skeletons in the family closet and programs such as "Who do you think you are" promote this, there is probably a reticence from aristocracy and royalty to do this. The Royal Archives is one of the richest, most interesting and significant collections ever created.  It could aptly be described as a pot of gold – an absolute treasure trove. The archivists were aware of this and some of the treasures within it.  The Royal Library at Windsor Castle was in a similar situation and also had extremely restricted access.  Because I have always championed access to archive and library collections I felt very sad whenever I thought of the treasures locked up and hidden (literally) at Windsor Castle.

I was very interested therefore to read about a new development at the Royal Archives timed to coincide with the Diamond Jubilee.  The Queen released this message:

“In this the year of my Diamond Jubilee, I am delighted to be able to present, for the first time, the complete on-line collection of Queen Victoria's journals from the Royal Archives. These diaries cover the period from Queen Victoria's childhood days to her Accession to the Throne, marriage to Prince Albert, and later, her Golden and Diamond Jubilees. Thirteen volumes in Victoria's own hand survive, and the majority of the remaining volumes were transcribed after Queen Victoria's death by her youngest daughter, Princess Beatrice, on her mother's instructions. It seems fitting that the subject of the first major public release of material from the Royal Archives is Queen Victoria, who was the first Monarch to celebrate a Diamond Jubilee. It is hoped that this historic collection will make a valuable addition to the unique material already held by the Bodleian Libraries at Oxford University, and will be used to enhance our knowledge and understanding of the past.”

I was intrigued by this and immediately found the website http://www.queenvictoriasjournals.org/home.do which tells us a lot more about Queen Victoria’s diaries and that this was a project undertaken in conjunction with the Bodleian Library at Oxford and Pro-Quest.  However on looking further it was a bit disappointing since although every page of all the journals has been scanned they have not all been transcribed.  Because they are all handwritten, they won’t be fully text searchable until they are all transcribed, a process which at present is most effectively done by the human hand and eye.  The website doesn’t give any indication of when or how they will be transcribed that I could see, although it says it ‘is in progress’. So far only the first diary has been transcribed, by whom I am not sure. I bet the project is only letting academics do it, who will be paid lots of money and progress very slowly. There is a lot to do: 1832- 1901 since Queen Victoria wrote her diary every day.

If ever I saw a collection that was so well-suited to crowdsourcing for public transcription this is it!  I could guarantee that in a few days or weeks all of Queen Victoria’s diaries would be transcribed by a willing and fascinated public. The handwriting is hard to decipher but with thousands of eyes, and amateur/professional genealogists and historians used to reading old writing, that are highly motivated I am sure it could be achieved. I feel excited just imagining it.  But why stop there?  What about the rest of the collection - the offical royal records and the personal records?  When is that going to come out of hiding? It’s just crying out for public description, tagging and transcribing. If only.  If only.

Extract of Queen Victoria’s diary.

Then thinking I would come back later and have another look I was most disappointed to read that following the example set by the British Library with its UK digitised newspapers the intent is to restrict access to the UK only, and to charge for access from July.  So loyal British subjects living in Commonwealth Countries, and academic researchers – you only have 14 more days to look at this for free, or at all.  Great shame!  But congratulations to whoever it was behind the scenes that convinced HRH to release the diaries from the Royal Archives, and who set up and managed the project with the Bodleian and Pro-Quest.  Bravo!!  Perhaps we just need to beg and grovel for more content and offer our unconditional help to get it for free.

Photo by Rose. June 2012. After participating in Trooping the Colour for the Queens Birthday in Canberra, Irish Guard Cliff Doidge (who plays the clarinet in the Royal Military Band and is on exchange from London to Australia for 4 months) stands beside Lake Burley Griffin with the National Library of Australia behind.