Showing posts with label text correction. Show all posts
Showing posts with label text correction. Show all posts

Thursday, 26 October 2017

National Digital Library of India

I was recently invited to travel to UNESCO HQ in New Delhi, India to give a keynote presentation at the UNESCO-National Digital Library of India International Workshop  ‘Knowledge Engineering for Digital Library Design’.  A small group of professional international experts had been invited to share their knowledge in the area of their expertise, mine being crowdsourcing in libraries, newspaper text correction and user led digital library design.

The Government of India along with the Ministry of Human Resource Development and the Indian Institute of Technology Kharagpur (IITK) are working on a project to develop the National Digital Library (NDL).  This is going to be an important part of their national academic infrastructure and is being led by Professor Partha Chakrabarti from IITK.  The question they really wanted me to answer for Trove and Australian Newspapers was “if you were doing it all again, starting now, what would you do differently”? This is so they can apply the knowledge that Australia learnt in the Indian Digital Library project.

Unlike other countries India has a growing young, rather than old population and they are not able to build and expand their universities in a timely way to meet educational demand.  Therefore the government are seeing the National Digital Library as an academic network for individual and community based learning. It mainly contains books and courses and they will shortly develop the module to deliver historic newspapers.  Academic libraries are also having their subscription resources harvested into it, since it is intended to be the backbone of the academic learning network.  It shares some similarities with the Australian equivalent Trove. This interesting explains the National Digital Library of India.


I was really pleased to find out that India are now utilising their technology expertise for themselves in this way.  If it was not for India we would not have Trove and Australian Newspapers.  The National Library of Australia has been sending the more technical workflow aspects of the historic newspaper digitisation out to India contractors for the last ten years. In 2008 I was overseeing this and I had the pleasure to visit the digitisation facilities and meet the hundreds of staff in Hyderabad, Chennai, and Delhi, as well as some of the call centres and technology companies in Kolcatta.  It was an eye opening experience for me to see such cutting edge technology and bright young people filled with hope and career aspirations, alongside with such extreme poverty and slums.  India is an experience that lets you see the whole of humanity in a single day, which can be quite overwhelming.

UNSW has just started to make a series of targeted and highly strategic investments in developing transformative partnerships in India. India represents a major priority for UNSW as part of its 2025 Strategy under the Global Impact pillar.  Building successful research and knowledge exchange partnerships in India will be key to the success of the UNSW India Strategy. India is a growing source of innovation and is home to some of the world’s most dynamic and innovative companies who are at the forefront of digital disruption, social enterprise and inclusive development. India’s research system is also growing as the Government of India considers its investment and capacity building strategy in higher education and research.

This week for Diwali the UNSW campus is being transformed and ‘The Festival of India 2017’ will be a stimulating, event-packed week that celebrates and promotes Australia’s partnership and friendship with the rapidly emerging global powerhouse – India. As the campus grounds transform into a little India – this unique festival will showcase not only the country’s rich, cultural offerings but also its ground breaking developments in innovation, finance, scientific research and economic growth.  In November there will be an inaugural Research Roadshow:


  • To enable UNSW researchers to travel to India to make new connections and/or strengthen existing relationships;
  • To showcase UNSW’s capabilities, especially in the following research areas: smart cities, energy, water, climate, health and social enterprise sectors to prospective Indian partners.
  • To initiate and nurture strategic industry partnerships that will lead to knowledge exchange outcomes.


Expected outcomes will be the identification or consolidation of opportunities that will lead to future collaborative research partnerships with academic, industry and government partners.

In the meantime as I contemplated preparing my presentation I unfortunately was one of the many Australian’s struck down with the virulent strain of influenza in the recent Australia wide outbreak of flu.  As the time approached to travel I realised I really was still not well enough.  This resulted in me asking the Creative Media Unit at UNSW Canberra for help. John Carroll used his wonderful skill and technologies to create a 40 minute video of my presentation which was delivered to New Delhi on video screen, link below.

I discuss the findings of my nine years of research into crowdsourcing based curation in libraries.  Using the digitised historic Australian Newspapers as an example, I look at how the functionality and interface was developed in close relationship with the users,  and how this led on to text correction of newspaper articles. It is nearly ten years since this pioneering project began and the motivations and achievements of the 50,000 volunteers are examined over this time. I question how successfully the goal of improving text quality and therefore search has been achieved?  I propose that if a similar project was begun now then artificial intelligence software would be used such as OverProof post OCR correction tool to improve the quality of the text.  OverProof has been trained on the manual corrections of the Australian newspaper corpus and trials demonstrate it is able to dramatically improve the quality of the corpus. Volunteer text correction could still continue afterwards for difficult text but the software would do the main donkey work, allowing users to have a better quality search.



The PowerPoint is on my slideshare account.




Tuesday, 2 April 2013

Crowdsourcing text correction and transcription of digitised historic newspapers: a list of sites


Last month two new websites were launched giving the public access to digitised historic newspapers.  The release of a new ‘old’ digitised newspaper site is becoming a regular monthly occurrence now, with a library somewhere in the world completing a newspaper digitisation project with astonishing regularity, after what seems like such a long wait. 

The two new sites this month were the Welsh Newspapers online and the Louiseville Leader.

The Welsh site has been several years in progress and seriously considered using the National Library of Australia software for text correction, before putting it in the ‘too hard basket’. The National Library of Wales is to be commended on making the Welsh Newspapers service free (unlike the English newspapers which are still in a subscription model from the British Library).
 
The Louiseville site delivers all the issues of a key African American community newspaper covering local, national, and international news published in Louisville, Kentucky from 1917-1950. Unfortunately the building which housed original copies of the paper was badly damaged by a fire. The remaining issues, loaned by Kentucky State University and the widow of the publisher, were microfilmed by the University of Louisville, with the digital files created from that microfilm. The long and winding road the texts have taken toward digital representation has made them less than ideal candidates for optical character recognition (OCR), which has difficulty transcribing faded, torn, or misaligned texts, even when they are readable to the human eye. For this reason the site has enabled public transcription to help improve the accuracy and searchability of the newspaper content.
 
It’s great to see both of these new sites and I fully understand the difficult process many libraries have gone through to get to this point, having been there and managed a newspaper digitisation project myself. I still have a particular interest in those newspaper sites which involve the public in text correction, which is another step perhaps just too challenging for many libraries to take.  After the worldwide library applaud of the Australian Newspapers/Trove text correction beta five years ago, now an internationally hailed success, and the stated intent of many libraries to follow suit with public text correction the question arises “how many actual did?”
 
There are many libraries internationally that now offer websites to search across digitised historic newspapers and I’m not going to list all of them, just the handful that give their users the text correction or transcription ability. With Australian text correctors, now addicted to text correction of newspapers and looking elsewhere to sate their ample appetites I thought it was time to compile a list specifically of text correction websites for historic newspapers. To the best of my knowledge there are 9 sites now.  Who will be the 10th?? If I have inadvertently missed a site perhaps let me know in the comments. Most of the sites are for English language content but it is interesting to see a few coming through for other languages.  As a note of interest there were several foreign language historic newspapers published in Australia (Chinese, Greek, Hebrew, German) but these were put in the ‘too hard basket’ for the first stage of Australian Newspapers/Trove and sadly did not make it into the second stage either.  They give a very interesting perspective on sub communities within a wider community.
 
Congratulations to all the libraries listed below who took the first difficult step to digitise and then the more challenging step to crowdsource. Happy text correcting to all the amazing people that volunteer their valuable time to help libraries make old newspapers more accessible, I hope you enjoy the list. The sites are all slightly different but work on the general basis of showing a digitised page and asking for public correction/transcription of the OCR text created from that page. If the OCR text is improved then keyword searching of the newspapers is improved.  It particularly helps to correct people’s names, especially in family notices, births and deaths, since these are often the first thing that users search on.
 
List of historic/old digitised newspaper sites that offer public text correction/transcription: March 2013
US Newspapers
 
Australian Newspapers
Finnish Newspapers
Vietnamese Newspapers
Russian
 
Useful Resource:
Frederick Zarndt’s recent PowerPoint on crowdsourcing in libraries with a particular focus on newspapers:

Tuesday, 2 October 2012

Digital Motor Archive available free for the month of October in return for…...


It came to my notice last week that a publisher in the UK had digitised both the current and back issues of their magazine the ‘Commercial Motor’ which in its early life was a newspaper. In the cut throat world of publishing there are few publishers left that are still publishing the same title they were 100 years ago, and who also have a complete set of back copies.  If they fall into this bracket they are in the unique position of being able to either digitise the content themselves for their readers (usually at a loss); offer it to a library to digitise (at no cost); or sell it to a commercial e-vendor to package with another product for academia (and make a profit).  Unfortunately most choose the latter, which makes this type of content only really accessible to academics and students via academic libraries.  E-vendors normally charge high subscription rates to digitised magazines and newspapers and package them up with other content, making it only viable for large universities and national libraries to purchase, and therefore severely restricting readership to the content.

Not a lot of publishers are digitising their own content because generally speaking the cost of preparation, digitisation and OCR, and building a good website to deliver the content outweigh the amount of money they would ever recuperate from reader subscriptions. Normally a subscription to the current copy would be packaged up with old copies.  And that’s where the model fails, because often current readers have no interest in the old stuff.  People who do have an interest in the old stuff are generally a different group of people – historians, researchers etc.  The exception to this rule appears to be anything to do with hobbies such as knitting, cooking, railways, cars, and stamps. 

The Commercial Motor Archive http://archive.commercialmotor.com/ came to my notice because it is available for free for the month of October and I wondered why. It is a rich archive going from 1905 to the present day, covering a complete century and two world wars, well illustrated and with everything you ever wanted to know about commercial vehicles.   

The search and browse mechanism is very impressive and works well.  For example articles on pages have been zoned so you can search and find the article easily within a page.  It has many similarities to the hugely popular Australian Newspapers http://trove.nla.gov.au/newspaper.  The page displays alongside the OCR text to make it easier to read. You can browse covers, browse by date, and zoom in on pages.  Results can be filtered. Users can add comments and tags.  The quality of the OCR text and therefore search is very good.

It has one thing that was never implemented on Australian Newspapers (though often asked for by users) which is a little box on each page called ‘Report an error’ and this is the reason for its free access.  The site owner is hoping that as people use and read the pages in the archive they will report errors as they see them, and for this they get free access to the content.  The known errors that need to be identified are incomplete articles (where the zoning has gone wrong); OCR text error in headlines of articles; and OCR errors in text. However readers can only report them, not actually fix them.  The site states:

“the archive is beta because it isn’t perfect at the moment and there are a few glitches to be ironed out. Every article page has a 'Noticed an error?' button you can use to report a problem. Please don’t expect an immediate change to the error - we gather all the reports together and prioritise them, fixing the most pressing errors first.”

It sounds like they don’t know how many errors there are, how many people will report errors, and who and how the errors will be fixed.  Interesting.

Although the ‘report an error’ button was never implemented on Australian Newspapers (it was mainly needed to report upside down and duplicate pages) it had already been decided that a ‘super user’ would be the person to review these reports and take action.  In a world where some volunteer text correctors wanted to take on extra responsibility and have special roles like the hierarchy in the Wikipedia editors community this would have been a good thing for trusted volunteers to do.

The Commercial Motor Archive has impressed me because they are clearly striving for perfection, they understand that the fewer mistakes there are the better the search will be and they have taken a brave step by asking the public to help in return for free access.  This is indeed unusual for a commercial publisher, belonging more in the realms of libraries and archives and referred to as crowdsourcing……

 

Saturday, 11 February 2012

Crowdsourcing: more cool sites to give libraries, archives and museums inspiration

Many people know of my interest in the relevance and application of online digital crowdsourcing for libraries, archives and museums, due to an article I wrote in 2010 called ‘Crowdsourcing: how and why should libraries do it?’, and my initiation of the Australian Newspapers public text correction. People therefore often send me links to sites they think may interest me.  This is really great. Sometimes sites which are nothing to do with libraries or archives may give us ideas. There is a ‘List of Crowdsourcing Projects’ in Wikipedia (which is separate to the main article on crowdsourcing). This is a useful starting point to get an overview of the sorts of activities going on. It goes without saying that Wikipedia is of course the greatest crowdsourcing project ever!
In this post I wanted to mention some newish crowdsourcing projects that I have been looking at that interest me, and that I haven’t written about before. 
1.       Star Wars Uncut (SWU) Released August 2011
About the project:  In 2009, Casey Pugh a web developer asked thousands of Internet users to remake "Star Wars: A New Hope" into a fan film, 15 seconds at a time. Contributors were allowed to recreate scenes from Star Wars however they wanted.  Multiple submissions were submitted for each scene, and votes were held to determine which ones would be added to the final film. Although the scenes reflect the dialogue and imagery of the original film, each scene is created in a separate distinct style, such as live-action, animation and stop-motion.  Within just a few months SWU grew into a wild success. The creativity that poured into the project was unimaginable. SWU has been featured in documentaries, news features and conferences around the world for its unique appeal. In 2010 it won a Primetime Emmy for Outstanding Creative Achievement in Interactive Media. Now the crowdsourced project has been stitched together and put online in YouTube and Vimeo. The "Director's Cut" is a feature-length film that contains hand-picked scenes from the entire StarWarsUncut.com collection.

Relevance for libraries and archives:  In the world of film, TV and radio fans and consumers are the subject experts.  They not only have in-depth knowledge, but also have the motivation and interest to share their knowledge with others in creative ways. This project really shows that.  The fans apparently had no trouble identifying specific seconds in a very long film.  This type of knowledge and interest is really useful for librarians and archivists when you want to open up discovery of audio items.  It is much more likely that a fan will know which series, episode, minute and second a subject came up, or a thing was said than the librarian who created the catalogue record. The knowledge could be used to help with the discovery process.  At the moment most audio is still catalogued and described at item level for example “it’s an interview with x”. It is still a costly and difficult process to convert speech from audio into text, and to manually add subject tags.  Most of our historic audio collections do not have this level of discoverability. A crowdsourcing project which taps into the crowd to help make films more discoverable by use of public tags is ‘Waisda’.   We know that the public like to consume by watching and listening, but they also want to create and share. There is potential for crowdsourcing to improve accessibility of historic digitised audio especially that which has a fan base or is iconic.

2.       What’s on the Menu (New York Public Library) Launched April 2011.
About the project:  With approximately 40,000 menus dating from the 1840s to the present, The New York Public Library’s restaurant menu collection is one of the largest in the world, used by historians, chefs, novelists and everyday food enthusiasts. But the menus cannot be searched for specific information about the dishes and prices. To solve this problem the NYPL is appealing for the public to transcribe the menus, dish by dish. Doing this will enable the collection to be accessed and researched in new ways, opening the door to new kinds of discoveries. The site was launched in late April 2011 and the original aim was to transcribe the 9,000 menus photographed several years before for inclusion in the NYPL Digital Gallery.   Volunteers transcribed all of these in the first three months, so more items have been scanned from the collection and are now awaiting transcription. As of 5 February 2012, there have been 758,748 dishes transcribed from 12,167 menus. The ultimate goal is to get the whole collection transcribed and to turn it into a powerful research tool.  NYPL are also looking into partnering with other libraries and archives with menu collections.

Researchers who use the collection for example historians, chefs, nutritional scientists, and novelists, are looking for a juicy period detail. They often have very specific questions they’re trying to answer for example:

“Where were oysters served in 19th century New York and how did their varieties and cost change over time?”
 “When did apple pie first appear on a menu? What about pizza?”
“What was the price of a cup of coffee in 1907?”

To find out these sorts of things more easily, the text on the cards needs to be transcribed.  Quotes on their website about the usefulness of the project:
Rich Torrisi, New York Chef:

What’s on the Menu is a tremendous educational resource that breathes life into our city’s most beloved restaurants and dishes.  It has been an indispensable and hugely inspirational tool in the ongoing development of my restaurant…”

Mario Batali, New York Chef, Author, Entrepreneur:

“Menu writing is an art form seldom appreciated, In our restaurants, we put an incredible amount of time and thought into crafting menus. It’s remarkable to see menus being preserved and documented, for them to become a resource for future chefs, sociologists, historians and everyone who loves food.  It’s not just What’s on the Menu, it reveals so much more.”

Relevance for libraries and archives:  Libraries love to collect and keep stuff and that includes things like menu’s, tickets, pamphlets, posters, invitations, theatre programs and greeting cards. We call this stuff ‘ephemera’. Ephemera is a Greek word and it means printed matter that it is intended to be transitory, short lived, or only last a day.  When the item is created it is not intended that it will be retained or preserved.  However I haven’t encountered a single library that did not have a large ‘ephemera’ collection and intend to keep it long-term. The National Library of Australia is no exception and collects ephemera because it is “a record of Australian life and social customs, popular culture, national events, and issues of national concern”. There are 2.3 million items of ephemera in the collection at the NLA. Nearly 170,000 of them have been digitised and are browsable by title.
However their full potential has still not been unlocked.  Ephemera is printed on a few pages which usually contain both words and pictures.  When ephemera is digitised it is scanned or photographed as an image file, and therefore the text is not indexed or searchable.  It would be very hard to apply OCR on the text because of the varying and usually fancy typefaces used.  The only way to make the text searchable, thereby unlocking the full discoverability potential is to manually transcribe it.  Librarians don’t have time for this, but an interested public do.  Give them a really interesting or topical ephemera collection like the menu cards and watch them go!
3.       Historypin Launched July 2011
About the project:  Historypin was launched in July 2011.  It allows people to upload historic and contemporary photos, videos and sounds to a specific geo location on a map of the world.  Well it’s actually not just any map, it’s a Google map and this is likely to make all the difference. It’s a combination of a crowdsourcing project (they want organisations and individuals to load content), a useful educational site, and a service that libraries and archives can hook into to expose their content and collections to new audiences (similar to Flickr Commons).  I’ve seen quite a few sites like this before, but on a small scale for specific locations. For example Sydney Sidetracks was launched in 2008 by the ABC in partnership with The Dictionary of Sydney, The National Film and Sound Archive, The City of Sydney, The Powerhouse Museum, The State Library of New South Wales and the Museum of Contemporary Art. There is a website and mobile app from which historic images, videos and sound are available for locations in central Sydney overlaid on a map.
The big difference with Historypin is that it has been developed by ‘We Are What We Do’,( a not for profit organisation that creates ways for millions of people to do more small, good things) in partnership with Google. Google is the main technology partner on the project and has helped with Google tools, including Google Maps, Google Street View, Picasa, Google App Engine and Android. Google has supported the development costs of the project with donations and sponsorship.  It has also given marketing support and created the video to promote the service:  a one minute introduction to Historypin. This means this is not some small scale project that may suffer from lack of budget, development, maintenance or marketing.  It is something likely to be around for a while and perhaps rival Flickr Commons. Google says “We share ‘We Are What We Do’s commitment to Historypin as a non-commercial, collaborative project that delivers social impact and contributes to digital inclusion.”
The marketing blurb says “Historypin is a way for millions of people to come together, from across different generations, cultures and places, to share small glimpses of the past and to build up the huge story of human history through a well-known medium - picture.”
Relevance for libraries and archives: Interestingly although the initial crowd Historypin were trying to attract was the public to contribute their photos and stories, it now appears that the crowd may actually be the libraries and archives community. This community has massive amounts of digitised content in image, video and sound format, and they want it more widely exposed, tagged, and used.  A service in which libraries and archives can do this, which they don’t have to develop and support themselves, and has no geographical boundaries is certainly a drawcard.  Batch upload has already been enabled, as has ‘make your own collection’ and ‘view slideshow’.  You can pin your content on any Google Street View scene, in any country of the world.  If you happen to be somewhere that Street View hasn’t yet been – don’t worry you can still pin your content down. It is a service that will be more valuable the more content there is.  I only wonder if they have under-estimated the interest that libraries and archives will have in joining, and the volume of content they will have.  If so it is advisable to get in early in case there is a three year waiting list like Flickr Commons had when it started. This is a crowdsourcing project that has a direct relevance to libraries and archives, no matter what their size or where they are located.
TEDx video: Nick Stanhope on mapping history  

4.       Ancient Lives  – Decoding Papyri Launched July 2011
About the project: The Ancient Lives project presents you with fragments of 1,000-year-old papyri to decode. The papyrus was discovered by researchers from Oxford University over a century ago in Oxyrhynchus (the city of the long-nosed fish).
With about 100 men from the local village, Grenfell and Hunt dug in the high winds roaring across the desert. In early January of 1897 a papyrus containing the apocryphal Gospel of Thomas was unearthed, and then a fragment of St. Matthew’s Gospel. The flow of papyri began. Within a few years not only Thucydides and Plato were delicately pulled from the sand, but also Greek lyric poetry that had not been seen or read in about 1000 years. Further, the private documents of this vanished city were collected en masse: private letters, accounts, wills, marriage certificates, land leases, etc. Ancient garbage became a modern treasure. By 1907 the digging ceased. 700 boxes of papyri, potentially carrying about 500,000 fragments, made the long journey back to Oxford University, where Grenfell and Hunt opened up a new branch of study: papyrology. A little over a century later, only a small percentage has been translated by scholars. The Oxyrhynchus collection is owned and overseen by the Egypt Exploration Society.”
The papyrus can be decoded easily by volunteers who match known characters from a grid to the unknown characters on the fragment.  Fragments can be matched by adding measurements of the fragments and the columns within them. The task is mammoth and before the arrival of the online tool could only be undertaken by scholars who were familiar with the code. A very difficult task has been effectively simplified, whilst retaining the challenge that is found in crosswords or code-breaking.
The project was launched in July 2011 and is part of the the Citizen Science Alliance, which is a transatlantic collaboration of universities and museums who are dedicated to involving everyone in the process of science. Growing out of the wildly successful Galaxy Zoo project, it builds and maintains the Zooniverse network of crowdsourcing projects, of which Ancient Lives is one of the newest. Nearly half a million people are contributing to the Zooniverse crowdsourcing projects.
Relevance for libraries and archives: This is a good example of a task that appears on the surface to be too difficult and extensive for a crowd to undertake.  By clever breaking down of the task and designing a simple user interface it becomes achievable.  It also demonstrates that private information about people is of eternal interest to the public. This project along with all the other Zooniverse projects has extensive public discussion forums to firstly foster the volunteer community and secondly let them know how their work helps new discoveries and knowledge grow and develop. We can learn much from how Zooniverse treats its volunteer community.
5.       Duolingo -  translate the web and learn a new language Launched November 2011
About the project: Luis Von Ahn of the Carnegie Mellon University is the creator of CAPTCHA and reCAPTCHA. Google bought both and reCAPTCHA has effectively helped Google Books improve the OCR in its digitised books word by word. Each year 750 million people are unwittingly converting the equivalent of 2.5 million books by using reCAPTCHA.  This is a crowdsourcing project where people don’t realise they are in a crowd or what they are doing. Luis is now working on a new project: Duolingo.  Luis says “Before the internet the biggest projects had 100,000 people involved and with that you could for example put a man on the moon.  My question is what can you achieve with the internet when you can have 100 million people working together on something?”  A good question.  Especially when you combine the number of people with all that ‘cognitive surplus’ that Clay Shirky is always talking about.
Duolingo will help people learn a new language and simultaneously (unwittingly) translate the Web.  He says “It is estimated that there are over 1 billion people learning a foreign language at any given time”. OK so this means a big potential crowd. The Google translator tool is quite good at translating websites but not as good as he thinks the new project Duolingo will be.  The site went live in beta mode in November 2011, but only a few road testers have been accepted.  There is a waiting list of 100,000 who want to join the site already. Luis says “Duolingo is a 100% free language learning site in which people learn by helping to translate the Web. That is, they learn by doing.” The difference to reCAPTCHA is that people will know what they are doing and consciously want to do it. Watch this space.
Relevance for libraries and archives: I’m not sure what the relevance for libraries and archives will be.  Although reCAPTCHA is a free program that is obviously very relevant for libraries and archives it has only been utilised by commercial companies so far, namely the New York Times historic newspaper archive and Google Books. No library has utilised it. I thought I should mention the new project Duolingo since the potential also seems big.  It’s a good idea to translate the web, but I also like the idea of something Luis didn’t mention which is translating books and newspapers into different languages. A question that the National Library of Australia was thinking about last week was “will our volunteer newspaper text correctors be as keen to correct Australian newspapers in foreign languages as they are the English ones? Will they correct them even if they don’t speak the language?” We are asking this because we will soon be adding Australian newspapers in foreign languages to Trove. If this content is classed as ‘part of the web waiting to be translated’, then I guess Duolingo holds big relevance for all national libraries. Duolingo is at an early stage of development so we will have to wait and see. That is unless libraries want to be really pro-active and actually make suggestions to the development team for things that would help them make their content more widely accessible and used……
The TEDx video:  Luis talking on CAPTCHA, reCAPTCHA and Duo-lingo

I hope you find some inspiration from these five crowdsourcing sites for your library, archive or museum.  If there is a newish site of relevance to libraries and archives that you think I’ve missed please add a comment to this post and share. Crowdsourcing sites I have previously reviewed are:
·         Picture Australia (National Library of Australia)
·         FamilySearchIndexing (Church of Latter Day Saints)
·         Distributed Proofreaders (contributes to Project Gutenberg)
·         Wikipedia  
·         UK MP's Expenses (The Guardian)
·         Galaxy Zoo  (Citizen Science Alliance)
·         BBC WorldWar2 Peoples War (BBC)
·         Digitalkoot (National Library of Finland)
·         Old Weather (National Maritime Museum and Citizen Science Alliance)
·         Remember Me: Displaced Children of the Holocaust (United States Holocaust Memorial Museum)
·         Trove Australian Newspapers (National Library of Australia)
·         Transcribe Bentham (University College of London)
·         Waisda (Netherlands Institute for Sound and Vision)

Read more - related posts by Rose Holley on crowdsourcing:
·         Gold star to text correctors for e-books, 13 December 2011
·         Software for journal and newspaper text correction, 18 December 2011
·         Digital cultural heritage awards for crowdsourcing, 4 February 2012

In March 2011 images of the digitised Australian Women's Weekly 1932- 1984 were projected onto the National Library of Australia building as part of the ‘Enlighten’ Festival in Canberra. Nearly 395,000 articles from the Australian Women's Weekly can be improved by public text correction in Trove.  Photograph by Paul Hagon.

Saturday, 4 February 2012

Digital Cultural Heritage Awards for Crowdsourcing (and thoughts on gamification)

Towards the end of last year I received an exciting letter informing me that Trove/Australian Newspapers had been nominated for an international crowdsourcing award for the text correction activity. I had never heard of the award before but the letter explained it:

 “The Digital Heritage Award is an initiative of the Dutch Institute for Cultural Heritage and the Digitaal Erfgoed Nederland Foundation (DEN). First introduced in 2008, it has since been awarded annually to a heritage institution or project that has used digital heritage in an innovative or inspiring way. In this year’s edition the award will go to the best digital heritage related crowdsourcing project.”

I considered it a great honour to be nominated and short listed.  Several years of my life have been committed day and night to developing, maintaining and promoting Australian Newspapers which is ‘my baby’. The five shortlisted nominees had been selected by a jury, consisting of five international experts on crowdsourcing and heritage. The jury consisted of Susan Hazan (Director of Digital Heritage UK and Curator of New Media and Head of the Internet Office at Israel Museum), Johan Oomen (Head of R&D at the Netherlands Institute for Sound and Vision), Josh Greenberg (Director of Alfred P. Sloan Foundations Digital Information Technology Programme in New York), Vincent Puig (Executive Director at IRI/Centre Pompidou) and Mia Ridge (UK Researcher, consultant, programmer, analyst).

To be shortlisted the crowdsourcing projects had to meet the following selection criteria:
·         Be in an advanced or finished state of development.
·         Hundreds or thousands of members of the public should have contributed.
·         Significant results already achieved in 2010 or 2011 and publicly visible.
·         Results exceeded expectations, are inspiring to others, and can be replicated.
·         Have had press coverage.
·         Have a clear project leader who can present the project at the awards ceremony.

The shortlisted projects also had the following criteria applied:
·         Long-term commitment to the activity
·         Continued progress
·         Motivation and rewards for the crowd
·         Effective design
·         Link to existing communities

This led to a list of five finalists for the Digital Heritage Award 2011:
  • Digitalkoot from the National Library of Finland. 50,000 volunteers are correcting OCR newspaper text to 99% accuracy.
  • Old Weather from the National Maritime Museum, UK. 700,000 - 97% of navy handwritten ships logs with temperatures have been transcribed by thousands of volunteers harnessed in Galaxy Zoo.
  • Remember Me: Displaced Children of the Holocaust from the Unites States Holocaust Memorial Museum. 61,000 people have viewed 1000 pictures of children lost in the Haulocaust. So far 180 have been identified and traced.
  • Trove Australian Newspapers from the National Library of Australia. 40,000 volunteers have corrected 51 million lines of OCR text in historic newspapers making them more searchable.
  • Transcribe Bentham from the University College of London. Volunteers subscribe 44 handwritten Bentham manuscripts per week.
The winner would be selected by audience voting. Over 500 digital culture heritage specialists attending the international conference Digital Strategies for Heritage (DISH2011) would watch presentations on the crowdsourcing projects, speak to the project leaders and then vote for their favourite on the first day of the conference - 7 December 2011.

So, rather belatedly I’m now going to tell you what happened next…

Unfortunately the National Library of Australia decided that it could not justify the cost of sending me to Holland for the conference.  This was perhaps rightly so since it would have cost over AU$5000 for me to attend and the Library is making travel cutbacks at a time of severe financial restraint.  The conference being in Europe was rather pricey, but did present a good professional development opportunity as well as the chance to win an award, and for me to meet face to face the other project managers.  Not attending or being able to present in person immediately reduced our chances of winning.  The conference organisers kindly let me send a video message instead. As it turned out there was only one conference attendee from Australia, and one from New Zealand, but a very strong contingent from Scandanavia.  Because the winner was based on audience voting, and on occasions such as this national pride and alliances run strong, things seemed to be against us from the start.

The winner who had a straight lead to the finishing post was DigitalKoot. I congratulate them and all the other nominees. It was a very hard choice to make with each project being really good.  Maybe that’s what the judges also thought, who came up with the idea of audience voting (devolving the responsibility!)

In discussions with people afterwards and by following the conference online I was interested to pick up that many of the digital culture specialists still seemed to think that you would stand little chance of getting thousands of volunteers to do something for you unless you made it into a game and it looked cool, hence perhaps their enthusiasm for DigitalKoot (a game to correct newspaper text that involves a mole).  This caused me pause for thought, because I don’t think I agree with this view, but then maybe I have got the whole thing wrong? It is also interesting that the year before in 2010 the Best Archives on the Web: Best Use of Crowdsourcing for Description’ Award was given to Waisda (What’s that) a Dutch project from Netherlands Institute of Sound and Vision. It also uses gaming technology. People tag videos with subjects and see if they match other peoples.

I realised there are some things we discussed on the Australian Newspapers project which I have never written about in my articles or mentioned in my presentations, which now seem very pertinent. Firstly when we began to design the Australian Newspapers site we employed a web design company who could not think why anyone would want to correct newspaper text unless they made it into a game. But the primary purpose of our newspaper project was to digitise and make online available for free and full-text searchable Australian Newspapers. The bit on the side was that it would be good to improve the quality of the text for searching if we had the time and means. Hence we never had a ‘crowdsourcing project’ and we never focused on that.  We told the web developers to focus first on getting the search and browse interface up and running and leave the text correction bit until last if they had time. They did this.  When doing public usability testing for ‘search and browse’ the developers were overwhelmed by the excited response they got from people off the street about the availability of Australian Newspapers.  None of these users were ‘library users’ or really considered that this was a library service. They all showed early signs of getting quite sucked into search and browse.  All the testers had no problem thinking up something they would like to look up in old newspapers.  Based on this high level of interest and motivation the developers thought maybe a ‘gaming strategy’ to attract users would not after all be required.  Also the library project team felt that anyone who wanted to improve the text would come from the user base i.e. newspaper searchers, so they would be in the site already and have an interest in improving the text of something they had just read. Lastly the image of the newspaper text was always visible so the need to match or verify someone’s corrections with someone else’s before accepting them was not really necessary, a strategy often employed in gaming technology.

The simple, explicit thing I have never said is that as far as we know our volunteer text correctors are a subset of our Trove search user base, not a separate group of people who simply want to crowdsource. That is they do not think “Oh I want to help with crowdsourcing, let’s find a site that does that”, they are already in our site thinking “this is a great site, I found what I wanted, oh look I can make it even better, I’ll do that to help”. Obviously some of the text correctors are doing vast amounts of work, but most are simultaneously undertaking research using the resource.  They seem to find both activities highly enjoyable and addictive without ‘gamification’

I’ve had so many people contact me and say “How can I set up a crowdsourcing project?” but we never came to it from that point of view and I don’t think that is how you should.  It is the wrong question to ask. You need to ask “What goal do I want to achieve, and how can I do that?” Ours was to improve the quality of our searching, and it happened that our solution was to get the public to help with this which became ‘crowdsourcing’. For the National Library of Australia the crowdsourcing activity is a side effect of its ongoing effort to deliver high quality services. We play it down actually, and never use the word ‘crowdsourcing’. We just say the public are helping us, or that the community is involved, or volunteers work online.  There is still great reticence about using the  “C” word itself, acknowledging the scale of the activity or the activity itself.  Three years in senior managers finally agreed we could put some text on the home page of Trove saying ‘contribute’ and ‘how to correct text’.  Before this a Trove user would only stumble across the fact they could correct text when they had actually reached the point in the newspaper search where they could do it. The Library has never formally appealed for the community to help, or had a strategy to do this, though it has acknowledged and congratulated the highest achieving volunteers.  

So firstly I am thinking that even with this lack of public appeal our results have been phenomenal. We have drawn on our existing Trove user base (5 million) and about 40,000 people have become volunteer text correctors, about 4,000 of them hugely committed and correcting each week. They have improved 56.8 million lines of text.  But then on the other hand I’m thinking “Would this have exponentially increased if we had used gaming technology from the start, targeted gamers rather than searchers, and introduced an animal like DigitalKoot have done?”  Of course our animal could not have been a mole it would have to have been a native – a kangaroo, koala, possum, sugar glider, or my favourite a bilby. Maybe our initial decision was wrong after all.  

There’s really not much written or researched on this topic.  In fact the term ‘gamification’ is only about 18 months old. There is quite a lot of negativity towards ‘gamification’. It is considered by some to be a stupid fad that will soon pass, and conversely others see it as important as the rise of social media. The ABC’s creative director of strategy thinks it is as stupid as it sounds because it limits creativity and goals.  Viewpoints on whether to use gaming technology or not on digital cultural heritage crowdsourcing sites, may also depend on what your goals are and whether or not you think the journey i.e. the level of social engagement and community building is as, or more important than the destination i.e. the result.  At the National Library of Australia I think it is fair to say we consider the journey ‘interesting’ but really we have our eyes on the destination only. As I said our results are phenomenal, but maybe they can be improved and increased.  I wonder if someone could do some research on this, or alternatively we could just change our interface, add a few different levels of competence and a couple of bilby’s and see what happens….

Photos of DISH 2011 from the DEN flickr stream
 Voting
 Finland's DigitalKoot receives the DISH2011 Crowdsourcing Award