Showing posts with label OCR. Show all posts
Showing posts with label OCR. Show all posts

Tuesday, 2 October 2012

Digital Motor Archive available free for the month of October in return for…...


It came to my notice last week that a publisher in the UK had digitised both the current and back issues of their magazine the ‘Commercial Motor’ which in its early life was a newspaper. In the cut throat world of publishing there are few publishers left that are still publishing the same title they were 100 years ago, and who also have a complete set of back copies.  If they fall into this bracket they are in the unique position of being able to either digitise the content themselves for their readers (usually at a loss); offer it to a library to digitise (at no cost); or sell it to a commercial e-vendor to package with another product for academia (and make a profit).  Unfortunately most choose the latter, which makes this type of content only really accessible to academics and students via academic libraries.  E-vendors normally charge high subscription rates to digitised magazines and newspapers and package them up with other content, making it only viable for large universities and national libraries to purchase, and therefore severely restricting readership to the content.

Not a lot of publishers are digitising their own content because generally speaking the cost of preparation, digitisation and OCR, and building a good website to deliver the content outweigh the amount of money they would ever recuperate from reader subscriptions. Normally a subscription to the current copy would be packaged up with old copies.  And that’s where the model fails, because often current readers have no interest in the old stuff.  People who do have an interest in the old stuff are generally a different group of people – historians, researchers etc.  The exception to this rule appears to be anything to do with hobbies such as knitting, cooking, railways, cars, and stamps. 

The Commercial Motor Archive http://archive.commercialmotor.com/ came to my notice because it is available for free for the month of October and I wondered why. It is a rich archive going from 1905 to the present day, covering a complete century and two world wars, well illustrated and with everything you ever wanted to know about commercial vehicles.   

The search and browse mechanism is very impressive and works well.  For example articles on pages have been zoned so you can search and find the article easily within a page.  It has many similarities to the hugely popular Australian Newspapers http://trove.nla.gov.au/newspaper.  The page displays alongside the OCR text to make it easier to read. You can browse covers, browse by date, and zoom in on pages.  Results can be filtered. Users can add comments and tags.  The quality of the OCR text and therefore search is very good.

It has one thing that was never implemented on Australian Newspapers (though often asked for by users) which is a little box on each page called ‘Report an error’ and this is the reason for its free access.  The site owner is hoping that as people use and read the pages in the archive they will report errors as they see them, and for this they get free access to the content.  The known errors that need to be identified are incomplete articles (where the zoning has gone wrong); OCR text error in headlines of articles; and OCR errors in text. However readers can only report them, not actually fix them.  The site states:

“the archive is beta because it isn’t perfect at the moment and there are a few glitches to be ironed out. Every article page has a 'Noticed an error?' button you can use to report a problem. Please don’t expect an immediate change to the error - we gather all the reports together and prioritise them, fixing the most pressing errors first.”

It sounds like they don’t know how many errors there are, how many people will report errors, and who and how the errors will be fixed.  Interesting.

Although the ‘report an error’ button was never implemented on Australian Newspapers (it was mainly needed to report upside down and duplicate pages) it had already been decided that a ‘super user’ would be the person to review these reports and take action.  In a world where some volunteer text correctors wanted to take on extra responsibility and have special roles like the hierarchy in the Wikipedia editors community this would have been a good thing for trusted volunteers to do.

The Commercial Motor Archive has impressed me because they are clearly striving for perfection, they understand that the fewer mistakes there are the better the search will be and they have taken a brave step by asking the public to help in return for free access.  This is indeed unusual for a commercial publisher, belonging more in the realms of libraries and archives and referred to as crowdsourcing……

 

Sunday, 11 March 2012

Crowdsourcing transcription of handwritten archives


One of the big differences between libraries and archives is that libraries tend to have more of ‘the printed word’ whilst archives have vast amounts of ‘the handwritten record’.  While some libraries are getting up to speed with mass digitisation of books and journals and then being able to offer users full text searchable digitised items, this is still a distant dream for most archives.  Some archives are undertaking mass digitisation, but the second step – making handwritten records full-text searchable is a massive challenge.  The reason for this is in the technology and processing steps.

After scanning a ‘printed word’ page into an image file a piece of software called Optical CharacterRecognition (OCR) converts the image into searchable text.  The OCR works best with clean, clear, black and white typeface such as a word document or a book, not quite so well on old books and journals, and very poorly on old newspapers.  When it comes to converting handwriting it fails miserably.  It just can’t distinguish and convert handwriting to text in the way the human eye can.  Therefore archives can’t easily automate the second part of the digitisation process using OCR software like libraries can for the printed word.

If you at least get some OCR text from print that is readable and therefore searchable you can offer a service to users to full-text search the books or journals such as Google does. If the OCR text is poor there are some things you can do to improve it. You can encourage users of your service to correct the OCR text with a text correction tool so that the searching is improved, such as Trove does with the Australian Newspapers.

Unfortunately the only viable option open to archives to convert digital images into full-text searchable text is to use a manuscript transcription tool, in combination with harnessing the power of a crowd to do the transcription work.  The transcription work for handwritten records is much harder than for example text correcting old newspapers because the handwriting is often difficult to read, old fashioned, barely legible and not necessarily structured in lines or columns. There is often nothing to go on.

I recently stumbled across a blog all about manuscript transcription tools that is written by a software developer Ben W Brumfield in Texas. Ben developed his own software to transcribe his great-great grandmother’s journal. ‘FromthePage’ is now being used by archives because Ben has made it available open source.

A year ago he wrote an in-depthblog post that covered manuscript transcription tools under development, manuscript transcription projects in archives, and made some predications for future directions of manuscript transcription.  I am not going to repeat what he said here, I suggest you read the post in full.  He notes that software development in this area is still fragmented and young with no particular tools taking dominance. Most developed applications are being made available open source. A standout is ‘Scribe’ from the Zooniverse team, currently being used by both the ‘Old Weather’ project to transcribe maritime weather records and by ‘What’s the score’ project to transcribe music scores at the Bodleian Library, Oxford.

Before an archive implements a manuscript tool it needs to find out what it’s users  would most like to be easily full-text searchable from the vast vaults of all the content it has. It is important to find this out, because the crowd will only be motivated and swell in numbers if they really feel what they are doing is very important to a broad group of people and really matters either right now, or in the long-term and is also interesting.  They have to feel this before they will join in.  Once they have joined in there are other motivational tips you can do to keep them going.  Just implementing a manuscript tool is simply not enough.  You need to engage, watch, understand and learn from your crowd, for they hold the passion and power in their hands to make your project successful or not.


Photo by Rose Holley, outside Canberra Bus Station

Saturday, 4 February 2012

Digital Cultural Heritage Awards for Crowdsourcing (and thoughts on gamification)

Towards the end of last year I received an exciting letter informing me that Trove/Australian Newspapers had been nominated for an international crowdsourcing award for the text correction activity. I had never heard of the award before but the letter explained it:

 “The Digital Heritage Award is an initiative of the Dutch Institute for Cultural Heritage and the Digitaal Erfgoed Nederland Foundation (DEN). First introduced in 2008, it has since been awarded annually to a heritage institution or project that has used digital heritage in an innovative or inspiring way. In this year’s edition the award will go to the best digital heritage related crowdsourcing project.”

I considered it a great honour to be nominated and short listed.  Several years of my life have been committed day and night to developing, maintaining and promoting Australian Newspapers which is ‘my baby’. The five shortlisted nominees had been selected by a jury, consisting of five international experts on crowdsourcing and heritage. The jury consisted of Susan Hazan (Director of Digital Heritage UK and Curator of New Media and Head of the Internet Office at Israel Museum), Johan Oomen (Head of R&D at the Netherlands Institute for Sound and Vision), Josh Greenberg (Director of Alfred P. Sloan Foundations Digital Information Technology Programme in New York), Vincent Puig (Executive Director at IRI/Centre Pompidou) and Mia Ridge (UK Researcher, consultant, programmer, analyst).

To be shortlisted the crowdsourcing projects had to meet the following selection criteria:
·         Be in an advanced or finished state of development.
·         Hundreds or thousands of members of the public should have contributed.
·         Significant results already achieved in 2010 or 2011 and publicly visible.
·         Results exceeded expectations, are inspiring to others, and can be replicated.
·         Have had press coverage.
·         Have a clear project leader who can present the project at the awards ceremony.

The shortlisted projects also had the following criteria applied:
·         Long-term commitment to the activity
·         Continued progress
·         Motivation and rewards for the crowd
·         Effective design
·         Link to existing communities

This led to a list of five finalists for the Digital Heritage Award 2011:
  • Digitalkoot from the National Library of Finland. 50,000 volunteers are correcting OCR newspaper text to 99% accuracy.
  • Old Weather from the National Maritime Museum, UK. 700,000 - 97% of navy handwritten ships logs with temperatures have been transcribed by thousands of volunteers harnessed in Galaxy Zoo.
  • Remember Me: Displaced Children of the Holocaust from the Unites States Holocaust Memorial Museum. 61,000 people have viewed 1000 pictures of children lost in the Haulocaust. So far 180 have been identified and traced.
  • Trove Australian Newspapers from the National Library of Australia. 40,000 volunteers have corrected 51 million lines of OCR text in historic newspapers making them more searchable.
  • Transcribe Bentham from the University College of London. Volunteers subscribe 44 handwritten Bentham manuscripts per week.
The winner would be selected by audience voting. Over 500 digital culture heritage specialists attending the international conference Digital Strategies for Heritage (DISH2011) would watch presentations on the crowdsourcing projects, speak to the project leaders and then vote for their favourite on the first day of the conference - 7 December 2011.

So, rather belatedly I’m now going to tell you what happened next…

Unfortunately the National Library of Australia decided that it could not justify the cost of sending me to Holland for the conference.  This was perhaps rightly so since it would have cost over AU$5000 for me to attend and the Library is making travel cutbacks at a time of severe financial restraint.  The conference being in Europe was rather pricey, but did present a good professional development opportunity as well as the chance to win an award, and for me to meet face to face the other project managers.  Not attending or being able to present in person immediately reduced our chances of winning.  The conference organisers kindly let me send a video message instead. As it turned out there was only one conference attendee from Australia, and one from New Zealand, but a very strong contingent from Scandanavia.  Because the winner was based on audience voting, and on occasions such as this national pride and alliances run strong, things seemed to be against us from the start.

The winner who had a straight lead to the finishing post was DigitalKoot. I congratulate them and all the other nominees. It was a very hard choice to make with each project being really good.  Maybe that’s what the judges also thought, who came up with the idea of audience voting (devolving the responsibility!)

In discussions with people afterwards and by following the conference online I was interested to pick up that many of the digital culture specialists still seemed to think that you would stand little chance of getting thousands of volunteers to do something for you unless you made it into a game and it looked cool, hence perhaps their enthusiasm for DigitalKoot (a game to correct newspaper text that involves a mole).  This caused me pause for thought, because I don’t think I agree with this view, but then maybe I have got the whole thing wrong? It is also interesting that the year before in 2010 the Best Archives on the Web: Best Use of Crowdsourcing for Description’ Award was given to Waisda (What’s that) a Dutch project from Netherlands Institute of Sound and Vision. It also uses gaming technology. People tag videos with subjects and see if they match other peoples.

I realised there are some things we discussed on the Australian Newspapers project which I have never written about in my articles or mentioned in my presentations, which now seem very pertinent. Firstly when we began to design the Australian Newspapers site we employed a web design company who could not think why anyone would want to correct newspaper text unless they made it into a game. But the primary purpose of our newspaper project was to digitise and make online available for free and full-text searchable Australian Newspapers. The bit on the side was that it would be good to improve the quality of the text for searching if we had the time and means. Hence we never had a ‘crowdsourcing project’ and we never focused on that.  We told the web developers to focus first on getting the search and browse interface up and running and leave the text correction bit until last if they had time. They did this.  When doing public usability testing for ‘search and browse’ the developers were overwhelmed by the excited response they got from people off the street about the availability of Australian Newspapers.  None of these users were ‘library users’ or really considered that this was a library service. They all showed early signs of getting quite sucked into search and browse.  All the testers had no problem thinking up something they would like to look up in old newspapers.  Based on this high level of interest and motivation the developers thought maybe a ‘gaming strategy’ to attract users would not after all be required.  Also the library project team felt that anyone who wanted to improve the text would come from the user base i.e. newspaper searchers, so they would be in the site already and have an interest in improving the text of something they had just read. Lastly the image of the newspaper text was always visible so the need to match or verify someone’s corrections with someone else’s before accepting them was not really necessary, a strategy often employed in gaming technology.

The simple, explicit thing I have never said is that as far as we know our volunteer text correctors are a subset of our Trove search user base, not a separate group of people who simply want to crowdsource. That is they do not think “Oh I want to help with crowdsourcing, let’s find a site that does that”, they are already in our site thinking “this is a great site, I found what I wanted, oh look I can make it even better, I’ll do that to help”. Obviously some of the text correctors are doing vast amounts of work, but most are simultaneously undertaking research using the resource.  They seem to find both activities highly enjoyable and addictive without ‘gamification’

I’ve had so many people contact me and say “How can I set up a crowdsourcing project?” but we never came to it from that point of view and I don’t think that is how you should.  It is the wrong question to ask. You need to ask “What goal do I want to achieve, and how can I do that?” Ours was to improve the quality of our searching, and it happened that our solution was to get the public to help with this which became ‘crowdsourcing’. For the National Library of Australia the crowdsourcing activity is a side effect of its ongoing effort to deliver high quality services. We play it down actually, and never use the word ‘crowdsourcing’. We just say the public are helping us, or that the community is involved, or volunteers work online.  There is still great reticence about using the  “C” word itself, acknowledging the scale of the activity or the activity itself.  Three years in senior managers finally agreed we could put some text on the home page of Trove saying ‘contribute’ and ‘how to correct text’.  Before this a Trove user would only stumble across the fact they could correct text when they had actually reached the point in the newspaper search where they could do it. The Library has never formally appealed for the community to help, or had a strategy to do this, though it has acknowledged and congratulated the highest achieving volunteers.  

So firstly I am thinking that even with this lack of public appeal our results have been phenomenal. We have drawn on our existing Trove user base (5 million) and about 40,000 people have become volunteer text correctors, about 4,000 of them hugely committed and correcting each week. They have improved 56.8 million lines of text.  But then on the other hand I’m thinking “Would this have exponentially increased if we had used gaming technology from the start, targeted gamers rather than searchers, and introduced an animal like DigitalKoot have done?”  Of course our animal could not have been a mole it would have to have been a native – a kangaroo, koala, possum, sugar glider, or my favourite a bilby. Maybe our initial decision was wrong after all.  

There’s really not much written or researched on this topic.  In fact the term ‘gamification’ is only about 18 months old. There is quite a lot of negativity towards ‘gamification’. It is considered by some to be a stupid fad that will soon pass, and conversely others see it as important as the rise of social media. The ABC’s creative director of strategy thinks it is as stupid as it sounds because it limits creativity and goals.  Viewpoints on whether to use gaming technology or not on digital cultural heritage crowdsourcing sites, may also depend on what your goals are and whether or not you think the journey i.e. the level of social engagement and community building is as, or more important than the destination i.e. the result.  At the National Library of Australia I think it is fair to say we consider the journey ‘interesting’ but really we have our eyes on the destination only. As I said our results are phenomenal, but maybe they can be improved and increased.  I wonder if someone could do some research on this, or alternatively we could just change our interface, add a few different levels of competence and a couple of bilby’s and see what happens….

Photos of DISH 2011 from the DEN flickr stream
 Voting
 Finland's DigitalKoot receives the DISH2011 Crowdsourcing Award