Showing posts with label archives. Show all posts
Showing posts with label archives. Show all posts

Thursday, 19 May 2022

Australians at War Film Archive- Project Stage 6 Completed-Full digital access.




I just wanted to say how happy I am that the new website for the Australians at War Film Archive is now available. The Archive project is a collaboration between UNSW Canberra and the Department of Veteran Affairs.  As project manager I have been honoured to work with a fantastic IT team who have helped realise the long-term vision of Michael Caulfield, the original instigator, project manager and film producer, 23 years ago, and still a great supporter of the Archive. 

The migration into a newly developed content management system; transcoding of the digital files into new formats; and a complete re-write of the delivery public system means that the Australians at War Film Archive is preserved, accessible and usable in new ways.  Researchers and family members of the interviewees can now search, view, save and download the interviews themselves, without needing mediation, help or permission from our team of archivists, which makes everyone happy. Michael Caulfield's vision of a national, open, freely accessible archive has reached the next level. We have been overwhelmed by the positive response from researchers and interviewee's relatives, particularly those family who have found their relatives interviews by surprise. One of the most useful new functions we added is the ability to easily make and save a clip. This has been widely used by museum curators for exhibitions, relatives for family events, school children for projects, and lecturers for presentations.

The archive project has spanned 23 years so far, and has involved over 500 project staff and 2,000 participants sharing their stories. The dedication and commitment of all is admirable. Having listened to several of the interviews and extracts, which often run for 5-9 hours each I am struck by the broad range and wealth of content, which was previously not easily accessible. It is wrong to assume the archive is only about war because of its branding and name. Recording the experiences of Australians in war and conflict was the driving force, but it is more accurate on reflection to say that the interview corpus covers the social history of Australia from Victorian times to present.

Interviewees describe in detail their lives, upbringing, families and places lived both before and after their war experiences.  The most fascinating information is offered up to researchers as the interviewees chat away. One man described his early swimming exploits as a young lad in hand knitted swimwear, and in detail the meals his mother made him for tea when he went home afterwards. These interviews give snapshots of not only people,  but also the places they grew up in, or passed through, with many small rural regional towns getting mentions. Each interview has a transcript which is full text searchable and synced to the audio and video recording. For example if you search on 'Queanbeyan' you will find 56 different interviewees describing their experiences in this town. It is good to hear some of their light hearted, fond memories from childhood and pre and post war as well as more difficult ones.  It goes without saying that the war recollections are an invaluable original research resource for military historians, holding great depth and context, more so as time passes by. 

I always describe the transcripts as being the 'gold' since they are what enable the full text search of the film/video/digital files. But more than this, the transcripts have had context and value added to the war recollections by a team of eminent historians over several years. We await with interest feedback from historians, researchers, authors and post graduate students on how they have used and interpreted the data. 

If you want to know more about how the project was carried out from 1999 to now, who was involved, and how it was planned a longer overview is now available.  









   

Friday, 29 August 2014

Audiovisual achievements - National Archives of Australia


I always get a sense of achievement from a job well done, and this week the audiovisual IT project I have managed at the National Archives of Australia (NAA) over the last 2 years has reached fruition – on time and under budget, which makes the achievement even better. 
The project was a big one costing several million and was the implementation of both an audiovisual asset management system, and an audiovisual digital preservation system.   It has been a long held ambition of the NAA to achieve these two goals. The concept crystallised into a firm plan in 2006.  Implementation commenced in 2012 and the project became the highest strategic objective of the NAA for the next two years, involving approximately half of the 400 NAA staff in some capacity. The Chester Hill office at Sydney took the lead because this office is the centre of expertise for audiovisual collections.  I feel fortunate to have had the opportunity to work with such a fantastic and knowledgeable group of people.

The project which is known internally at NAA as ‘AVAMS’ (audiovisual asset management system project), and its achievements is described in more detail in my AVAMS presentation available on slideshare.

The chosen software that has been implemented is Mediaflex from a UK based company called TransMedia Dynamics.  The National Archives is the second Archives client to install the Asset Management Software as the Collection Management System for both physical and digital audiovisual assets, and the first client in the world to install the Mediaflex digital preservation platform.  Other Australian clients include the National Film and Sound Archive, and DAMsmart an audiovisual digitisation contractor.
 
The project has been important to the NAA because firstly audiovisual is a significant part of the collection amounting to nearly 1 million items, and secondly there is a need to increase capability and capacity to ingest born digital audiovisual from transferring agencies.  One of the main agencies transferring audiovisual material to the NAA is the Australian Broadcasting Corporation (ABC) who creates all radio and TV programs digitally now and has done so for some time.  Because older parts of the NAA audiovisual collection are analogue, and these formats deteriorate quickly there has been an active and ongoing NAA audiovisual digitisation program to convert analogue to digital formats for at least the last 10 years in state-of-the-art digitisation labs onsite at Sydney.
This youtube video gives a small glimpse behind the scenes at the Sydney Office, and a sample of a very  deteriorated analogue film now digitised is available to view on youtube in 'a cautionary tale'.  
For these two reasons the NAA already holds a sizable store of digital AV assets. These are now being migrated into the digital preservation system ‘the AV Digital Archive’, which will replace the previous rather clunky and very slow system that was based on a system backup procedure.  It will give increased surety that important digital assets are secure and preserved into the future. It is a giant leap forward to have a robust and easy to use digital preservation system.  The screenshot below shows the console that an archivist would use to manage the digital preservation copies.  The traffic light system is particularly easy to use.






The requirements to manage audiovisual digitisation workflows, storage of physical items, ingest and digital preservation of items are much more complex than those for other format types such as photographs or paper.  In order to better manage and search on collection items a data model with multi-layers is needed.  It is usual in libraries to have 3 data layers, and for archives to have 5 or 6 for paper formats, however in the case of audiovisual the ideal data model has 12 layers.  This is the new model that has now been implemented at the National Archives. It has caused great excitement for those who understand the complexities of audiovisual metadata and realise the benefits this will bring long-term to the management of the collection.  However it has been a steep learning curve for staff to become familiar with the audiovisual data model.
 
An ambition for Archivists has been to expose more of the audiovisual collection to public searchers, because at the moment for various reasons it is largely invisible.  The new data model means that can be changed and improved.  In addition Mediaflex can automatically create low resolution digital access copies on the fly, which brings the potential to make more of the collection digitally available.  There is still more work to be done in this area since RecordSearch is remaining the front end for public searchers for the foreseeable future.  Therefore a fair amount of configuration work has already been undertaken to enable exchange of metadata between the Audiovisual Asset Management System Mediaflex and RecordSearch.   
An immediate benefit that Mediaflex has brought is the ability to much better manage storage of audiovisual items.  These items require repositories of different temperatures e.g. cold and cool, and conditioning rooms between for the gradual movement of items into room temperature for access or digitisation.  In addition there are a variety of different shelving configurations for different sizes and types of items.  Mediaflex allows the management of all this, but in addition 'capacity management'.  A visual interface shows where spare space is and how full shelves are in real time.  This really helps to micro manage over 30 km of audiovisual repository space in multiple locations.

It is rewarding to see how the project achievements - the implementation of an audiovisual asset management system and digital preservation system are already having positive benefits for the NAA. As I reflect on the last 2 years (which feel as if they have passed in the blink of an eye) I attribute the success to the fantastic project team members at both the NAA and Transmedia Dynamics, as well as NAA making the right choice of software. The core project teams contributed their audiovisual expertise and worked diligently under my direction with enthusiasm and total commitment towards the end result.  There is no doubt it was challenging at times, but everyone rose to the challenge with tenacity, determination and persistence.
The National Archives of Australia is now strongly and ably positioned in the audiovisual digital arena.  It has the capability to undertake its core business much better, as well as do groovy and amazing things with the new software.  It’s very unfortunate that the current tight fiscal constraints may now hamper the capacity of the NAA to uptake the new benefits as quickly as it would like, but I am assured it will happen in time. This project achievement has boosted the confidence of the National Archives of Australia and is indeed a job well done!
 Mediaflex in use in the sound preservation lab at National Archives of Australia, Sydney Office.


Saturday, 10 November 2012

National Archives of Australia embraces crowdsourcing and releases ‘The Hive’.


 
The National Archives of Australia (NAA) has made a bold step into the cultural heritage crowdsourcing arena with ‘The Hive’ which was released two weeks ago. The brand makes a clever play on the word ‘Archive’ combined with the idea of a hive of working bees (the public).  The site encourages the public to transcribe archive records.

Early this year when David Fricker became Director General of the NAA he was quick to encourage staff to think innovatively, embrace change, and to harness opportunities such as crowdsourcing to improve access to our collections. He publicly spoke in favour of  crowdsourcing and a changing business model for archives at the International Council of Archives Congress in August:

“Another key development in expanding access is crowdsourcing. As many of us are now seeing, by allowing the public to contribute to the description of archival resources we are enhancing the ability of future generations to discover and learn from our archives. I also think it is a wonderful opportunity for the public to be more engaged with us as archives and to share in the work we do – preserving the memories of our nations. There is still some work to do here, in order to maximise the value of contributions and to maintain the integrity of our archives as authentic and accurate. However, I do not believe these problems are insurmountable, and indeed I believe these systems can to some extent be self-correcting.

This is a type of the co-design, citizen first activity… drawing on the interest and enthusiasm of the community to bring more of our archives into view – discoverable and retrievable…Access will be online and everywhere, improved by rich new data visualisation techniques and expanded descriptive contributions from an engaged citizenry”.

The Hive is the Archives pilot and experimentation into the potential of large scale transcription crowdsourcing to improve access to records.  Staff have looked closely at other crowdsourcing sites on offer and attempted to build on their knowledge and techniques, to provide a site that could be used as a large scale platform for a variety of transcription crowdsourcing projects.

At present the site offers just over 800 lists for the public to transcribe. Some of these are typed and some handwritten.  They are rated in difficulty as easy, medium or hard.  Part of the difficulty with this project is that the public need to have some understanding of how archives receive and describe their records to make sense of what they are being asked to do.  In simple terms archives receive vast amounts of records (referred to as consignments).  Each consignment comes with a list of the items in it.  However because of the large volume of records being received it is usual that only the consignment record is entered into the catalogue e.g. ‘100 boxes of plans and drawings’ from x government agency, rather than all the individual items on the consignment list being described in the catalogue.  The ideal scenario for users of the archives is that every item e.g. plan and drawing is described on the catalogue so that it can be found.  Without this a lot of guess work goes into finding relevant things, or alternatively personal visits are required to view the hard copy consignment lists.
 
The project that the archives is undertaking is to digitise consignment lists and then make them available for transcription by the public. Once transcribed they become searchable and the items within them can be found more easily.  Because so many of the lists are old and handwritten it is virtually impossible to get good OCR on them.  That’s where the public come in who can read them with the human eye. Also the time of the public is needed to speed up the access. Projections on the time it would take archives staff to describe the lists without public help currently stand at 210 years.  It is anticipated that a member of the public could with relative ease describe several hundred items per hour with the Hive tool, which would make a big difference, especially if there was a swarm.

The consignment lists in the pilot are those that have proved most popular with researchers and contain items in the ‘open period’, that is older than 30 years and now open to the public.  The top interest is lists of architectural drawings and historic buildings. This is closely followed by PNG patrol officer records, maritime incidents, personal records from the war office, prisoners of war, meteorology and cyclones, WW1 intelligence, and oil drilling on the Great Barrier Reef.

In the first 2 weeks 300 records have been transcribed of the 800. There is a definite preference for the lists rated hard (handwritten) and ones that involve names.

The site is well presented and gives volunteer transcribers things we know they want such as progress chart, recent activity, points scoring system, rewards, optional login using Open ID e.g. their Google ID, ability to search and choose items, or just take the next one served up, to pick easy or difficult items, to add a marker for where they got to if they are interrupted, and to favourite records.  The only slight drawback is the placing of the transcription window at the bottom of the screen rather than right or left, which often means it is hard to see the transcription window and the content you are transcribing at the same time. Also the OCR text in the transcription window and the cursor is not hooked directly to the text in the image so it is easy to get lost whilst transcribing sometimes.  This is largely because most of the lists are in tables, and the table rows and columns have not been retained in the OCR, so the OCR is somewhat muddled.  Further development of the site will largely depend on feedback given by the public users, and the ability of the archives to keep up a steady supply of new, interesting digitised consignment lists to the Hive.  The Archives is still considering how it may be able to integrate the public content back into its main catalogue RecordSearch, or integrate the Hive into RecordSearch. In the meantime the list content will remain searchable in the Hive.

There is obviously an expectation from the Archives that by making its content more discoverable it will lead to more access requests.  This is why at point of transcription there is a button which enables the user to request a copy of the item.  These requests are being met by digitising the item, and then uploading them into the main catalogue ‘RecordSearch’ with the full item description.

I congratulate the National Archives of Australia Access Team on the development of this exciting new site, which holds so much potential to improve access to records and engage with our citizens in new ways.

The screenshots below show the site in action:



 Easy level transcription- Archived drawings

Medium Level Transcription - ABC Drama Scripts
 


Difficult level transcription - Plans
 


Sunday, 17 June 2012

If only they would crowdsource! – Diamond Jubilee - Royal Archives at Windsor Castle


Many years ago I worked for a software company installing the first archive management systems into large UK archives such as the London Metropolitan Archives, Cumbria Archives at Carlisle Castle and the Royal Archives at Windsor Castle.  It was a challenging time for archives going from paper systems to computer systems, in fact very similar to the challenges archives now face transitioning from managing paper records to born digital records.  Ironically I have just returned again to the archives sector and am now working at the National Archives of Australia on the second challenge.

When the first computer systems were installed in archives it often came as a shock to archivists to discover that when the system was installed it would be ‘empty’ and their records would not somehow miraculously appear in the system. This was the first piece of news I usually had to convey in training before showing an online process for acquisitions. I particularly remember that at the Royal Archives they estimated with their current staff of 4 it would take them 700 years to record their archive collection into their new system, and they were somewhat despondent to say the least. Nevertheless the Queen was pleased with the install of the first computerised system at Windsor Castle and awarded the software company I worked for the Royal Warrant, which meant we could use the Royal Coat of Arms on our letterhead.  The warrant is more often seen on pots of jam and pickle than on software. The celebration of implementation party at Windsor Castle with members of the Royal Household and staff was one to remember. 

The Round Tower at Windsor Castle contained every hand written record every monarch and members of their household had ever created. Queen Victoria’s collection was particularly large.  The Royal Archives could only be contacted by letter and each year less than 10 well vetted members of the public were allowed to access a very restricted and pre-agreed part of the collection under strict supervision.  Because it was largely uncatalogued, described or known there was a terrible fear of what a member of the public might find in the archives. This was understandable since household records such as the cost of banquets were intermingled with personal letters and diaries.  From the public's point of view the archive is that of our Kings and Queens and we would like to access it, but from HRH's view it is her private family archive. Although it is now more acceptable to expose skeletons in the family closet and programs such as "Who do you think you are" promote this, there is probably a reticence from aristocracy and royalty to do this. The Royal Archives is one of the richest, most interesting and significant collections ever created.  It could aptly be described as a pot of gold – an absolute treasure trove. The archivists were aware of this and some of the treasures within it.  The Royal Library at Windsor Castle was in a similar situation and also had extremely restricted access.  Because I have always championed access to archive and library collections I felt very sad whenever I thought of the treasures locked up and hidden (literally) at Windsor Castle.

I was very interested therefore to read about a new development at the Royal Archives timed to coincide with the Diamond Jubilee.  The Queen released this message:

“In this the year of my Diamond Jubilee, I am delighted to be able to present, for the first time, the complete on-line collection of Queen Victoria's journals from the Royal Archives. These diaries cover the period from Queen Victoria's childhood days to her Accession to the Throne, marriage to Prince Albert, and later, her Golden and Diamond Jubilees. Thirteen volumes in Victoria's own hand survive, and the majority of the remaining volumes were transcribed after Queen Victoria's death by her youngest daughter, Princess Beatrice, on her mother's instructions. It seems fitting that the subject of the first major public release of material from the Royal Archives is Queen Victoria, who was the first Monarch to celebrate a Diamond Jubilee. It is hoped that this historic collection will make a valuable addition to the unique material already held by the Bodleian Libraries at Oxford University, and will be used to enhance our knowledge and understanding of the past.”

I was intrigued by this and immediately found the website http://www.queenvictoriasjournals.org/home.do which tells us a lot more about Queen Victoria’s diaries and that this was a project undertaken in conjunction with the Bodleian Library at Oxford and Pro-Quest.  However on looking further it was a bit disappointing since although every page of all the journals has been scanned they have not all been transcribed.  Because they are all handwritten, they won’t be fully text searchable until they are all transcribed, a process which at present is most effectively done by the human hand and eye.  The website doesn’t give any indication of when or how they will be transcribed that I could see, although it says it ‘is in progress’. So far only the first diary has been transcribed, by whom I am not sure. I bet the project is only letting academics do it, who will be paid lots of money and progress very slowly. There is a lot to do: 1832- 1901 since Queen Victoria wrote her diary every day.

If ever I saw a collection that was so well-suited to crowdsourcing for public transcription this is it!  I could guarantee that in a few days or weeks all of Queen Victoria’s diaries would be transcribed by a willing and fascinated public. The handwriting is hard to decipher but with thousands of eyes, and amateur/professional genealogists and historians used to reading old writing, that are highly motivated I am sure it could be achieved. I feel excited just imagining it.  But why stop there?  What about the rest of the collection - the offical royal records and the personal records?  When is that going to come out of hiding? It’s just crying out for public description, tagging and transcribing. If only.  If only.

Extract of Queen Victoria’s diary.

Then thinking I would come back later and have another look I was most disappointed to read that following the example set by the British Library with its UK digitised newspapers the intent is to restrict access to the UK only, and to charge for access from July.  So loyal British subjects living in Commonwealth Countries, and academic researchers – you only have 14 more days to look at this for free, or at all.  Great shame!  But congratulations to whoever it was behind the scenes that convinced HRH to release the diaries from the Royal Archives, and who set up and managed the project with the Bodleian and Pro-Quest.  Bravo!!  Perhaps we just need to beg and grovel for more content and offer our unconditional help to get it for free.

Photo by Rose. June 2012. After participating in Trooping the Colour for the Queens Birthday in Canberra, Irish Guard Cliff Doidge (who plays the clarinet in the Royal Military Band and is on exchange from London to Australia for 4 months) stands beside Lake Burley Griffin with the National Library of Australia behind.



Tuesday, 8 May 2012

Libraries harnessing the cognitive surplus of the nation

It was with great pleasure that I accepted an invitation to lunch at the Parliament of New South Wales last week with Her Excellency Marie Bashir, the Governor of NSW.  The lunch was in memory of Jean Arnot, forward thinking librarian. In her memory each year a female librarian is awarded a prize for the best essay on librarianship. This year I was the JeanArnot Memorial Fellowship prize winner for my essay:  ‘Harnessing the cognitive surplus of the nation: new opportunities for libraries in a time of change’.

The judges said

“your essay was energetic and passionate, and argued cogently for your position, which obviously has significant import for the Library profession”.

The essay was an amalgamation of my ideas, research and practice over the last 4 years into crowdsourcing in libraries.  Although it is aimed at librarians it is equally relevant to archivists. The essay focuses on the idea of cognitive surplus and how and why libraries urgently need to tap into this opportunity. ‘Cognitive surplus’ is a phrase coined by the author and academic Clay Shirky (whose mother is a librarian). It means the free time that people have in which they could be creative or use their brain.  Many people spend their ‘cognitive surplus’ time by watching hours of television, gaming, surfing the internet or reading. However, due to the increased availability of the internet in households, the rise of social media technology, and the desire of people to be creative rather than consumptive, there is now a major change in use of cognitive surplus time. People want to produce and share just as much if not more than consume. Due to new forms of online collaboration and participation, people are seeking out and becoming very productive in online social endeavours. Clay Shirky hypothesizes in his books that there is huge potential for creative human endeavour if the billions of hours that people watch TV are channelled into useful causes instead.

I suggest that libraries can and should harness this cognitive surplus to save themselves. Four powerful examples of libraries harnessing cognitive surplus are:

2008. The National Library of Australia set an international example of how to harness the cognitive surplus of the nation with the Australian Newspapers service. The community is able to improve the computer generated text in digitised historic newspapers by a ‘text correction’ facility, thereby improving the search results in the service. 40,000 people have corrected 52 million lines of text.

2010. The National Library of Finland was the second library to implement community newspaper text correction in their Digitalkoot crowdsourcing project. So far 50,000 people have corrected the text to 99% accuracy.

2011. The New York Public Library released ‘What’s on the menu?’, a crowdsourcing project where the community transcribe text from digitised menus held in the library’s collection. So far 800,000 dishes have been transcribed from 12,000 menus, making them full-text searchable.

2012. The Bodleian Library released the fourth large scale library crowdsourcing project this year. ‘What’s the score?’ is a project where the community can help describe the vast music score collection at Oxford.
If the library profession leverages our expertise with technology and collaboratively harnesses the cognitive surplus of the community we will be able to develop, expand, and open our collections. We will be able to enhance and preserve the social history of the nation while meeting the ever-changing needs of our society. By engaging the community, libraries can develop projects of equal scale, quality and output of commercial endeavours.

The survival of libraries is under threat and I believe that gaining the help of our community with their ideas, knowledge, skills, time and money is the answer. To remain relevant and valued in society libraries must look at their collections and communities in new, imaginative and open ways.  We have the technology to do whatever we want. We must change our culture and thinking to embrace new opportunities such as crowdsourcing on a mass scale.  The value and relevance of libraries is two-fold. It lies in both our collections and in the community that creates, uses, and values these collections. Let us demonstrate this and our place in it. Let us hold onto our original values of open access to all, and do whatever it takes to remain core, valued and relevant in society.

I would encourage you to read the full essay, pass it onto your colleagues, think about this idea deeply and work out how you can harness cognitive surplus to help your profession and organisation in the immediate future.

Photo: Women reach for the skies, big opportunities are out there….
This Andrew Rodgers sculpture was unveiled at Canberra airport on 2 April 2012.  It is the largest bronze figurative sculpture in Australia and is called ‘Perception and Reality 1’.

Monday, 7 May 2012

Church archive starts crowdsourcing: help tag sermon podcasts

There are many church and cathedral archives around the world but a particular one that has just caught my eye and held my interest is All Souls Anglican Church at Langham Place, London. This is because of a crowdsourcing project it has started.  It is setting a fine example for other cathedral and church archives to follow. In a blog post last week the church appealed for Christian volunteers to help make the archive more accessible and used.  The church upholds the principles of information access, strongly believing that resources it generates should be free and open to the community. The church puts its current sermons and talks up on its website as podcasts.  However they have a large back archive of sermons: 3,600 to be precise.  As far as I can see these are all available as podcasts.  To increase their usage and make them more findable they want the community to add subject tags to them. There is a webpage explaining how to do this.  

I followed through to see how simple the process would be.  It is pretty simple and easy to do, but there are a couple of surprising things here.  Firstly it is assumed that only one person needs to allocate tags to a sermon and they will put the ‘right’ tags on. Because of this once someone has ‘grabbed’ a series of sermons to tag no-one else can pick them as far as I could see.  This may have been set up like this because they may have thought that not enough people would sign up to help. However even though the call for help only went out last week, there are very few sermon series left that haven’t been grabbed. I think they have under anticipated the interest and enthusiasm of the crowd here.  Personally I think it may be helpful to encourage more than one person to add tags to the same sermon.  The general premise in crowdsourcing is to use the wisdom of the crowd. This is particularly relevant for tagging.  In order to choose tags the sermon or talk needs to be listened to first. This takes about 30 minutes for each one.

The next interesting thing is that the volunteers can pick 3-4 tags from a very small controlled list and then they have a chance to add one tag of their own choosing that is not on the list.  That one tag will be moderated by the archivist (and presumably added to the list if deemed suitable and often used).  This is the first time that I have seen a combination tagging approach. Again I’m not quite sure about the thinking behind this. I would like to know more. This is a very interesting project to me because firstly it is a small controlled experiment into crowdsourcing where it will be very easy to report back to the community on results, levels of activity and lessons learned.  If successful as I am sure it will be, it could easily be replicated in other church archives, or widened for other item types in the church archive.  It is also a demonstration of how to make audio-visual content better searchable, as well as calling on a specific group of the community – Christians. 

I am really interested to hear more about the results and lessons learned from this small experiment.

Photo: I was lost and parked the car to consult the map when I noticed the car in front of me, it gave me a chuckle…



Sunday, 11 March 2012

Crowdsourcing transcription of handwritten archives


One of the big differences between libraries and archives is that libraries tend to have more of ‘the printed word’ whilst archives have vast amounts of ‘the handwritten record’.  While some libraries are getting up to speed with mass digitisation of books and journals and then being able to offer users full text searchable digitised items, this is still a distant dream for most archives.  Some archives are undertaking mass digitisation, but the second step – making handwritten records full-text searchable is a massive challenge.  The reason for this is in the technology and processing steps.

After scanning a ‘printed word’ page into an image file a piece of software called Optical CharacterRecognition (OCR) converts the image into searchable text.  The OCR works best with clean, clear, black and white typeface such as a word document or a book, not quite so well on old books and journals, and very poorly on old newspapers.  When it comes to converting handwriting it fails miserably.  It just can’t distinguish and convert handwriting to text in the way the human eye can.  Therefore archives can’t easily automate the second part of the digitisation process using OCR software like libraries can for the printed word.

If you at least get some OCR text from print that is readable and therefore searchable you can offer a service to users to full-text search the books or journals such as Google does. If the OCR text is poor there are some things you can do to improve it. You can encourage users of your service to correct the OCR text with a text correction tool so that the searching is improved, such as Trove does with the Australian Newspapers.

Unfortunately the only viable option open to archives to convert digital images into full-text searchable text is to use a manuscript transcription tool, in combination with harnessing the power of a crowd to do the transcription work.  The transcription work for handwritten records is much harder than for example text correcting old newspapers because the handwriting is often difficult to read, old fashioned, barely legible and not necessarily structured in lines or columns. There is often nothing to go on.

I recently stumbled across a blog all about manuscript transcription tools that is written by a software developer Ben W Brumfield in Texas. Ben developed his own software to transcribe his great-great grandmother’s journal. ‘FromthePage’ is now being used by archives because Ben has made it available open source.

A year ago he wrote an in-depthblog post that covered manuscript transcription tools under development, manuscript transcription projects in archives, and made some predications for future directions of manuscript transcription.  I am not going to repeat what he said here, I suggest you read the post in full.  He notes that software development in this area is still fragmented and young with no particular tools taking dominance. Most developed applications are being made available open source. A standout is ‘Scribe’ from the Zooniverse team, currently being used by both the ‘Old Weather’ project to transcribe maritime weather records and by ‘What’s the score’ project to transcribe music scores at the Bodleian Library, Oxford.

Before an archive implements a manuscript tool it needs to find out what it’s users  would most like to be easily full-text searchable from the vast vaults of all the content it has. It is important to find this out, because the crowd will only be motivated and swell in numbers if they really feel what they are doing is very important to a broad group of people and really matters either right now, or in the long-term and is also interesting.  They have to feel this before they will join in.  Once they have joined in there are other motivational tips you can do to keep them going.  Just implementing a manuscript tool is simply not enough.  You need to engage, watch, understand and learn from your crowd, for they hold the passion and power in their hands to make your project successful or not.


Photo by Rose Holley, outside Canberra Bus Station