Showing posts with label user generated content. Show all posts
Showing posts with label user generated content. Show all posts

Saturday, 10 November 2012

National Archives of Australia embraces crowdsourcing and releases ‘The Hive’.


 
The National Archives of Australia (NAA) has made a bold step into the cultural heritage crowdsourcing arena with ‘The Hive’ which was released two weeks ago. The brand makes a clever play on the word ‘Archive’ combined with the idea of a hive of working bees (the public).  The site encourages the public to transcribe archive records.

Early this year when David Fricker became Director General of the NAA he was quick to encourage staff to think innovatively, embrace change, and to harness opportunities such as crowdsourcing to improve access to our collections. He publicly spoke in favour of  crowdsourcing and a changing business model for archives at the International Council of Archives Congress in August:

“Another key development in expanding access is crowdsourcing. As many of us are now seeing, by allowing the public to contribute to the description of archival resources we are enhancing the ability of future generations to discover and learn from our archives. I also think it is a wonderful opportunity for the public to be more engaged with us as archives and to share in the work we do – preserving the memories of our nations. There is still some work to do here, in order to maximise the value of contributions and to maintain the integrity of our archives as authentic and accurate. However, I do not believe these problems are insurmountable, and indeed I believe these systems can to some extent be self-correcting.

This is a type of the co-design, citizen first activity… drawing on the interest and enthusiasm of the community to bring more of our archives into view – discoverable and retrievable…Access will be online and everywhere, improved by rich new data visualisation techniques and expanded descriptive contributions from an engaged citizenry”.

The Hive is the Archives pilot and experimentation into the potential of large scale transcription crowdsourcing to improve access to records.  Staff have looked closely at other crowdsourcing sites on offer and attempted to build on their knowledge and techniques, to provide a site that could be used as a large scale platform for a variety of transcription crowdsourcing projects.

At present the site offers just over 800 lists for the public to transcribe. Some of these are typed and some handwritten.  They are rated in difficulty as easy, medium or hard.  Part of the difficulty with this project is that the public need to have some understanding of how archives receive and describe their records to make sense of what they are being asked to do.  In simple terms archives receive vast amounts of records (referred to as consignments).  Each consignment comes with a list of the items in it.  However because of the large volume of records being received it is usual that only the consignment record is entered into the catalogue e.g. ‘100 boxes of plans and drawings’ from x government agency, rather than all the individual items on the consignment list being described in the catalogue.  The ideal scenario for users of the archives is that every item e.g. plan and drawing is described on the catalogue so that it can be found.  Without this a lot of guess work goes into finding relevant things, or alternatively personal visits are required to view the hard copy consignment lists.
 
The project that the archives is undertaking is to digitise consignment lists and then make them available for transcription by the public. Once transcribed they become searchable and the items within them can be found more easily.  Because so many of the lists are old and handwritten it is virtually impossible to get good OCR on them.  That’s where the public come in who can read them with the human eye. Also the time of the public is needed to speed up the access. Projections on the time it would take archives staff to describe the lists without public help currently stand at 210 years.  It is anticipated that a member of the public could with relative ease describe several hundred items per hour with the Hive tool, which would make a big difference, especially if there was a swarm.

The consignment lists in the pilot are those that have proved most popular with researchers and contain items in the ‘open period’, that is older than 30 years and now open to the public.  The top interest is lists of architectural drawings and historic buildings. This is closely followed by PNG patrol officer records, maritime incidents, personal records from the war office, prisoners of war, meteorology and cyclones, WW1 intelligence, and oil drilling on the Great Barrier Reef.

In the first 2 weeks 300 records have been transcribed of the 800. There is a definite preference for the lists rated hard (handwritten) and ones that involve names.

The site is well presented and gives volunteer transcribers things we know they want such as progress chart, recent activity, points scoring system, rewards, optional login using Open ID e.g. their Google ID, ability to search and choose items, or just take the next one served up, to pick easy or difficult items, to add a marker for where they got to if they are interrupted, and to favourite records.  The only slight drawback is the placing of the transcription window at the bottom of the screen rather than right or left, which often means it is hard to see the transcription window and the content you are transcribing at the same time. Also the OCR text in the transcription window and the cursor is not hooked directly to the text in the image so it is easy to get lost whilst transcribing sometimes.  This is largely because most of the lists are in tables, and the table rows and columns have not been retained in the OCR, so the OCR is somewhat muddled.  Further development of the site will largely depend on feedback given by the public users, and the ability of the archives to keep up a steady supply of new, interesting digitised consignment lists to the Hive.  The Archives is still considering how it may be able to integrate the public content back into its main catalogue RecordSearch, or integrate the Hive into RecordSearch. In the meantime the list content will remain searchable in the Hive.

There is obviously an expectation from the Archives that by making its content more discoverable it will lead to more access requests.  This is why at point of transcription there is a button which enables the user to request a copy of the item.  These requests are being met by digitising the item, and then uploading them into the main catalogue ‘RecordSearch’ with the full item description.

I congratulate the National Archives of Australia Access Team on the development of this exciting new site, which holds so much potential to improve access to records and engage with our citizens in new ways.

The screenshots below show the site in action:



 Easy level transcription- Archived drawings

Medium Level Transcription - ABC Drama Scripts
 


Difficult level transcription - Plans
 


Sunday, 29 April 2012

Mobilising and archiving social metadata (user generated content).


It is fantastic to see members of our library communities adding their own knowledge and opinions to our content through use of features such as tags and comments, and social media tools such as Twitter and Facebook.  More libraries are opening their content and sites to their communities through these tools and features than ever have before. We call this content user generated content (UGC) or social metadata.

But if we think about it for too long it gives us a big headache. Being of the ‘collecting’ mind we really want to care for and keep the UGC in the same way we care for our collection content.  Caring for it means:

  • knowing how much has been added and keeping meaningful statistics.
  • keeping the UGC in context with the data the users meant it to be related to.
  • archiving it for the long-term.
  • being able to migrate it along with our own content as our services and interfaces change in the future.
  • being able to mobilise it to share with other services.
  • being able to easily supply it back to the original creators if they want it.
Doing any one of these things is currently difficult, let alone all of them together. We really haven’t got our act together yet for managing UGC content and social metadata, only enabling the facility for the community to add it.

Firstly let’s take a simple concept. A member of the community is actively engaged with your site.  They are contributing a lot of data to it in the form of comments and descriptions.  After a while they want to get all of ‘their’ data out so they can use it for something else they are working on. Let’s call this ‘user takeout’. Seems reasonable, seems simple, but I don’t know of any library site that does this.  For example a ‘user takeout’ option in Trove newspapers would let a contributor get a copy of all of the comments and tags they have added to historic newspaper articles. You may ask “Do people want to do this?”  Contributors seem to accept that content they add to sites will be locked to that site.  I’m not sure they even think about it very much when they start to add stuff, or check the user licence for the terms. Many don’t intend to add the volume of stuff that they do.  But suddenly they think about it when either a better site comes along that they would like to transfer or copy their content to, or the site they are adding to is unexpectedly taken down or frozen.  Recent examples in the news are Facebook users wanting to be able to transfer or ‘user takeout’ their photographs from the site.  Although of course it is easily technically possible to implement this social media sites such as Facebook are reluctant to let users do this, for fear they will take their content and move to competitors sites. However in the library world it is reasonable that users may want to share their value added data around multiple library sites, and yet we still don’t enable it.  Another item in the news was the suddenclosure of poetry.com. Over 7 million users were given 15 days notice that the 14 million poems they had added would be taken down when the site was sold.  They were not given an easy option to ‘takeout’ their poems, but instead it was suggested that they could copy and paste their poems if they had time. This infuriated many users who read the message, and many others who didn’t read the message in time.  It’s worth pointing out here that a lot of sites people use frequently and think are for the common good are actually commercial sites that can do exactly what they like, and do not ever promise to keep, manage or archive content in the same way libraries do. Although the new owners restored the poetry.com site, it appears that the 14 million poems added prior to 2012 are still not restored hence the large pink box at the top ‘Where’s my poem?’

If we think about measuring our user activity and data through all channels i.e. our own site as well as Twitter, Facebook etc we hit a brick wall.  Providing useful statistics on both volume and value of data social metadata is difficult.  For social media sites such as Twitter and Facebook your options are to either buy costly software and do it yourself, or employ a company (many of which are springing up) to do it for you. These companies however would have great difficulty integrating measurements on the value and content from social media sources with those that go directly to your site i.e. your own comments, tags, blogs.  Doing measurements separately is difficult, but combining them even more so.

Many libraries are part of central or local government so have requirements to archive records and content they create, which should also include social metadata and media.  But does anyone know the best way to do this and are our archives agencies telling us how to do it?  The simple answer is no. The National Archives in USA (NARA) say they are working on it as a matter of urgency. They are due to explain how it should be done by this July. The National Archives of Australia website states that “The Archives Act 1983 does not define a record by its format. Generally, records created as a result of using social media are subject to the same business and legislative requirements as records created by other means.” But the guidelines on the NAA website as to how this should be done simply say “Methods of capturing social media content as a record may vary according to the tools being used”.  This month the Public Record Office of Victoria released an issues paper for comment: ‘Recordkeeping implications ofsocial media’.

An extract of the PROV proposed guidelines for archiving social metadata follows:

How should the record be captured?

Currently printing screenshots to .pdf and registering the resulting document in an Electronic Document and Record Management System (EDRMS) to record the necessary metadata is the  most accessible and expedient method of creating social media records. Necessary metadata includes who sent it (username and real name), date and time of sending, context and purpose of content, name of tool used to create it.

My first reaction on reading this was ‘this is mad!’ Perhaps the archives are under-estimating the amount of social metadata and media activity that is going on.  Taking a screenshot of every tweet for example would assume that you are not going to get thousands, whereas successful sites and topics such as Trove do get thousands and millions of interactions, which makes this unworkable from a staff resourcing point of view. Twitter is notorious for ‘disappearing tweets’ after a very short amount of time – sometimes less than a week because of the volume of activity that takes place.  This also puts pressure on to archive tweets at the time of creation.  You don’t have the luxury to go back and archive later. This suggested form of archiving only gives a screen-based image, which is not in context, not searchable, has no metadata, no timestamp, and is not authenticable. It seems there is money to be made if someone develops a simple software system to mechanically capture the tweet, its response and its components and safely and uniformly archives/indexes them along with descriptive metadata. The tool could also render the page "as it appears" and save it as a PDF if that is required.  

In April 2010 The Library of Congress rather bravelyannounced that it intended to archive all tweets since they began in 2006 to record the social fabric of the world and signed an agreement with Twitter and Google. In 2010 the Twitter archive was growing rapidly with users sending 50 million tweets a day. A year and a half later several news agencies tried to get a progressreport from LC without much success.  Other than trying to transfer the data from Twitter servers to LC servers the LC weren’t giving any detail on what technological developments they were creating to do the mammoth task. The task seemed to be growing bigger by the day with usage of Twitter increasing.  Currently 140 million tweets are sent every day.

A core element of the archive process should be that the data is kept in context with that it was referring to, and other elements surrounding it.  Most libraries that are keeping UGC and social metadata are keeping it in a separate layer to their own content in the database to protect the provenance, but may integrate it for public display.  If it is kept separate it can easily be stored, managed, and moved, but is at risk of becoming separated from the context it is related to. This is something libraries need to work out.  This will become more pressing in a few years time when existing services are migrated as part of their maintenance.  The UGC needs to be migrated in context with them.

On this topic I have more questions than answers. I think libraries and archives need to work together to take an active role in firstly encouraging mobilisation of social metadata -‘user takeout’, and secondly demonstrating how social metadata and social media activity can be archived. I see massive opportunities for start-ups to create archiving tools to bolt onto Facebook, Twitter, Youtube and Blogging software to meet the requirements of government archiving.
 
Photo: Prime Ministers Chiefly and Curtin chat on the way to work 1945. Bronze sculpture by Peter Corlett outside the National Archives of Australia, Canberra. Rose Holley

Tuesday, 17 January 2012

Social metadata and sharing stuff

Since 2009 I have been undertaking research for OCLC Research (formally known as the Research Library Group RLG). Library and archive professionals from partner institutions around the world contribute to research groups that focus on topics of interest to the information community. There are currently around 50 research activities in progress. 

I am a member of the group called ‘sharing and aggregating social metadata’. This group is quite large and there are 21 of us  from different institutions in 5 countries.  It has been great to work with other professionals from institutions such as Yale, Stanford, Berkeley, and Getty. In the normal course of my day job I have a high level of contact with other national libraries, and institutions at the cutting edge of digital technology, but to have this opportunity to research a specific topic in detail and have ongoing discussions with the group over 2 years was really rewarding.

The group started by setting itself questions to answer for example:
  • What are the objectives for social metadata and how do we measure success?
  • What user contributions would most enrich existing metadata created by libraries, archives, and museums?
  • What are examples of successful social media sites and what factors contribute to their success?
  • What best practices currently exist, or need to be developed, that can guide institutions in managing user contributions and various related issues?
  • To what extent is moderation necessary or desirable?
  • How are cultural institutions integrating social metadata into formal taxonomies?
The research was divided up into chunks and mini groups formed.  We quickly realised we had to establish agreed terminology and definitions of what we were researching. These were:

Social media/networking: Ways for people to communicate online with each other e.g. Twitter, Facebook, Blogs.
User Generated Content (UGC): Things produced by users rather than owners of the site e.g. image, video, text AND metadata – tags, comments, notes.
Social Metadata: Additional information about a resource given by online users e.g. tags, comments.
Social Media Features: Interactive features added to a site that enable virtual groups to build and communicate with each other and social metadata to be added.
Social Engagement:   User interaction online e.g. communication between users, from users to site owners, from users with objects/resources.
Web 2.0: Online applications that facilitate interactive rather than passive experiences.

We used Basecamp project management software to work together. Most of us never met other members of the group face-to-face, just online or by telephone. I would certainly enjoy meeting the whole group face to face sometime in the future.

The research activity of the group and the volume of output was much larger than expected, so rather than ending up with a single report we have written three. I am writing this blog post now because the first two reports have recently been published and the third is expected to be released next month (February 2012).

Our first report, Social Metadata for Libraries, Archives, and Museums, Part 1: Site Reviews, provides an environmental scan of sites and third-party hosted social media sites relevant to libraries, archives, and museums. We provide a brief overview of each site and why it was of interest to us. We noted which social media features each site supported, such as tagging, comments, reviews, images, videos, ratings, recommendations, lists, links to related articles, etc. The report also contains a very useful and interesting section written by Cyndi Shein on use of third-party sites and blogs by libraries, archives and museums. The third party sites include LibraryThing, LibraryThing for Libraries, Flickr, Flickr Commons, YouTube, Facebook, Twitter and Wikipedia. We particularly focused on institutions that were doing cool, groovy or unusual things.  The good thing about use of third party sites is that the cost is minimal or nothing, so if you have plenty of ideas and a bit of time but very little budget you can still do some really interesting things by tapping into some of their better features. I strongly recommend a read of this part (pages 37 to 67). ‘Regardless of the challenges in using third party sites to host content and relate to users, most LAMs believe their efforts are well spent’.

Our second report Social Metadata for Libraries, Archives, and Museums, Part 2: Survey Analysis  is our analysis of the results from a social metadata survey of site managers conducted from October to November 2009. In here we find that engaging new or existing audiences is used as a success criteria more frequently than any other criteria; only a small minority of survey respondents are concerned about the way the site’s content is used or repurposed outside the site; spam and abusive user behavior are sporadic and easily managed; engagement is best measured by quality, not quantity.

The upcoming third report Social Metadata for Libraries, Archives and Museums, Part 3: Recommendations and Readings and the Executive Summary for which I gave my final edits last week provides recommendations on application of social metadata features for libraries, archives, and museums and factors which contribute to success. It also contains an annotated bibliography. I will blog more on the recommendations once it is published but in the meantime I have two things to share:

1.      The social metadata research group believes it is riskier to do nothing and become irrelevant to your user communities than to start using social media features. A major question to consider before you start is ‘What are your objectives for using social media?’ 

2.      Whilst sharing our experiences within the group and also analysing the results of the survey we realised there is often a tension between the organizational desire to have “one voice” in the media, with social media as an important marketing tool, and the information specialists drive to communicate - in both directions with multiple voices - in various channels. We thought that distinguishing between using social media to create community around your organization (the province of public relations offices) and using social media to create community around collections was important. Publicity and participation are at different ends of the spectrum. Although it is important to develop the patron base for the institution through good use of social media publicity tools, it is equally important to give those patrons a voice - and therefore a sense of ownership - in the materials and content curated by the institution.

Now that our research is over there is an empty space for me.  I miss the share of information amongst the group and particularly the emails titled ‘you must read or watch this!’ Having being the person responsible for compiling the bibliography in the third report I know that we all read or watched over 200 items of interest in a 12 month period.  Many of these were blog posts. My favourite YouTube videos that were shared in the group happen to be both the first and last items we sent: ‘you must watch this!’


The Machine is Us/ing us (1.4 million views)  




Gotta Share: The Musical (1.5 million views)