Showing posts with label social engagement. Show all posts
Showing posts with label social engagement. Show all posts

Saturday, 10 November 2012

National Archives of Australia embraces crowdsourcing and releases ‘The Hive’.


 
The National Archives of Australia (NAA) has made a bold step into the cultural heritage crowdsourcing arena with ‘The Hive’ which was released two weeks ago. The brand makes a clever play on the word ‘Archive’ combined with the idea of a hive of working bees (the public).  The site encourages the public to transcribe archive records.

Early this year when David Fricker became Director General of the NAA he was quick to encourage staff to think innovatively, embrace change, and to harness opportunities such as crowdsourcing to improve access to our collections. He publicly spoke in favour of  crowdsourcing and a changing business model for archives at the International Council of Archives Congress in August:

“Another key development in expanding access is crowdsourcing. As many of us are now seeing, by allowing the public to contribute to the description of archival resources we are enhancing the ability of future generations to discover and learn from our archives. I also think it is a wonderful opportunity for the public to be more engaged with us as archives and to share in the work we do – preserving the memories of our nations. There is still some work to do here, in order to maximise the value of contributions and to maintain the integrity of our archives as authentic and accurate. However, I do not believe these problems are insurmountable, and indeed I believe these systems can to some extent be self-correcting.

This is a type of the co-design, citizen first activity… drawing on the interest and enthusiasm of the community to bring more of our archives into view – discoverable and retrievable…Access will be online and everywhere, improved by rich new data visualisation techniques and expanded descriptive contributions from an engaged citizenry”.

The Hive is the Archives pilot and experimentation into the potential of large scale transcription crowdsourcing to improve access to records.  Staff have looked closely at other crowdsourcing sites on offer and attempted to build on their knowledge and techniques, to provide a site that could be used as a large scale platform for a variety of transcription crowdsourcing projects.

At present the site offers just over 800 lists for the public to transcribe. Some of these are typed and some handwritten.  They are rated in difficulty as easy, medium or hard.  Part of the difficulty with this project is that the public need to have some understanding of how archives receive and describe their records to make sense of what they are being asked to do.  In simple terms archives receive vast amounts of records (referred to as consignments).  Each consignment comes with a list of the items in it.  However because of the large volume of records being received it is usual that only the consignment record is entered into the catalogue e.g. ‘100 boxes of plans and drawings’ from x government agency, rather than all the individual items on the consignment list being described in the catalogue.  The ideal scenario for users of the archives is that every item e.g. plan and drawing is described on the catalogue so that it can be found.  Without this a lot of guess work goes into finding relevant things, or alternatively personal visits are required to view the hard copy consignment lists.
 
The project that the archives is undertaking is to digitise consignment lists and then make them available for transcription by the public. Once transcribed they become searchable and the items within them can be found more easily.  Because so many of the lists are old and handwritten it is virtually impossible to get good OCR on them.  That’s where the public come in who can read them with the human eye. Also the time of the public is needed to speed up the access. Projections on the time it would take archives staff to describe the lists without public help currently stand at 210 years.  It is anticipated that a member of the public could with relative ease describe several hundred items per hour with the Hive tool, which would make a big difference, especially if there was a swarm.

The consignment lists in the pilot are those that have proved most popular with researchers and contain items in the ‘open period’, that is older than 30 years and now open to the public.  The top interest is lists of architectural drawings and historic buildings. This is closely followed by PNG patrol officer records, maritime incidents, personal records from the war office, prisoners of war, meteorology and cyclones, WW1 intelligence, and oil drilling on the Great Barrier Reef.

In the first 2 weeks 300 records have been transcribed of the 800. There is a definite preference for the lists rated hard (handwritten) and ones that involve names.

The site is well presented and gives volunteer transcribers things we know they want such as progress chart, recent activity, points scoring system, rewards, optional login using Open ID e.g. their Google ID, ability to search and choose items, or just take the next one served up, to pick easy or difficult items, to add a marker for where they got to if they are interrupted, and to favourite records.  The only slight drawback is the placing of the transcription window at the bottom of the screen rather than right or left, which often means it is hard to see the transcription window and the content you are transcribing at the same time. Also the OCR text in the transcription window and the cursor is not hooked directly to the text in the image so it is easy to get lost whilst transcribing sometimes.  This is largely because most of the lists are in tables, and the table rows and columns have not been retained in the OCR, so the OCR is somewhat muddled.  Further development of the site will largely depend on feedback given by the public users, and the ability of the archives to keep up a steady supply of new, interesting digitised consignment lists to the Hive.  The Archives is still considering how it may be able to integrate the public content back into its main catalogue RecordSearch, or integrate the Hive into RecordSearch. In the meantime the list content will remain searchable in the Hive.

There is obviously an expectation from the Archives that by making its content more discoverable it will lead to more access requests.  This is why at point of transcription there is a button which enables the user to request a copy of the item.  These requests are being met by digitising the item, and then uploading them into the main catalogue ‘RecordSearch’ with the full item description.

I congratulate the National Archives of Australia Access Team on the development of this exciting new site, which holds so much potential to improve access to records and engage with our citizens in new ways.

The screenshots below show the site in action:



 Easy level transcription- Archived drawings

Medium Level Transcription - ABC Drama Scripts
 


Difficult level transcription - Plans
 


Sunday, 26 August 2012

Digital vandalism or just good fun? The Prado comes to Brisbane


Since I was attending the International Congress of Archives (ICA 2012) in Brisbane last week I had the opportunity to visit the Queensland Art Gallery which is a stone’s throw from the Brisbane Convention Centre.  As an added bonus all conference attendees got a discounted entry to the ‘Portrait of Spain’ Exhibition, which has 100 paintings from the Prado, Madrid on show.

I was surprised to see that visitors were encouraged to get their iphones out at the start of the tour. The primary reason was to take a photo of yourself against a backdrop of a Prado gallery, so that you could pretend to friends you had actually been to the Prado. The ticket collector obliged and took my photo:
 
As I approached the first painting a security guard then warned me that photos were not allowed, with or without flash, and I should put my iphone away.  At the next painting when I commented to another visitor how little description was provided beside paintings another security guard overheard and said ‘oh you can use your iphone’.  Now somewhat perplexed I asked again if I could take a photo.  ‘No, but you can scan the QR codes beside the paintings which tell you more information about them’. [If you don't know what QR codes are and how galleries and museums use them read this short explanation].

I couldn’t be bothered with that.  That is, until I got to a breathtaking painting that had absolutely no explanation about why it took your breathe away.  The painting appeared to be of a man (with hairy forearms, moustache, thick neck) but dressed in formal women’s court clothing with a bust. The description had no mention about this surprising phenomena.  At that point I got my iphone out and checked the QR code.  Disappointingly still no info on the surprise, simply that the artist had ‘a good eye for detail and had painted the hair well. The woman was graceless and had a less than feminine appearance’.  I made a note of the painting to look it up afterwards and see if anyone else had more information on the person in it that they were willing to candidly share. Perhaps a Wikipedia entry? The painting is called Senora de Delicado de Imaz by Vincent Lopez Portana from the Spanish Court of 1833.

On leaving the exhibition I felt in dire need of a cup of tea so headed towards what I thought was the café, only to be blocked by a guard who demanded to see my entry ticket (yes there are a lot of guards).  I queried why I should need to show my entry ticket to partake of tea and was told that this café was themed with Spanish food (?). I was about to turn away when she also mentioned that the photo booths were included in the ticket price.  Being a sucker for a photo I couldn’t resist so passed through the checkpoint having no idea what the photo booth would do. 

Well – talk about surprise!! You are absolutely not allowed to take a photo of a painting in an art gallery.  The answer usually given is because of copyright or mis-use or inappropriate use of images. Clearly the images in the Prado exhibition are out of copyright dating from 1500-1800. Images of them are sold in the gift shop.  However the photo-booth allowed me to stick my own face into a selection of the Prado portraits (in a similar way to sticking your head through a funfair cardboard cut out of Popeye and Olive) and create my own digital image.  After all the high brow gallery poppycock I have heard over the years about galleries digitising images and rules they have made up around this, I was staggered and thrilled to be able to do this fun activity (which I think is probably aimed at children!).  It was the most fun I have had for a while and my first foray into what is commonly called ‘digital vandalism’ or mis-appropriate use of digital artworks.

It made me recall the battle that Wikipedia had with art galleries and use of images over a long duration.  On one particular occasion Liam Wyatt the VP of Wikimedia Australia spoke about how the public were not allowed to take or share photos of artworks for invalid reasons of copyright ownership or inappropriate or commercial use of images, but then the galleries or museums in question would use the same images themselves to make things like ties or mugs for a profit in the gift shop. This was such a case exactly, but more extreme than any I have yet seen. In case you are in doubt – I am endorsing the activity offered in the photo booths at Queensland Art Gallery.  The e-mail I received with my bastardised portrait also had an animated version… where things like my hand and head moved – crikey! The experience led  me to formally write to Queensland Art Gallery (via their online form) and ask them why if I can do this I was not allowed to take a photo without flash of the same painting in the exhibition?  That was 7 days ago and I still have not had a reply….

Incidentally I cannot find a Wikipedia entry for the portrait of Senora de Delicado de Imaz, or anything about her life and circumstance which I am sure is most interesting.  I could also not find a really good digital print online that matched the real life experience of seeing the painting. I did manage however to take a photo of the painting myself via the photo booth.  For some strange reason the photo booth kept offering me this painting as the perfect match to put my face into.

However I preferred to go with a much more regal match.

Me with my head digitally stuck into the portrait of La infanta Isabel Clara Eugenia Magdalena Ruiz by Alonso Sanchez Coello 1588.
The animation had me fiddling with my minature and the monkeys.....

 

Saturday, 25 August 2012

Crowdsourcing and Social Media at US National Archives (NARA). The Citizen Archivist Dashboard


Last week I attended the International Congress of Archives (ICA 2012) which was held in Brisbane. Over 1,000 Archivists from 93 countries attended.

The much anticipated opening keynote on the first day was given by David Ferriero
head of US National Archives.  He is the first librarian to become a National Archivist, previously being in charge of New York Public Library and known for promoting use of social media and relationships with Google and Wikipedia.  His talk was called ‘A world of social media’. I was looking forward to hearing what the US National Archives are doing with social media and crowdsourcing.  People were generally of the opinion that this organisation will/is leading by example in this field.

David Ferriero took to the stage and took us by surprise.  He only used 20 minutes of his 40 minute slot, gave no presentation, instead reading from his notes at breakneck speed and bombarding us with statistics that were largely out of context. At the end he took no questions and dashed off the stage.  He left a surprised and bewildered audience behind.  I for one was immensely disappointed not to see and hear more about some of the exciting US Archives activities. He of course may have had mitigating circumstances that I am totally unaware of.  He did however give small tasters of what his organisation is doing. There was brief mention of large scale crowdsourcing on unspecified projects, a citizen archivists dashboard, and a relationship with Wikipedia which peaked my interest.

So I decided to follow up online and find out for myself what may be happening at NARA. I took me quite some time to search the internet and blogs and get the information I had hoped David would give in his keynote, but it was worth it. Here is what I found:

1. Citizen Archivist Dashboard Webpage http://www.archives.gov/citizen-archivist/

In January 2012 the US National Archives launched the Citizen Archivist Dashboard. This is a great webpage bringing all the online and physical social engagement and crowdsourcing activities together.  It is easy for someone to see what options they may have to help the US National Archives. It is very clearly designed and I like it a lot.

 
2. Transcription Projects

There are two transcription projects going on for handwritten records. Firstly the National Archives Transcription Pilot Project. It appears still to be in ‘pilot’ mode (started in January 2012) since only 300 documents (about 1,000 pages) are available for transcription. They have been very carefully selected from a collection of billions of pages and graded by colour codes according to how difficult the handwriting is to read. This pre-selection must have taken very valuable staff time. You can browse or search by difficulty of transcription, year, and the status of transcription: “Not Yet Started,” “Partially Transcribed,” and “Completed.” You then choose a page to work on and then that page is blocked to other users, so it’s not being edited by multiple users at the same time.  The interface is very simple, much like the Australian Newspapers. In a free text box beside the image you can transcribe what you see. No login is required, though you do have to complete a captcha. 

The missing part is that I can’t see how many people have transcribed what.  It’s not clear if the documents disappear from here when fully transcribed, and how and where they become full text searchable in the collection.  It also seems to be a time consuming process for NARA staff to do the pre-selection and difficulty rating of the documents. This is of course a very small pilot and hopefully lessons will be learnt and the site will be developed further to reach it’s full potential. Also it would be good if more documents became available for transcription. This is one of the easiest handwritten transcription tools I have seen.  I could not find any information about who developed the tool and if it is available open source.

Interestingly David Ferriero says that many US school children are no longer taught cursive handwriting and therefore cannot read handwriting. He says ‘Help us transcribe records and guarantee that school children can make use of our documents’. I’m not quite clear if he thinks this is a potential crowdsourcing exercise for school children to learn handwriting and become better educated, or if adults are supposed to do it so that school children can just read the finished text.

The National Archives have developed a relationship with the Wikipedia Community and currently have a Wikipedian in residence. As part of that program they have shared some primary handwritten national documents into ‘Wikisource’ for transcription via the Wikisource Tool. These documents are mostly at the beginner level in terms of difficulty. I’m not clear if they are the same ones in being used in the Archives own pilot, or different documents. I’m also not clear why they are piloting two different methods for transcription, or what the initial results are compared to each other. Wikisource offers more than transcription however, Wikipedians (if they can get access to original documents or copies) can also scan documents and OCR them.

3. Scanning Projects

  • Scanathons
For reasons I don’t understand the US National Archives has only digitised 750,000 of its 40 million images. This is a very low figure for an organisation like this. They seem to be focusing quite a lot of effort on getting physical volunteers to come in person to the Archives to digitise/scan images for them at ‘Scanathons’. This started in 2011. In January 2012 there was a 4 day Wikipedia ExtravaSCANza. Over the 4 days a group of Wikipedians met in the Still Pictures Research Room and scanned 500 images on desktop scanners. Each day there was a theme: NASA, women’s history, Chile, and battleships.

NARA encourages readers to take their own photos of records in the reading rooms and upload them to a special group in Flickr.  The important thing here is that they should also be described with title, series, and record group if possible so they can be found. So far only 20 people have joined the group and 133 photos have been uploaded (most of these by the same person). I’m not clear how NARA intends to link these digital images back to the item descriptions in their collections but this is a great idea to tackle large scale digitisation of images.

 

The tagging facility, unlike the other pilots seems to me to be unlikely to succeed in its objectives. This is perhaps because of the tight controls that have been placed around it and the isolation of the activity from normal search and browse behaviour. Whilst anyone can easily transcribe a record without needing to login the process for tagging is difficult.

The activity is focused on Tuesdays and themed around a topic.  Records for the topic are pre-selected by the Archives and available in an online group e.g. Elvis, Titanic.  Volunteers must register and follow a set of guidelines; Tags will be reviewed by NARA staff before being accepted and going live on the database. I looked at the topics and it was unclear to me why if the Archives had already identified the items as being about Elvis they couldn’t simply generate an automatic tag for ‘Elvis’. In my opinion tagging is not actually a crowdsourcing activity because individuals are motivated to add tags to help themselves find things, it is a by product of search. Research shows it is rare for users to have concensus on tag terms and use. Crowdsourcing activities achieve a big clear goal that could not be achieved by individuals alone, and everyone in the crowd should be aware of how they are helping the ultimate goal.  

5. Indexing the 1940 Census

On April 2, 2012, NARA released the digital images of the 1940 United States Federal Census after a 72 year embargo. The census images will be uploaded and made available on Archives.com, FindMyPast.com, National Archives, ProQuest, and FamilySearch.org. The entire 1940 census data will be indexed by a community of volunteers and made available for free. The free index of the census records and corresponding images will be available to the public for perpetuity.

6. Useful Links

I found a recent presentation given this year by Pamela Wright – Chief Digital Access Strategist at NARA which gives screenshots of what I have talked about above. ‘From access to engagement’

7. Social Media

NARA are active users of social media channels and they have started to monitor their activity. The Social media statistics from NARA May 2012 may be interesting reading for some.

I would be interested in reading more presentations or articles about the citizen archivist pilot projects from NARA and finding out what they have achieved and learnt so far. I hope this information is made available to the archives and library community soon.  Please reply in comment if you have any more information on the pilot activities.

Tuesday, 8 May 2012

Libraries harnessing the cognitive surplus of the nation

It was with great pleasure that I accepted an invitation to lunch at the Parliament of New South Wales last week with Her Excellency Marie Bashir, the Governor of NSW.  The lunch was in memory of Jean Arnot, forward thinking librarian. In her memory each year a female librarian is awarded a prize for the best essay on librarianship. This year I was the JeanArnot Memorial Fellowship prize winner for my essay:  ‘Harnessing the cognitive surplus of the nation: new opportunities for libraries in a time of change’.

The judges said

“your essay was energetic and passionate, and argued cogently for your position, which obviously has significant import for the Library profession”.

The essay was an amalgamation of my ideas, research and practice over the last 4 years into crowdsourcing in libraries.  Although it is aimed at librarians it is equally relevant to archivists. The essay focuses on the idea of cognitive surplus and how and why libraries urgently need to tap into this opportunity. ‘Cognitive surplus’ is a phrase coined by the author and academic Clay Shirky (whose mother is a librarian). It means the free time that people have in which they could be creative or use their brain.  Many people spend their ‘cognitive surplus’ time by watching hours of television, gaming, surfing the internet or reading. However, due to the increased availability of the internet in households, the rise of social media technology, and the desire of people to be creative rather than consumptive, there is now a major change in use of cognitive surplus time. People want to produce and share just as much if not more than consume. Due to new forms of online collaboration and participation, people are seeking out and becoming very productive in online social endeavours. Clay Shirky hypothesizes in his books that there is huge potential for creative human endeavour if the billions of hours that people watch TV are channelled into useful causes instead.

I suggest that libraries can and should harness this cognitive surplus to save themselves. Four powerful examples of libraries harnessing cognitive surplus are:

2008. The National Library of Australia set an international example of how to harness the cognitive surplus of the nation with the Australian Newspapers service. The community is able to improve the computer generated text in digitised historic newspapers by a ‘text correction’ facility, thereby improving the search results in the service. 40,000 people have corrected 52 million lines of text.

2010. The National Library of Finland was the second library to implement community newspaper text correction in their Digitalkoot crowdsourcing project. So far 50,000 people have corrected the text to 99% accuracy.

2011. The New York Public Library released ‘What’s on the menu?’, a crowdsourcing project where the community transcribe text from digitised menus held in the library’s collection. So far 800,000 dishes have been transcribed from 12,000 menus, making them full-text searchable.

2012. The Bodleian Library released the fourth large scale library crowdsourcing project this year. ‘What’s the score?’ is a project where the community can help describe the vast music score collection at Oxford.
If the library profession leverages our expertise with technology and collaboratively harnesses the cognitive surplus of the community we will be able to develop, expand, and open our collections. We will be able to enhance and preserve the social history of the nation while meeting the ever-changing needs of our society. By engaging the community, libraries can develop projects of equal scale, quality and output of commercial endeavours.

The survival of libraries is under threat and I believe that gaining the help of our community with their ideas, knowledge, skills, time and money is the answer. To remain relevant and valued in society libraries must look at their collections and communities in new, imaginative and open ways.  We have the technology to do whatever we want. We must change our culture and thinking to embrace new opportunities such as crowdsourcing on a mass scale.  The value and relevance of libraries is two-fold. It lies in both our collections and in the community that creates, uses, and values these collections. Let us demonstrate this and our place in it. Let us hold onto our original values of open access to all, and do whatever it takes to remain core, valued and relevant in society.

I would encourage you to read the full essay, pass it onto your colleagues, think about this idea deeply and work out how you can harness cognitive surplus to help your profession and organisation in the immediate future.

Photo: Women reach for the skies, big opportunities are out there….
This Andrew Rodgers sculpture was unveiled at Canberra airport on 2 April 2012.  It is the largest bronze figurative sculpture in Australia and is called ‘Perception and Reality 1’.

Sunday, 26 February 2012

Crowdsourcing Australian Climate Change


In my last blog post I described how knitters and yarn enthusiasts were crowdsourcing knitting patterns from digitised Australian newspapers in Trove for use in a crowdsourcing site called Ravelry http://www.ravelry.com.  This week I wanted to give another example of the Australian Newspapers in Trove giving leverage to yet another crowdsourcing site. This time it’s for research into climate change and the site is called OzDocs. http://ozdocs.climatehistory.com.au/

Australian newspapers hold unique content, for example convict records and climate records. In Australia official weather records only began in 1908 when the Bureau of Meteorology was established. However there are weather tables and forecasts appearing in Australian newspapers from 1803 onwards.  The newspapers therefore provide 200 years of weather records.  Newspapers not only give tables with statistics of temperature, rainfall, winds etc, but also eye witness accounts of weather conditions such as floods, droughts and fire.

The citizens and politicians of Australia have a high and ongoing level of interest in climate change and how it is affecting our nation. A project investigating climate change is SEARCH: SouthEastern Australian Recent Climate History. It spans the sciences and the humanities, drawing together a team of leading climate scientists, water managers and historians in Australia to better understand south-eastern Australian climate history over the past 200–500 years. The digital newspapers in Trove are a fundamental part of SEARCH’s research process. However even though the newspapers are full-text searchable it is still a challenge to find and bring together in context the eye-witness accounts and the weather tables so that the temperatures, rainfall and other statistics can be transcribed into a research database. This is why the SEARCH project has established this month the OzDocs citizen science project, with a $10,000 grant from the University of Melbourne so that the public can help them. Basically the public are asked to find and tag historic newspaper articles on weather conditions, and transcribe useful weather statistics from historic newspapers into a database.

In 2010 Joelle Gergis, the lead SEARCH investigator spoke to me and said:

“Having all this information online and being able to quickly access it has been amazing. Being able to find weather tables from 1803 onwards in the Sydney Gazette is crucial to our research. Official records from the Bureau of Meteorology only began 100 years ago so being able to access the newspaper records which are earlier than this is really useful. The sources in Trove also show how weather events have affected society, with eye witness accounts of floods and bushfires. For example we have been researching the 1851 Black Thursday bushfires in Victoria.”

In 2010 as a locust plague swept across the south-eastern side of Australia the pilot volunteers for the project working at the State Library of New South Wales noted that weather conditions in 1825 were very similar (heavy rainfall followed by nice weather then a terrible locust plague) and found eye-witness accounts in The Sydney Gazette and New South Wales Advertiser, 24 March 1825 of a similar locust plague.

“Prior to the late rains the caterpillar, that old enemy to the agriculturist interest of the Colony, made its appearance ; but, upon the visitation of the heavy showers, their ranks were consider ably thinned. However, since the present enchanting fine weather has again set in, the number of these destructive insects has increased to an unparalleled extent, covering whole fields in their course, which in some spots seemed to be towards the South, in a line from East to West. Wherever they make their appearance, the most complete destruction immediately follows. Upon Captain Campbell’s estate, in the district of Cooke, they were supposed to be at least two inches in height.”

This month I caught up with Joelle again to see how the newly released crowdsourcing part of the project is going. She said:

“We now have over 100 volunteers who have contributed over 4000 articles. The database will be searchable in our next release. It will be the country’s first publicly searchable database of climate information using a diverse collection of pre-20th century historical records. The database will give easy access to information for researchers, organisations, government departments and the public.”

The scope of OzDocs work has now been expanded to include not only the digitised newspapers but other resources that are held in the State Library of Victoria, State Library of New South Wales and the National Library of Australia.  These pre-date the newspapers by another 100 years and go back to the 1700’s. Joelle said:

“Our OzDocs volunteers will be working their way through logbooks of the first European explorers, governors’ correspondence, early settlers’ diaries, newspapers and the works of 18th and 19th century scholars.”

On of the questions the SEARCH project hopes to answer is what the South East region of Australia’s ‘natural’ climate has been like since 1788. This may ultimately help to refine current climate models, allowing more accurate climate change estimates to be developed for the future.  The lack of records in a consistent accessible format before 1900 is currently making this difficult.

For the project to be successful the volunteer numbers really need to increase a lot.  At the moment this is still quite a small scale effort compared to the knitting project Ravelry that I reported on in my last blog.  Volunteers can join by accessing the OzDocs site. http://ozdocs.climatehistory.com.au/



Tuesday, 17 January 2012

Social metadata and sharing stuff

Since 2009 I have been undertaking research for OCLC Research (formally known as the Research Library Group RLG). Library and archive professionals from partner institutions around the world contribute to research groups that focus on topics of interest to the information community. There are currently around 50 research activities in progress. 

I am a member of the group called ‘sharing and aggregating social metadata’. This group is quite large and there are 21 of us  from different institutions in 5 countries.  It has been great to work with other professionals from institutions such as Yale, Stanford, Berkeley, and Getty. In the normal course of my day job I have a high level of contact with other national libraries, and institutions at the cutting edge of digital technology, but to have this opportunity to research a specific topic in detail and have ongoing discussions with the group over 2 years was really rewarding.

The group started by setting itself questions to answer for example:
  • What are the objectives for social metadata and how do we measure success?
  • What user contributions would most enrich existing metadata created by libraries, archives, and museums?
  • What are examples of successful social media sites and what factors contribute to their success?
  • What best practices currently exist, or need to be developed, that can guide institutions in managing user contributions and various related issues?
  • To what extent is moderation necessary or desirable?
  • How are cultural institutions integrating social metadata into formal taxonomies?
The research was divided up into chunks and mini groups formed.  We quickly realised we had to establish agreed terminology and definitions of what we were researching. These were:

Social media/networking: Ways for people to communicate online with each other e.g. Twitter, Facebook, Blogs.
User Generated Content (UGC): Things produced by users rather than owners of the site e.g. image, video, text AND metadata – tags, comments, notes.
Social Metadata: Additional information about a resource given by online users e.g. tags, comments.
Social Media Features: Interactive features added to a site that enable virtual groups to build and communicate with each other and social metadata to be added.
Social Engagement:   User interaction online e.g. communication between users, from users to site owners, from users with objects/resources.
Web 2.0: Online applications that facilitate interactive rather than passive experiences.

We used Basecamp project management software to work together. Most of us never met other members of the group face-to-face, just online or by telephone. I would certainly enjoy meeting the whole group face to face sometime in the future.

The research activity of the group and the volume of output was much larger than expected, so rather than ending up with a single report we have written three. I am writing this blog post now because the first two reports have recently been published and the third is expected to be released next month (February 2012).

Our first report, Social Metadata for Libraries, Archives, and Museums, Part 1: Site Reviews, provides an environmental scan of sites and third-party hosted social media sites relevant to libraries, archives, and museums. We provide a brief overview of each site and why it was of interest to us. We noted which social media features each site supported, such as tagging, comments, reviews, images, videos, ratings, recommendations, lists, links to related articles, etc. The report also contains a very useful and interesting section written by Cyndi Shein on use of third-party sites and blogs by libraries, archives and museums. The third party sites include LibraryThing, LibraryThing for Libraries, Flickr, Flickr Commons, YouTube, Facebook, Twitter and Wikipedia. We particularly focused on institutions that were doing cool, groovy or unusual things.  The good thing about use of third party sites is that the cost is minimal or nothing, so if you have plenty of ideas and a bit of time but very little budget you can still do some really interesting things by tapping into some of their better features. I strongly recommend a read of this part (pages 37 to 67). ‘Regardless of the challenges in using third party sites to host content and relate to users, most LAMs believe their efforts are well spent’.

Our second report Social Metadata for Libraries, Archives, and Museums, Part 2: Survey Analysis  is our analysis of the results from a social metadata survey of site managers conducted from October to November 2009. In here we find that engaging new or existing audiences is used as a success criteria more frequently than any other criteria; only a small minority of survey respondents are concerned about the way the site’s content is used or repurposed outside the site; spam and abusive user behavior are sporadic and easily managed; engagement is best measured by quality, not quantity.

The upcoming third report Social Metadata for Libraries, Archives and Museums, Part 3: Recommendations and Readings and the Executive Summary for which I gave my final edits last week provides recommendations on application of social metadata features for libraries, archives, and museums and factors which contribute to success. It also contains an annotated bibliography. I will blog more on the recommendations once it is published but in the meantime I have two things to share:

1.      The social metadata research group believes it is riskier to do nothing and become irrelevant to your user communities than to start using social media features. A major question to consider before you start is ‘What are your objectives for using social media?’ 

2.      Whilst sharing our experiences within the group and also analysing the results of the survey we realised there is often a tension between the organizational desire to have “one voice” in the media, with social media as an important marketing tool, and the information specialists drive to communicate - in both directions with multiple voices - in various channels. We thought that distinguishing between using social media to create community around your organization (the province of public relations offices) and using social media to create community around collections was important. Publicity and participation are at different ends of the spectrum. Although it is important to develop the patron base for the institution through good use of social media publicity tools, it is equally important to give those patrons a voice - and therefore a sense of ownership - in the materials and content curated by the institution.

Now that our research is over there is an empty space for me.  I miss the share of information amongst the group and particularly the emails titled ‘you must read or watch this!’ Having being the person responsible for compiling the bibliography in the third report I know that we all read or watched over 200 items of interest in a 12 month period.  Many of these were blog posts. My favourite YouTube videos that were shared in the group happen to be both the first and last items we sent: ‘you must watch this!’


The Machine is Us/ing us (1.4 million views)  




Gotta Share: The Musical (1.5 million views)