Sunday, 29 April 2012

Mobilising and archiving social metadata (user generated content).


It is fantastic to see members of our library communities adding their own knowledge and opinions to our content through use of features such as tags and comments, and social media tools such as Twitter and Facebook.  More libraries are opening their content and sites to their communities through these tools and features than ever have before. We call this content user generated content (UGC) or social metadata.

But if we think about it for too long it gives us a big headache. Being of the ‘collecting’ mind we really want to care for and keep the UGC in the same way we care for our collection content.  Caring for it means:

  • knowing how much has been added and keeping meaningful statistics.
  • keeping the UGC in context with the data the users meant it to be related to.
  • archiving it for the long-term.
  • being able to migrate it along with our own content as our services and interfaces change in the future.
  • being able to mobilise it to share with other services.
  • being able to easily supply it back to the original creators if they want it.
Doing any one of these things is currently difficult, let alone all of them together. We really haven’t got our act together yet for managing UGC content and social metadata, only enabling the facility for the community to add it.

Firstly let’s take a simple concept. A member of the community is actively engaged with your site.  They are contributing a lot of data to it in the form of comments and descriptions.  After a while they want to get all of ‘their’ data out so they can use it for something else they are working on. Let’s call this ‘user takeout’. Seems reasonable, seems simple, but I don’t know of any library site that does this.  For example a ‘user takeout’ option in Trove newspapers would let a contributor get a copy of all of the comments and tags they have added to historic newspaper articles. You may ask “Do people want to do this?”  Contributors seem to accept that content they add to sites will be locked to that site.  I’m not sure they even think about it very much when they start to add stuff, or check the user licence for the terms. Many don’t intend to add the volume of stuff that they do.  But suddenly they think about it when either a better site comes along that they would like to transfer or copy their content to, or the site they are adding to is unexpectedly taken down or frozen.  Recent examples in the news are Facebook users wanting to be able to transfer or ‘user takeout’ their photographs from the site.  Although of course it is easily technically possible to implement this social media sites such as Facebook are reluctant to let users do this, for fear they will take their content and move to competitors sites. However in the library world it is reasonable that users may want to share their value added data around multiple library sites, and yet we still don’t enable it.  Another item in the news was the suddenclosure of poetry.com. Over 7 million users were given 15 days notice that the 14 million poems they had added would be taken down when the site was sold.  They were not given an easy option to ‘takeout’ their poems, but instead it was suggested that they could copy and paste their poems if they had time. This infuriated many users who read the message, and many others who didn’t read the message in time.  It’s worth pointing out here that a lot of sites people use frequently and think are for the common good are actually commercial sites that can do exactly what they like, and do not ever promise to keep, manage or archive content in the same way libraries do. Although the new owners restored the poetry.com site, it appears that the 14 million poems added prior to 2012 are still not restored hence the large pink box at the top ‘Where’s my poem?’

If we think about measuring our user activity and data through all channels i.e. our own site as well as Twitter, Facebook etc we hit a brick wall.  Providing useful statistics on both volume and value of data social metadata is difficult.  For social media sites such as Twitter and Facebook your options are to either buy costly software and do it yourself, or employ a company (many of which are springing up) to do it for you. These companies however would have great difficulty integrating measurements on the value and content from social media sources with those that go directly to your site i.e. your own comments, tags, blogs.  Doing measurements separately is difficult, but combining them even more so.

Many libraries are part of central or local government so have requirements to archive records and content they create, which should also include social metadata and media.  But does anyone know the best way to do this and are our archives agencies telling us how to do it?  The simple answer is no. The National Archives in USA (NARA) say they are working on it as a matter of urgency. They are due to explain how it should be done by this July. The National Archives of Australia website states that “The Archives Act 1983 does not define a record by its format. Generally, records created as a result of using social media are subject to the same business and legislative requirements as records created by other means.” But the guidelines on the NAA website as to how this should be done simply say “Methods of capturing social media content as a record may vary according to the tools being used”.  This month the Public Record Office of Victoria released an issues paper for comment: ‘Recordkeeping implications ofsocial media’.

An extract of the PROV proposed guidelines for archiving social metadata follows:

How should the record be captured?

Currently printing screenshots to .pdf and registering the resulting document in an Electronic Document and Record Management System (EDRMS) to record the necessary metadata is the  most accessible and expedient method of creating social media records. Necessary metadata includes who sent it (username and real name), date and time of sending, context and purpose of content, name of tool used to create it.

My first reaction on reading this was ‘this is mad!’ Perhaps the archives are under-estimating the amount of social metadata and media activity that is going on.  Taking a screenshot of every tweet for example would assume that you are not going to get thousands, whereas successful sites and topics such as Trove do get thousands and millions of interactions, which makes this unworkable from a staff resourcing point of view. Twitter is notorious for ‘disappearing tweets’ after a very short amount of time – sometimes less than a week because of the volume of activity that takes place.  This also puts pressure on to archive tweets at the time of creation.  You don’t have the luxury to go back and archive later. This suggested form of archiving only gives a screen-based image, which is not in context, not searchable, has no metadata, no timestamp, and is not authenticable. It seems there is money to be made if someone develops a simple software system to mechanically capture the tweet, its response and its components and safely and uniformly archives/indexes them along with descriptive metadata. The tool could also render the page "as it appears" and save it as a PDF if that is required.  

In April 2010 The Library of Congress rather bravelyannounced that it intended to archive all tweets since they began in 2006 to record the social fabric of the world and signed an agreement with Twitter and Google. In 2010 the Twitter archive was growing rapidly with users sending 50 million tweets a day. A year and a half later several news agencies tried to get a progressreport from LC without much success.  Other than trying to transfer the data from Twitter servers to LC servers the LC weren’t giving any detail on what technological developments they were creating to do the mammoth task. The task seemed to be growing bigger by the day with usage of Twitter increasing.  Currently 140 million tweets are sent every day.

A core element of the archive process should be that the data is kept in context with that it was referring to, and other elements surrounding it.  Most libraries that are keeping UGC and social metadata are keeping it in a separate layer to their own content in the database to protect the provenance, but may integrate it for public display.  If it is kept separate it can easily be stored, managed, and moved, but is at risk of becoming separated from the context it is related to. This is something libraries need to work out.  This will become more pressing in a few years time when existing services are migrated as part of their maintenance.  The UGC needs to be migrated in context with them.

On this topic I have more questions than answers. I think libraries and archives need to work together to take an active role in firstly encouraging mobilisation of social metadata -‘user takeout’, and secondly demonstrating how social metadata and social media activity can be archived. I see massive opportunities for start-ups to create archiving tools to bolt onto Facebook, Twitter, Youtube and Blogging software to meet the requirements of government archiving.
 
Photo: Prime Ministers Chiefly and Curtin chat on the way to work 1945. Bronze sculpture by Peter Corlett outside the National Archives of Australia, Canberra. Rose Holley

Sunday, 25 March 2012

Crowdsourcing: the crowd ‘rations’ its experience to make it last

When I am at work I feel I have far too much to do and that my ambitions can never all be achieved. This tends to make me feel despondent. However in crowd sourcing projects it is well known that providing way too much work and impossible goals is a very powerful motivator. Rather than leaving individuals in the crowd feeling despondent it drives them to put in even more hours.

I know from experience that this is the case because online volunteers working on the Australian Newspapers text correction have told me this.  Also after adding thousands of new pages to the service, surges in text correction would be observed.  This was a regular pattern.

Last week I was alerted to a great article in the Guardian about crowdsourcing and two good blog posts on crowdsourcing in cultural heritage.  All three articles are well worth reading and give some fascinating background to specific crowdsourcing projects. They all touch on the fact that the crowd wants to be given as much work as possible.

Crowdsourcing Cultural Heritage: the objectives are upside down, by Trevor Owens 10 March 2012  

Ben Brumfield noticed that in his transcription project one of his most valuable power users was slowing down on their transcriptions. The user had started to cut back significantly in the time they spent transcribing this particular set of manuscripts. Ben reached out to the user and asked about it. Interestingly, the user responded to explain that they had noticed that there weren’t as many scanned documents showing up that required transcription. For this user, the 2-3 hours they spent each day working on transcriptions was such an important experience, such an important part of their day, that they had decided to cut back and deny themselves some of that experience. The user needed to ration out that experience. It was such an important part of their day that they needed to make sure that it lasted.


Galaxy Zoo and the new dawn of citizen science, by Tim Adams, 18 March 2012

The volunteers only worry that the source of their obsession will dry up, and that they will run out of visible galaxies to classify. "In the beginning," Alice Sheppard said, "we all were enjoying it so much that we didn't like the idea of getting to the end." As it has worked out, more data sets have kept becoming available just as one tranche of images has been classified; now Sheppard believes that the work will continue to expand like the objects of its attention, "though no one seems quite sure how many galaxies are in the Hubble database?"

So the lesson we can learn from this is that we must give our crowd as much work and new data as we can.  We don’t want our crowd to have to ‘ration themselves’ because we haven’t left them enough work to do.

Photo: a worker bee eats the last crumbs of my sticky date pudding.

Sunday, 11 March 2012

Crowdsourcing transcription of handwritten archives


One of the big differences between libraries and archives is that libraries tend to have more of ‘the printed word’ whilst archives have vast amounts of ‘the handwritten record’.  While some libraries are getting up to speed with mass digitisation of books and journals and then being able to offer users full text searchable digitised items, this is still a distant dream for most archives.  Some archives are undertaking mass digitisation, but the second step – making handwritten records full-text searchable is a massive challenge.  The reason for this is in the technology and processing steps.

After scanning a ‘printed word’ page into an image file a piece of software called Optical CharacterRecognition (OCR) converts the image into searchable text.  The OCR works best with clean, clear, black and white typeface such as a word document or a book, not quite so well on old books and journals, and very poorly on old newspapers.  When it comes to converting handwriting it fails miserably.  It just can’t distinguish and convert handwriting to text in the way the human eye can.  Therefore archives can’t easily automate the second part of the digitisation process using OCR software like libraries can for the printed word.

If you at least get some OCR text from print that is readable and therefore searchable you can offer a service to users to full-text search the books or journals such as Google does. If the OCR text is poor there are some things you can do to improve it. You can encourage users of your service to correct the OCR text with a text correction tool so that the searching is improved, such as Trove does with the Australian Newspapers.

Unfortunately the only viable option open to archives to convert digital images into full-text searchable text is to use a manuscript transcription tool, in combination with harnessing the power of a crowd to do the transcription work.  The transcription work for handwritten records is much harder than for example text correcting old newspapers because the handwriting is often difficult to read, old fashioned, barely legible and not necessarily structured in lines or columns. There is often nothing to go on.

I recently stumbled across a blog all about manuscript transcription tools that is written by a software developer Ben W Brumfield in Texas. Ben developed his own software to transcribe his great-great grandmother’s journal. ‘FromthePage’ is now being used by archives because Ben has made it available open source.

A year ago he wrote an in-depthblog post that covered manuscript transcription tools under development, manuscript transcription projects in archives, and made some predications for future directions of manuscript transcription.  I am not going to repeat what he said here, I suggest you read the post in full.  He notes that software development in this area is still fragmented and young with no particular tools taking dominance. Most developed applications are being made available open source. A standout is ‘Scribe’ from the Zooniverse team, currently being used by both the ‘Old Weather’ project to transcribe maritime weather records and by ‘What’s the score’ project to transcribe music scores at the Bodleian Library, Oxford.

Before an archive implements a manuscript tool it needs to find out what it’s users  would most like to be easily full-text searchable from the vast vaults of all the content it has. It is important to find this out, because the crowd will only be motivated and swell in numbers if they really feel what they are doing is very important to a broad group of people and really matters either right now, or in the long-term and is also interesting.  They have to feel this before they will join in.  Once they have joined in there are other motivational tips you can do to keep them going.  Just implementing a manuscript tool is simply not enough.  You need to engage, watch, understand and learn from your crowd, for they hold the passion and power in their hands to make your project successful or not.


Photo by Rose Holley, outside Canberra Bus Station

Saturday, 3 March 2012

The digital game: helping librarians get digital jobs

A question I am often asked is “How can I become a digital librarian and get a job in this field?”
I wish I could say “Well the subject is well covered in a Library/Information Science courses, and there are lots of opportunities for you to gain experience.” But unfortunately neither is the case in Australia or New Zealand.  I have been mentoring and helping Masters in Library Studies students with their assignments on digital topics for the last 10 years.  I do this for two reasons. Firstly because I am naturally curious about the assignments and how close to reality they are, and secondly because I live in the hope that some or even just one of these new graduates may be inspired rather than discouraged with digital,  and end up becoming a digital specialist like me.  We are so short of digital specialists. 
I am disappointed that most library courses and degrees still offer the digital bit as an optional rather than compulsory part of the program, even though these days most libraries would be doing something they call ‘digital’, in the same way they all catalogue.  There is no Australian University course I am aware of that actually covers the whole breadth of digital topics at degree level for cultural heritage specialists (museums, galleries, libraries, archives) i.e. digitisation, digital delivery, digital preservation, data sharing.  However I am encouraged because the digital assignments from library courses that I am asked about are increasingly becoming more realistic and practical. They are moving on from theoretical questions about online catalogues and digitisation to topics such as utilising social media and digital preservation. But it is still hard for new graduates to find jobs, when they may have theoretical knowledge only and no practical experience in the field.  It is also hard to up-skill our existing librarians.
I was very interested therefore to hear about a new board game focusing on digital topics that was road tested at the DISH2011 conference.  I thought it held immense value as a tool for three things: graduate teaching; for up-skilling staff in an organisation; and for interview practice to get some of those tricky digital questions right.  The game is based on monopoly and covers the whole digital life cycle, including digitisation and digital preservation.  I was interested to see that some of the questions are ones I have actually been asked at interview.  Things like “What would you do if half way through your digitisation project the funding was cut?”  The game is created by the European DigCurv Project.  DigCurV brings together a network of partners to address the availability of vocational training for digital curators in the library, archive, museum and cultural heritage sectors in Europe.  These skills are needed for the long-term management of digital collections. There is a very good blog post with pictures of the game being played and some of the questions, so I won’t repeat them here.
At the moment the game is being refined and will be only available to European partners of DigCurv (some of whom would like it translated from English into their own language).  It would be great if copies could be obtained for the national Australian Cultural Heritage Institutions and Australian Universities offering Library/Archive/Museum degree courses.
There have been a number of organisations set up in Europe in the past to address training issues in digitisation and digital preservation.  Not all of these survived, many being based on short term funding.  The earliest I am aware of was in the UK in the year 2000, funded through revenue from the National Lottery.  50 million pounds was given away as ‘nof-digitise’ for organisations to start digitisation projects.  However it was quickly realised that training would be required before the digitisation and delivery could start and so short term national training courses were set up.  In 2001 the UK was the place to be if you were working as an information professional and wanted to learn about digitisation on the job and had got your hands on some of the nof-digi money.  Sadly in Australia and New Zealand we are still awaiting a financial windfall for digitisation on the scale we have seen from the European Union, French and Scandinavian Governments and UK Lottery Funds.  This means that we also haven’t developed the training we need and have no such equivalent organisation as DigCurv. I’m still hoping the proposed National Cultural Policy may address some of these things in 2012.
Photo from DEN Flickr stream:

Sunday, 26 February 2012

Crowdsourcing Australian Climate Change


In my last blog post I described how knitters and yarn enthusiasts were crowdsourcing knitting patterns from digitised Australian newspapers in Trove for use in a crowdsourcing site called Ravelry http://www.ravelry.com.  This week I wanted to give another example of the Australian Newspapers in Trove giving leverage to yet another crowdsourcing site. This time it’s for research into climate change and the site is called OzDocs. http://ozdocs.climatehistory.com.au/

Australian newspapers hold unique content, for example convict records and climate records. In Australia official weather records only began in 1908 when the Bureau of Meteorology was established. However there are weather tables and forecasts appearing in Australian newspapers from 1803 onwards.  The newspapers therefore provide 200 years of weather records.  Newspapers not only give tables with statistics of temperature, rainfall, winds etc, but also eye witness accounts of weather conditions such as floods, droughts and fire.

The citizens and politicians of Australia have a high and ongoing level of interest in climate change and how it is affecting our nation. A project investigating climate change is SEARCH: SouthEastern Australian Recent Climate History. It spans the sciences and the humanities, drawing together a team of leading climate scientists, water managers and historians in Australia to better understand south-eastern Australian climate history over the past 200–500 years. The digital newspapers in Trove are a fundamental part of SEARCH’s research process. However even though the newspapers are full-text searchable it is still a challenge to find and bring together in context the eye-witness accounts and the weather tables so that the temperatures, rainfall and other statistics can be transcribed into a research database. This is why the SEARCH project has established this month the OzDocs citizen science project, with a $10,000 grant from the University of Melbourne so that the public can help them. Basically the public are asked to find and tag historic newspaper articles on weather conditions, and transcribe useful weather statistics from historic newspapers into a database.

In 2010 Joelle Gergis, the lead SEARCH investigator spoke to me and said:

“Having all this information online and being able to quickly access it has been amazing. Being able to find weather tables from 1803 onwards in the Sydney Gazette is crucial to our research. Official records from the Bureau of Meteorology only began 100 years ago so being able to access the newspaper records which are earlier than this is really useful. The sources in Trove also show how weather events have affected society, with eye witness accounts of floods and bushfires. For example we have been researching the 1851 Black Thursday bushfires in Victoria.”

In 2010 as a locust plague swept across the south-eastern side of Australia the pilot volunteers for the project working at the State Library of New South Wales noted that weather conditions in 1825 were very similar (heavy rainfall followed by nice weather then a terrible locust plague) and found eye-witness accounts in The Sydney Gazette and New South Wales Advertiser, 24 March 1825 of a similar locust plague.

“Prior to the late rains the caterpillar, that old enemy to the agriculturist interest of the Colony, made its appearance ; but, upon the visitation of the heavy showers, their ranks were consider ably thinned. However, since the present enchanting fine weather has again set in, the number of these destructive insects has increased to an unparalleled extent, covering whole fields in their course, which in some spots seemed to be towards the South, in a line from East to West. Wherever they make their appearance, the most complete destruction immediately follows. Upon Captain Campbell’s estate, in the district of Cooke, they were supposed to be at least two inches in height.”

This month I caught up with Joelle again to see how the newly released crowdsourcing part of the project is going. She said:

“We now have over 100 volunteers who have contributed over 4000 articles. The database will be searchable in our next release. It will be the country’s first publicly searchable database of climate information using a diverse collection of pre-20th century historical records. The database will give easy access to information for researchers, organisations, government departments and the public.”

The scope of OzDocs work has now been expanded to include not only the digitised newspapers but other resources that are held in the State Library of Victoria, State Library of New South Wales and the National Library of Australia.  These pre-date the newspapers by another 100 years and go back to the 1700’s. Joelle said:

“Our OzDocs volunteers will be working their way through logbooks of the first European explorers, governors’ correspondence, early settlers’ diaries, newspapers and the works of 18th and 19th century scholars.”

On of the questions the SEARCH project hopes to answer is what the South East region of Australia’s ‘natural’ climate has been like since 1788. This may ultimately help to refine current climate models, allowing more accurate climate change estimates to be developed for the future.  The lack of records in a consistent accessible format before 1900 is currently making this difficult.

For the project to be successful the volunteer numbers really need to increase a lot.  At the moment this is still quite a small scale effort compared to the knitting project Ravelry that I reported on in my last blog.  Volunteers can join by accessing the OzDocs site. http://ozdocs.climatehistory.com.au/



Friday, 24 February 2012

Crowdsourcing knitting patterns


Wordle showing the most popular search terms in Trove.
I have been project managing the Australian Newspapers Digitisation Program, The Australian Newspapers service and Trove at the National Library of Australia for the last 5 years.  The content of Trove – a free discovery service for Australian content, is now massive with a total of 250 million items from different organisations around Australia.  This includes archives, pictures, books, music and newspapers. There are 62 million full-text articles coming from Australian Newspapers 1803-1954 and the Australian Women’s Weekly 1932-1982, which the National Library of Australia has digitised. About 100,000 new newspaper articles are being added each week to Trove. 

Despite the content of Trove being varied over 80% of Trove usage and engagement still revolves around digitised historic Australian newspapers. They are the most used content the National Library has ever had, eclipsing everything else with usage continuing to increase. One fifth of the Australian population are regular users of Trove (4 million people). There are about 50,000 searches every hour. Over 40,000 online volunteers have corrected over 58 million lines of newspaper text to help improve the searching.  This is the crowdsourcing aspect. Today there will be more than 100,000 lines of newspaper text corrected by users, this week more than 10,000 items tagged by users and this month 2,000 comments added to items by users.
I’m often asked what the most accessed items are in Trove.  Unfortunately I cannot answer this because it isn’t logged on our servers.  However I do know that newspapers are used more than any other content. I keep an eye on the most used search terms using Google Analytics to get a feel for what people are looking for.  In fact anyone can look at search terms as they happen, second by second, by clicking the link immediately above the Trove search box. For a very long time the top search terms have been Smith, George, death, birth, hanging, suicide, murder, cricket and gold, and also anything topical e.g. Lionel Logue (the Kings Speech). The most popular articles seem to be births, deaths and marriages and articles on murders.  However I have seen a recent trend whereby the terms “knitting pattern” and “knit + cast on” have knocked “death” off the top spot.  I decided to look into this a bit further. I was fascinated to discover that the crowdsourcing in Trove is inter-connecting with another crowdsourcing project.  It’s for knitters and is called Ravelry http://www.ravelry.com

Ravelry is a place for knitters, crocheters, designers, spinners, weavers and dyers to keep track of their yarn, tools, project and pattern information, and look to others for ideas and inspiration. The content on the site is user- driven and created by the knitting community. Ravelry lets you keep notes about your projects, see what other people are making, find the perfect pattern and connect with people who love to play with yarn from all over the world in forums.
The site was started by Jess who had been a knitter and a blogger for a while. She knew that there was all this great information out there from other fiber lovers – but with the growing number of crochet and knitting blogs, finding that information just kept getting harder. It was getting frustrating for her to try and find information about the patterns and yarns that she was interested in using. Her programmer partner Casey thought that he would be able to build a website that could solve her problems, so they started working on it together, introducing it to a few friends at a time.

A key part of the site is the database of knitting patterns, gathered together by the community, described and catalogued by them, and then knitted by them. The user community can favourite them, add comments, add patterns to projects and lists to do. They often photograph the end results and add these to the database.  They can seek help from other knitters on patterns, yarns, techniques and designs.
The site is free but does require a login to look at the patterns.  It is proving immensely popular. In the first weekend 15,000 knitters had signed up.  Apparently quite a lot of these happened to be librarians. * In July 2010 Ravelry appealed to the community to both find and describe patterns. In one week 23,500 users categorised and assigned metadata to 160,000 patterns. The advanced search which draws on these fields for faceted searching is quite amazing, and quite frankly leaves most library catalogues for dead.  Facets include availability, category, yardage, gender, source, fibre, needle size, rating, fibre, difficulty, language and more.  I was quite stunned by this because this it is one of the few crowdsourcing projects I have seen that has very successfully engaged a crowd to help assign metadata to records to the highest possible level.  Cataloguing is a task that most cataloguers and librarians think cannot be done well by anyone except themselves, besides which it would be far too boring to attract people’s interest.  However the knitters can clearly see the value in adding descriptive metadata.  For example by adding yardage  or meterage required they can easily find out by searching on that field what patterns they can knit when they only have x yards left of wool. 

So how does Trove come into all this?  Well, as knitting regains popularity and we see the resurgence in ‘retro’ fashion from yesteryear the knitting community are falling with glee on digitised historic Australian newspapers and the Australian Women’s Weekly, particularly from the 1950’s.  Someone has helpfully added the instructions into Ravelry for how to find vintage knitting patterns in Trove, which is search for knit+"cast on" or knitting patterns (now one of the top search terms – see the wordle above).  If you do this in Trove you get nearly 73,000 results for knitting patterns.  Most newspaper included at least one pattern a week. Of all those patterns the community has chosen to add some of the more popular ones into Ravelry so more community engagement can happen.  So far 290 have been added from Australian newspapers and the Australian Women’s Weekly.
The two screenshots from Ravelry below show firstly a classic number ‘a cosy cardigan’ which appeared in the Sydney Morning Herald of 1953. It has been favourited by 145 people, one has knitted it and 93 people have added it to their queue of things to knit next. The person who has knitted it has added notes and instructions on how they did it with a colour picture of the finished garment.

The second shot shows that the most favourited pattern added to Ravelry from the National Library of Australia’s digitised Australian Women’s Weekly collection is… wait for it…… the ‘Elegant Elephant’. It has been favourited by 690 people,  rated 4 out of 5 and easy to knit, knitted by 21 people in a variety of colours, and 174 more people intend to knit it soon.  If you click on the pattern you will see uploaded photos of finished knitted elephants….


If you don’t want to log in to Ravelry get an overview by watching this 6 min video on Ravelry.
This is a very interesting example of re-use of material from old newspapers, one that was not even considered when newspapers were digitised.  Ravelry is an outstanding site offering community engagement and crowdsourcing that has really impressed me. I love the advanced pattern search by facets.  It clearly shows that for some items users don’t want a dumb it down simple search box. They want ADVANCED SEARCHING, MORE DESCRIPTIVE METADATA AND FACETS!  They are prepared to add the descriptive metadata themselves.
The only thing I have ever knitted myself was a pink and blue tea-cosy for my mother as a present when I was 14. My mother was a teapot collector then, but interestingly only had 2 tea cosies.  Knitted tea cosies are becoming popular again.  However I am contemplating knitting the Australian Women’s Weekly ‘Elegant Elephant’, maybe in pink, my favourite colour? At least if I get stuck I know I will be able to get online help in Ravelry, and it looks easier than a tea cosy!
* I acknowledge the use of Nyssa Parkes article ‘Fibre FRBRisation’ in the November 2011 issue of Incite Magazine.  The statistics on Ravelry user engagement come from this article.

Sunday, 19 February 2012

Digital protection of indigenous knowledge



The governments of New Zealand and Australia are unfortunately not as advanced as India in respect to protection of indigenous knowledge.  It is hardly believable but India has not only tackled the issue but also partially solved it digitally with an online database.  This post will have a look at this important topic in more detail.  It is relevant for libraries, archives, galleries and museums because they are collecting, storing, describing, handling,  loaning, digitising and displaying items which are classified as indigenous knowledge such as:
  •  moveable cultural property
  • literary and artistic works (including music, dance, song, ceremonies, symbols and designs, narratives and poetry)
  • scientific, agricultural, technical and ecological knowledge
  • human remains
  • sacred sites, burials and sites of historical significance
  • documents of Indigenous peoples' heritage (including film, photographs, video and audio recordings, and archival collections).
The problem with protecting indigenous knowledge is that most countries in the world have difficulties reconciling locally indigenous traditions, laws and cultural norms with predominantly western legal systems, effectively leaving indigenous peoples' individual and communal intellectual property rights unprotected. Also most countries do not have specific legislation or systems in place to protect and therefore prevent misuse, or commercial use of indigenous knowledge. It is usually up to individuals or tribal groups to take costly, long running and often unsuccessful court action to protect their knowledge and intellectual property. This is not how it should be.
The Wikipedia entry on Indigenous Intellectual property gives a brief overview of the subject which is a topic of international concern.  It notes two important declarations:

New Zealand: Mataatua Declaration on Cultural and Intellectual Property Rights of Indigenous Peoples (June 1993)
150 delegates from fourteen countries, including indigenous representatives from Japan, Australia, Cook Islands, Fiji, India, Panama, Peru, Philippines, Surinam, USA and New Zealand
·         Affirmed indigenous peoples' knowledge is of benefit to all humanity.
·         Recognised indigenous peoples are willing to offer their knowledge to all humanity provided their fundamental rights to define and control this knowledge is protected by the international community.
·         Insisted the first beneficiaries of indigenous knowledge must be the direct indigenous descendants of such knowledge.
·         Declared all forms of exploitation of Indigenous knowledge must cease.
Section 2 of the declaration asks State, National and International Agencies to:
·         Recognise that Indigenous peoples are the guardians of their customary knowledge and have the right to protect and control dissemination of that knowledge.
·         Recognise that indigenous peoples also have the right to create new knowledge based on cultural tradition.
·         Accept that the cultural and intellectual property rights of Indigenous peoples are vested with those who created them.

Australia: Julayinbul Statement on Indigenous Intellectual Property Rights (November 1993)
A meeting of indigenous and non-indigenous specialists agreed that indigenous intellectual property rights are best determined from within the customary laws (Aboriginal common laws) of the indigenous groups themselves.  These laws must be acknowledged and treated as equal to any other systems of law.
·         Indigenous Peoples and Nations reaffirm their right to define for themselves their own intellectual property, acknowledging the uniqueness of their own particular heritage
·         Indigenous Peoples and Nations declare that we are willing to share [our intellectual property] with all humanity provided that our fundamental rights to define and control this property are recognised by the international community
·         Aboriginal intellectual property, within Aboriginal Common Law, is an inherent, inalienable right which cannot be terminated, extinguished, or taken .. Any use of the intellectual property of Aboriginal Nations and Peoples may only be done in accordance with Aboriginal Common Law, and any unauthorised use is strictly prohibited.

Examples of indigenous knowledge being protected
1. New Zealand: The haka
Maori’s have been trying to defend their rights to the haka dance for over 10 years.  Ka Mate is the most widely known haka because it has traditionally been performed by All Blacks rugby teams at the opening of international games. It is agreed that Ka Mate was composed by Te Rauparaha, war leader of the Ngāti Toa tribe (iwi) of the North Island of New Zealand.
Between 1998 and 2006, the Ngati Toa iwi attempted to trademark Ka Mate to prevent its use by commercial organisations without their permission, but in 2006 the Intellectual Property Office of New Zealand turned their claim down on the grounds that Ka Mate had achieved wide recognition in New Zealand and abroad as representing New Zealand as a whole and not a particular trader.  However in 2009, as a part of a wider settlement of grievances, the New Zealand government agreed to:

"...record the authorship and significance of the haka Ka Mate to Ngāti Toa and ... work with Ngāti Toa to address their concerns with the haka... [but] does not expect that redress will result in royalties for the use of Ka Mate or provide Ngāti Toa with a veto on the performance of Ka Mate...".
In March 2011 a few months before the Rugby World Cup started in New Zealand the NZ Rugby Union came to an amicable agreement with the Ngati Toa not to bring the mana of Ka Mate into disrepute. In one of the final games France were fined $10,000 by the International Rugby Board for advancing towards the haka, but I think this was more to do with protecting rugby rules than the haka itself. 
Matiu Rei, head of the Ngati Tao Maori tribe said in October 2011
“We are not seeking compensation, we are seeking recognition.”
2. India: Yoga
For more than 10 years India watched as western governments granted patents, trademarks, and copyrights to what was India’s indigenous knowledge, for example yoga and herbal cures. The U.S. Patent and Trademark Office alone issued 150 yoga-related copyrights, 134 patents on yoga accessories, and 2,315 yoga trademarks. Yoga is big business around the world and is estimated to make $3 billion a year in America alone. The Indian Government did not stand by and do nothing, they decided to take action in the form of the Traditional Knowledge Digital Library (TKDL).

The aim of the TKDL is to make digitally available the indigenous knowledge of India in multiple languages, so that Patent offices can search the knowledge and then reject patents applications that are actually traditional knowledge.  India is quite lucky because things like yoga and herbal medicine are actually well described in ancient texts.  But these are in different  languages such as Sanskrit, Urdu, Arabic, Persian, Tamil (usually not English), hard to get hold of in hard copy, and not widely understood by the western world or patent examiners.  TKDL breaks the language and format barrier and makes available this information in English, French, Spanish, German and Japanese in patent application format, which is easily understandable by patent examiners. TKDL is thus a tool providing defensive protection to the rich traditional knowledge of India. In June 1999 the World Intellectual Property Organization (WIPO) and the Standing Committee on Information Technology (SCIT) recognised the need for developing countries to create Traditional Knowledge (TK) data bases. The concept of the Indian TKDL was formed in 2001.

By August 2011, 150 books on yoga, ayurveda, unani, siddha and natural medicines from multiple Indian languages had been digitised and transcribed into 4 European languages and Japanese.  In addition 1,300 yoga 'asanas' had been documented making them public knowledge. Around 250 of these `asanas' have also been made into video clips with an expert performing them.

But this hasn’t stopped self-styled yoga gurus such as Bikram Choudhury in the USA still trying to patent ‘Hot Yoga’, a set of 26 sequences practised in a heated room. Bikram and the TKDL made headline news last year because of this. Boingboing reported in March 2011 that apparently Mr. Choudhury was threatening to sue people teaching a popular style of yoga he claims to have invented and copyrighted. He also reputedly said "Because I have balls like atom bombs, two of them, 100 megatons each. Nobody fucks with me."   Perhaps he has not finished reading the full library of indigenous yoga knowledge, since I’m sure this is an attitude not encouraged by yogis?
Dr V P Gupta, who created TKDL, said "All the 26 sequences which are part of Hot Yoga have been mentioned in Indian yoga books written thousands of years ago." He added, "However, we will not legally challenge Choudhury. By putting the information in the public domain, TKDL will be a one-stop reference point for patent offices across the world. Every time, somebody applies for a patent on yoga, the office can check which ancient Indian book first mentioned it and cancel the application."

These two examples of defending indigenous knowledge and the quotes by the leaders reinforce the statements in the declarations, namely:

“indigenous peoples are willing to offer their knowledge to all humanity provided their fundamental rights to define and control this knowledge is protected by the international community.”
Libraries, archives, galleries and museums are part of that community and are often seen as gatekeepers of knowledge. We should be pro-actively supporting, encouraging and enabling this outcome.

What is happening in Australia?

In Australia there are many examples of non-indigenous Australians inappropriately using indigenous knowledge and artworks without permission. Most recently this has revolved around sacred rock art.  In April 2009, the Australian Government adopted the UN’s Declaration on the Rights of Indigenous Peoples. It states that Indigenous people have the right to maintain, control, protect and develop cultural heritage, traditional knowledge and traditional cultural expressions including oral traditions, literature, designs, visual and performing arts. It also includes the right for Indigenous people to maintain, control, protect and develop their intellectual property over such cultural heritage, traditional knowledge and traditional cultural expressions. However we currently don’t have the infrastructure or systems in Australia to do this easily or effectively.

In April 2008 at the Australia 2020 Summit Terri Janke proposed the establishment of a National Indigenous Knowledge Centre (NIKC). This idea was followed through with a feasibility study into a NKIC which was submitted to FaHCSIA in October 2011. In 2009 Terri Janke wrote her own report ‘Beyond Guarding Ground: A vision for a National Indigenous Cultural Authority’. Both of these reports require legislation and a system to be established for the protection of indigenous intellectual and cultural knowledge. Part of this infrastructure could be a digital library database.
Some of the suggestions for the NKIC are that it would build on the existing role of the Australian Institute for Aboriginal and Torres Strait Islanders Studies (AIATSIS):
·         Become a reference point for Aboriginal and Torres Strait Islander culture.
·         Engage in research to harness traditional knowledge to support sustainable management of country.
·         Support the education and understanding of indigenous culture and affairs across Australia and preserve indigenous heritage.
·         Become a national gathering place for the celebration and discussion of indigenous culture in a physical or virtual sense.

The AIATIS response to the National Cultural Policy quite rightly questions if and how protecting indigenous knowledge will fit into the proposed NCP.
Further Reading:
Michael Davis, Indigenous Peoples and Intellectual Property Rights, 1996 Research Report

ATSILERN Protocols: Guidance for libraries, archives and information services in appropriate ways to interact with Aboriginal and Torres Strait Islander people, their culture and heritage.  
The Australian Wattle, photo by Rose Holley