Monday, 8 September 2008

Ten Years, Five Lessons

Google is 10 today(ish) and has just pushed it's digitisation programme even further. After a decade of ubiquitous multicolour rule what does Google's success teach us? I've listed five quick lessons which libraries can draw from Larry and Sergey's approach:

1: Services not Products

Google has taught us that you don't actually need to sell a product to make money, only a service. It's especially good if this service is advertising. In fact, selling a simple service (Ads) through multiple services (search, email, calendar, etc) greatly increases both your revenue and your respect in the community. 

Libraries are already ahead of the game in this area as we're all about fantastic services aimed directly at our customer, what we have to remember is that we should avoid attempts to comodify those services into 'products' which limit our users and they way in which they can interact with them.

2: People Are Interested - if you are too.

Ever heard a rumour about a 'new' Google product? Or closed beta? The reaction from Android developers is a prime example of how keen people are to use Google's technology. The public are willing to invest time and energy into testing incomplete products primarily because they think that their feeback might be listened to and lead to a better results at the end of the day. 

The lesson here is simple - listen! It's easy enough to pay lip service to user feedback but true user engagement requires concerted effort but yields real results, we should always be happy to receive comments and suggestions, take them seriously and give credit where it's due.

3: Give Your Staff Space to Develop

Google is famous for being a great employer, and part of their approach is to give staff '20% time' in which to pursue their own projects. This doesn't mean slacking off, staff have to present on this work and take comments from colleagues. 

Libraries know full well that training is key to keeping staff providing great service but how much time do we give our staff to explore their own work-related interests? Let's not pretend that any of us have Google's capacity or funding but the lesson still stands - great staff need space to develop themselves and we should encourage that development whilst also ensuring that it's relevant by providing an environment where staff can present their work and comment on others'.

4: Employ Creative and Driven People

Ok, now here's a no-brainer - Google employs a large number of PhDs because they feel that they're creative and driven. Because their entire service portfolio is built around a core of information searching creativity is key, and new ways of working are essential to preserving their edge - the technology and development are essentially just adjuncts to the core information base. 

Libraries still need to pay attention here - it's just not about hiring chartered Librarians, it's about recognising the need to hire and develop creative people who are driven to work with information.  Because it's information, and we providing creative ways of working with it, that is at the centre of our business and everything else is adjunct

5: Change, but Don't.

Isn't it great that Google offers new services all the time? That the homepage logo reflects different events and that your gmail space is constantly ticking higher and higher? But deep down, under this ever changing fascia, is a stable core. A large part of the respect Google has is because it hardly ever compromises on it's central values. 
Now this is tough for any public institution - not least libraries. We all have to deal with changing wider strategic frameworks which sometimes make it hard for us to pursue our own long-term plans. However, we need to make sure that we retain our integrity whilst adapting and evolving to meet new challenges. 



Comments?


Sunday, 7 September 2008

What About Web Archiving?


Some of my role involves working on a large web archiving project. One of the questions that's often asked by those who use the UK Web Archive, Internet Archive and the like, or are looking to start web archiving themselves, is about the 'quality' of archived web pages in comparison to the original.

The key issue to get across is that website aren't 'things' you can locate, pin down and put in a box - even the most static collection of HTML pages exists in the continuum between server, network, browser and screen. More simply, there's no such thing as a website1.

Web archiving is more akin to photography; most of what we do is to make copies of sites, trying to make them fit with our own individual construction of them at the particular moment and conditions of archiving. Similarly, the problems associated with archiving dynamic content can be compared to taking photographs of a city: images of 1940s Cardiff don't capture the entirety of the city but do provide insights and representations of a transient experience, just as an archive of Amazon.com will only give a glimpse into the constantly moving community beneath the surface.

The question we, as the custodians of information, have to ask ourselves is: is this enough? Should we strive for completeness even if that means emulation of servers and proprietary software and when all we'd be keeping is a shell of the original (akin to a digital St Fagans)?

With web archiving we need to engage with the content more directly. At the moment it seems cheaper to keep things than to select them but this isn't always going to be the case. If we take an object in to our collections we're looking to keep it in perpetuity - and the lifetime costs of keeping that single file will be infinite (this is true, of course, for physical things as well).

For domain level harvests it's clear that captures will only ever be superficial. It's my view that there will always be a place for selective archiving - both of sites which fall within collection policies of organisations but outside of the domain and in terms of effort put into the processing and checking of specific sites selected to be of special interest.

It's tempting to equate domain harvesting to the collection of printed material through legal deposit however there are some key differences. Firstly, the range of material (and therefore preservation requirements) of printed material is limited, not necessarily small but limited whereas web content will contain every obscure file format you can think of stored in all kinds of different ways. Secondly, there is a significant cost in printing material which reduces duplication (although, of course, it also reduces the range of material - for better or worse).

By putting our information on the web we're all becoming our own librarian-archivists (hooray!) - although we've not yet taken the next step and become records managers. (My Gmail currently tells me that I'm currently using 534 MB (7%) of my 7081 MB - if this keeps up I'll never need to delete anything and, as long as search technologies keep up, I will be able to find them again too.)

There is a movement within our information society towards keeping all the iterations of content as part of the services themselves. We're constantly seeing statistics which suggest that the amount of information available is increasing exponentially but, if Wikipedia and blogs are anything to go by, most of this will be drafts, previous versions and backups!

Similarly the web is notoriously difficult to time. Print and even 'formal' e-Journals are published on a specific schedule whereas blogs and other forms of social content can be started, flourish and die out in a short period of time (perhaps even between large-scale harvests). Part of our engagement with content has to be a realistic and positive in our engagement with content creators.

If you look at the short history of the web you can begin to see a shift from silos of content (personal and corporate websites, each with their own domain) to portals (where the silos are connected by indexes and directories) to dispersed content (where photos live on flickr, blogs live on wordpress.com, updates on twitter, videos on youtube and links on delicious)- it's easy to believe that this dispersal was the inevitable consequence of hypertext.

Of course this changes the way we archive and - more importantly - provide access to archived sites. If all the information on particular subjects is spread across user pages on multiple websites which are reliant on the major search indexes like google to link them together then we need to think about not only capturing the separate content but also the links between them. At the moment we work primarily at site level, preserving the links between pages but treating them as immutable objects - in the future we will need to let the harvesting agents roam freely, capturing snippets of content which make up a web2.0 website.

Ultimately, the archiving of the web is a positive thing and it's certainly a great area to work in, I know that every site we archive is a resource preserved for future research. The challenges that face us shouldn't put us off adding these important cultural assets to our collections but we do need to begin to engage with them before they seem insurmountable.

1 In fact it's this facet that really attracted me to working on Web Archiving in the first place. A substantial portion of my doctoral research was centred around the social construction of the web and the fallacy of the monolithic websites. If you ever want a long and overly detailed conversation about constructivism and the web you know where to come...

Wednesday, 3 September 2008

Antisocial E-Books?


On the day that Sony launch their E-Reader in the UK (via Waterstones stores and online) I read this wired article which suggests that E-Books are 'antisocial' - admittedly only as a turn off to girls.

Interestingly enough, I've been conducting my own 'long term test' on the iLiad E-Book platform - not just me, in fact, but also my wife. We both devour books but she's found the ability to buy books online, on demand, very useful. Perhaps the problem for Charlie Sorrel is that he thinks that a iPhone is a good way to read....

Tuesday, 2 September 2008

SOPAC, NOPAC

It's not uncommon for users to mix up a library's OPAC and website and why not? Gone are the days when a database of any description was kept distinct from it's host website. In fact, libraries are some of the worst offenders in terms of drawing artificial lines between services.

So it's of real interest to me that Darien Library has just launched it's combined website/OPAC. Essentially they've built a Drupal module that ties, via a connector, to any ILS - more here. It's a really neat move, and one which I'm sure we'll all be watching with interest.

Microsoft Surface


Engadget has reported that Sheraton Hotels have started to install a few Microsoft Surface stations in their hotels. I think these are one of the coolest things ever for libraries. The Object Recognition element should be fantastic for anyone with digital content which can be used to enhance physical material.

Forget the capacity for actually accessing, flipping pages and interacting with digital content for the moment (although this does take the 'Turning the Pages' idea to a whole new level). Instead I can see the system recognising ISBN'd material lain on top of it from the barcode and unique material from the (equally unique) back surfaces.

From there you can see how interpretive data created by libraries and user comments can be lain around the surface, with the physical item linked to all kinds of complementary digital content - whether it be "See also..." type information, biographies or user reviews. In addition, library services can be marketed by laying them out next to relevant material.

There's not much information on how cultural organisations might use this kind of technology - Microsoft seems to be pitching this at Bars and, obviously, Hotels. It's a shame because I can honestly see museums loving this as a concept - and libraries should too.

Google Chrome Announced

Apparently today will mark the launch of Google Chrome - Google's own open source web browser which purports to be lightweight with better resistance to tab crashes taking down the whole browser. This is interesting for three reasons:
  • Google has put significant support into the Mozilla Foundation (home of Firefox).
  • Google chose to develop their own open source browser to fix what seems to me as a quite 'small issue' - why not fix it in one of the existing open source browsers?
  • The announcment echoes Google's web-app philosophy; " To most people, it isn't the browser that matters. It's only a tool to run the important stuff -- the pages, sites and applications that make up the web." which is could be an interesting look at the future of any browser wars: IE as a 'fully featured' browser, Chrome as a lightweight 'window to the web' only and Firefox and Opera somewhere in the middle (although, with all the user add-ons bloat is a severe possibility for Firefox).
All in all, there'll be a lot of talk on the web about a new browser war even though this may well be another short-term 'product to fix a problem' from Google.

(Typical Google, they launched with a comic.)


UPDATE: Ok, Chrome is fast! And on my little 12" laptop it's very nice - I can really see this as the start of a web-OS. (Oh, and all my websites work with it - yay!).

Monday, 1 September 2008

Online Poster Goodness

A fair few library bloggers have been singing the praises of the ALA Online Poster Maker which allows anyone and everyone to create their own poster to compliment the ALA's own promotional efforts.

What a great idea! Whilst 'outline' posters are available for lots of different library promotional programmes, allowing the public to make use of them is a new thing. The National Year of Reading website (for England only) allows you to play with the logo and 'design your own' - which is a nice touch - but the ALA interface really benefits from its simplicity (although, a few more designs would be nice).

Plus, "READ" is such a great slogan (if such a simple thing counts as a slogan) - it gets round the whole "shouldn't we say that libraries have more than books?" temptation that seems to raise it's head whenever more than three librarians get together. More please!