Showing posts with label code4lib. Show all posts
Showing posts with label code4lib. Show all posts

Wednesday, June 2, 2010

code4lib NW: digital initiatives presentation

I'm going to be giving a talk at the code4lib Northwest conference on Monday and I thought that I'd put together my thoughts for it here. The theme is creating a digital initiatives program at a liberal arts college.

I believe that we're in the middle of a paradigm shift for academic libraries. As content shifts to the network and as discovery is disintermediated from the library, the work needed to support the library's traditional roles as buyer, archiver, and gateway to information is slowly diminishing. Concurrent with this trend is the rise of potential opportunities for new roles. The new roles that we hear most about, perhaps, are instruction in information literacy/fluency and the management of digital assets.

Within the digital assets area, I think that the development of thematic digital projects is a particularly fertile area. When faculty approach teaching, their own research or a collaborative research project, they now have opportunities to do sophisticated things with information resources. Some examples include:
  • a ceramics professor uses his connection with hundreds of artists around the world to create digital collection of ceramic art images (accessCeramics)
  • an environmental studies program uses social bookmarking software to develop a virtual library of research resources on environmental issues at various geographic sites (L&C Environmental Studies)
  • using an online digital collection, students observe the process of writing a poem by viewing its evolution from an initial draft through various stages of revision and eventual publication (William Stafford Archive)
  • students do primary historical research on a topic and contribute their work to a publicly available online collection of historical research (History Engine)
  • a biology professor studying hundreds of different spider species wants displays her findings geographically on Google Earth and Google Maps
I think academic libraries should develop the expertise and capacity to support these kinds of thematic projects. This is relatively new territory for libraries. We are more comfortable managing our own internal catalogs and collections and providing standardized services. The institutional repository is really not much of a leap for us: it's an attempt to go out and uniformly collect and catalog objects. Many digital collections don't go much further than digitize and catalog a defined set of photographs, manuscripts, etc. held in the library.

But I think there are more interesting opportunities when we actually wade out into the messy world of teaching and research and offer up our expertise at organizing information. A way of doing this is to establish some kind of a digital initiatives program that faculty can engage with directly. We see this at large institutions such as University of Virginia and Columbia, but also now increasingly at liberal arts colleges like Hamilton, The University of Richmond, and Kenyon. The programs at these institutions in one way or another offer support to faculty for teaching or research related digital projects.

At Lewis & Clark, it is my goal to develop a digital initiatives program from the library that provides support for academic projects that involve sophisticated information management problems. Our library has been ramping up support for digital initiatives for the last seven years or so. We began with initiatives to digitize student theses and put some of our archival collections online. We also developed a digital image collection to support arts and humanities instruction. As a side project a few years ago we started working with an studio art professor to develop a collection of images of contemporary ceramics: accessCeramics. The project has involved all sorts of interesting technical and organizational challenges. Above all, it has been a a truly collaborative effort between several library staff and the faculty.

The success of this project made me want to expand the library's digital initiatives to include more collaborative academic projects. However, I wasn't really sure if demand existed for these kinds of projects. To find out, I made an effort this spring to speak with at least one faculty member in each department at Lewis & Clark about potential digital collaborations with the library.

After talking to about 15 faculty, I have a good handful of ideas for projects, some more imminent than others. They include: web mapping wine and foi gras regions in Oregon and France for an anthropologist, digitizing, translating, and annotating collection of published documents from an early 20th century Moroccon Jewish community for another anthropologist, creating an online map to accompany a book about Mount Fuji by a historian of modern Japan, developing an online archive of historical depictions of Lewis & Clark's slave York for a communications scholar, building maps and phylogenetic tree of various spider species for a biologist, developing an archiving system for recitals for the Music department.

A few trends emerged in my conversations. Even though I unearthed some real cool ideas for projects, traditional scholarly communication methods still seemed to dominate faculty thinking when discussing their academic work. Digital work in non traditional forms is bonus work in their eyes and those working for tenure are hard pressed to find the time for it. Geocoding and web mapping were popular projects. There is a seemingly insatiable demand for assistance with generic web design, development, and upkeep by faculty for their personal pages and pages related to their research and other academic endeavors like conferences. Among the handful of scientists I spoke with, I had been expecting to find a need for long term preservation and access of scientific data, but I found that most of these researchers had disciplinary level destinations on the internet for the research data that they believed was critical to preserve and disseminate. The scientists' interests were in making their scientific work more accessible to a broader audience via the web rather than than data preservation.

Right now, I'm still in the process of talking to more faculty to get a feel for the types of projects that might be useful for their teaching/research. My goal is to come up with a digital initiatives program that offers certain services based on what faculty want and what we can do. As of now, I'm thinking that our digital initiatives program should prioritize projects that involve a sophisticated information management problem and provide a relatively broad impact, especially among students of our institution.

I'm hoping that we can provide a couple different levels of support: one level would be consulting. We would offer our expertise to jump start someone on an information management project: building a Google map, organizing a wiki or a database, teaching a class about mashups. A second level of support would be project based: we would take responsibility for a finite project: developing part of a website, a database etc. The top level of support would involve making a long term commitment to a digital resource or collection, the kind of commitment we now have to accessCeramics.

The other critical aspect of this initiative is to find the human resources internal to the library to support it. To do so, we need staff with expertise in a few key areas:
  • metadata profiling
  • metadata assignment
  • information architecture
  • web design
  • web programming
  • project management
  • grant writing
In our shop, we have a Digital Services Coordinator at Watzek with voracious appetite for web development and programming who likes nothing better to be given free reign on a challenging, creative project. As our portfolio of projects has grown, however, he is stretched for time. Fortunately, we have a cataloging assistant who is emerging as a talented designer and creative metadata strategist, and due to some increased efficiencies in our tech services operations, we've been able to reallocate some of her position to design and digital project metadata work. Special Collections and Visual Resources do most of the metadata work on digital projects in their areas. I do some web programming but recently, more project management and grant writing.

Nevertheless, to expand further, we'll need to find more time and expertise among our staff in a budget neutral environment. As manager of our cataloging/acquisitions (Collection Management Services) operations, I'm looking hard to find ways that we can save time on things like copy cataloging, serials checkin, and govt. docs. Even if we can free up time, developing the skills among staff in areas needed to support digital projects remains a challenge. Making the switch from traditional cataloging to digital project metadata work isn't a huge leap. But finding staff who can do work like web design and web programming is a bigger shift. As a manager, if there are any signs of that kind of aptitude in existing employees, I'm looking to foster it.

There are some potential pitfalls to a digital initatives program like this. The first one is overloading library staff: we have a lot to do just keeping our own house in order, and my faith in network level services lightening that load may be a bit naive. Many institutions would view this kind of function as more appropriately filled by a unit independent from the library, perhaps in IT or instructional technology. In our case, I want to keep IT in the loop at all stages, make sure that we don't overlap our services, and partner when possible. Another criticism is that the services provided by this program could represent a very uneven distribution of library resources to support the interests of particular faculty. As a library, we're used to providing relatively broad based transactional services like checking out books and answering reference questions that benefit a wide swath of the institution. This is something to be conscious of when designing this type of program. The other side of it is that these projects are often immediately beneficial to current students and faculty whereas many library-centered or institutionally-oriented digitization projects have somewhat diffuse, long term benefits.

I think developing this kind of program has some compelling advantages. These kinds of services can really advance faculty research and teaching and help assure the continued relevance of the library. If we help faculty do the heavy lifting needed to do more advanced projects and get more grants and recognition, they will back the library that much more. Administrators like deans and presidents love to see this kind of innovation as it can truly advance an institution's core mission, and in fact many of the centers I mentioned above have been started in top-down fashion by such officials.

This kind of creative work can help energize a library and connect library staff to the academic mission of the institution. I also think there is a tie in with library liaison work and library instruction. Collaborative work with faculty on collection development, instruction sessions, and information literacy can foster connections that lead to some of these digital projects. I'm hoping that reference/instruction librarians can be partners in this endeavor as well, especially when it comes to recruiting interested faculty.

To put this in perhaps more tangible terms for the code4lib crowd: starting this type of program can lead to more cool projects. I think the code4lib phenomenon is about the library becoming that much more of a creative organization, and starting a digital initiatives program will move it further in that direction.

Thursday, March 6, 2008

code4lib presentation on Flickr/ceramics database

This is the code4lib lightning talk that Jeremy McWilliams and I did.

I'm including the first slide, which advertises our pretty well attended trail run up on Wildwood trail.

We're hoping that this project can demonstrate some of the advantages of using a digital asset management system "in the cloud," one that is part of a greater, participatory network.

Tuesday, March 4, 2008

the next generation library catalog phenomenon

When Casey Bisson gave an update of the Scriblio project as a lightning talk at code4lib last week, he made one comment that struck me: "Do we need another OPAC?" He was referring to the bevy of open source projects out there that support some kind of next generation search interfaces for library catalogs: eXtensible catalog, VUFind, FacBac, Scriblio, etc.

I recalled the first code4lib in 2006 when Casey's presentation on WPOPAC was one of the hot topics of the conference. NC State had just come out with their new Endeca based catalog then and people were pretty fired up about this idea of applying modern search features like faceting and relevance ranking to library catalog data.

Since then, many vendors and open source coders have jumped in to this area to make their play. Last year's code4lib conference also featured SOLR prominently and its potential role in the next generation OPAC.

I spent much of my time over the past couple years as Summit Catalog Committee Chair arguing for the need to move the Summit catalog over to a next generation platform. (Still hasn't happened.)

It's funny what a couple years can do. This area that was so cutting edge two years ago now seems overcrowded and almost passé.

Wednesday, February 27, 2008

broke breakout

I proposed a breakout at code4lib today but I didn't really get enough takers for it to fly. So here are some notes on it, DOA:

The topic was to be on "cloud computing and network level services" a discussion of the Nick Carr Big Switch thesis and its application to library environments; in addition, consideration of what library applications, services, and databases should be provided as network level services.

Cloud computing questions:
  • Is Nick Carr's thesis about utility style computing applicable to the library software world?
  • Do some of our open source projects (eg LibraryFind, Vufind, Scriblio, eXtensible catalog, Evergreen, Koha) miss out on the network effects of a more centralized model of data/services provision?
  • Should open source projects be approached differently from a cloud computing perspective?
  • Who in the library world is positioned to provide utility-level services? OCLC? Talis? Internet Archive? What should they provide?
  • Are there two visions of cloud computing out there, one more commercial another more open?
  • How does the cloud computing model intersect with the semantic web?
  • Are people using utility style computing as they build applications (Talis platform, OCLC web services, bibliocommons, Amazon S3, etc.)

Examples of network level services floating around the conference:
  • OCLC grid services
  • Talis platform
  • LibraryThing
  • OpenLibrary
  • BiblioCommons
  • Zotero 2.0

code4lib 2008 trends

Some observations regarding the first day of code4lib 2008:

There are a couple big players here from outside the library and library vendor worlds that are doing big, important library-related things: The Internet Archive and the Center for History and New Media.

The Internet Archive considers itself a library, a basically altruistic, nonprofit institution. But unlike any library I know, it has ambitions that are as far reaching as Google. It wants to create a digital archive of as much of the human record as it can get its hands on. I gotta say that I'm impressed with what they've done so far with the Internet Archive, and I support their other efforts with the OpenLibrary.

The Center for History and New Media comes at the perspective of research in the digital age from that of the humanities scholar. They've brought us Zotero, which I was impressed to learn has a user base of over 500K already. I'm also eager to here about their digital collections/digital exhibit software Omeka, which I see as a possible replacement for ContentDM at our shop.

It's really evident that there are a lot of competing solutions out there for solving various library technology problems. Throughout the day we heard of several solutions for catalog search/metasearch: the WorldCat API, VUFind, eXtensible catalog, and tangentially, LibraryFind. I can only wonder if there'll be a shakeout here sometime in the near future.

Thursday, March 8, 2007

data stores

At code4lib, Talis was promoting their platform. It's based on the concept of "stores", which are basically large bodies of data stored on Talis' computers. The advantage of putting your data in these stores is that it can be queried, searched, and related to other data in numerous ways.

Some of what they say about their platform:
Large-Scale Content Stores

The Talis Platform is designed to smoothly handle enormous datasets, with its multiple content stores providing a zero-setup, multi-tenant content and metadata storage facility, capable of storing and querying across numerous large datasets. Internally, the technology behind these content stores is referred to as Bigfoot, and there is an early white paper on this technology here.

Content Orchestration

The Talis Platform also comprises a centrally provided orchestration service which enables serendipitous discovery and presentation of content related according to arbitrary metadata. This service makes it easy to combine data from across different Content Stores flexibly, quickly and efficiently.


This all seems rather nebulous when you first think about it, but slowly, the usefulness of the concept begins to reveal itself. They discussed a little bit about how this platform is supporting Interlibrary Loan at UK libraries because it provides a way to query across different libraries.

My question is, do libraries really have enough of their own content to leverage a platform like this? All we really have is generic data about books and journals and specific data about what libraries holds them.

I wonder whether this kind of service would most useful if a player like Google offered it. Why Google and not Talis? Because they have huge amount of data already amassed from web crawling, publisher relationships, not to mention scanning books. Think about the opportunities that would present themselves if you could query specific slices of Google's content alongside your organization's own data? What if Google hosted research databases as stores and you could slice them up, query them, and relate them ala the Talis platform?

Essentially, a library could create its own, highly tailored searching/browsing/D2D systems.

Maybe I'm asking for too much.

Friday, March 2, 2007

standing on the sholders of giants

Casey Durfee's presentation on "Open Source Endeca in 250 lines or less" was pretty cool. How could he create a "next-gen" faceted catalog with such little code...by relying on Solr and Django to do the heavy lifting. Because Solr indexes XML natively, no relational database is even necessary. One of the things, generally speaking, I'm looking for at this conference is ways that we can leave the complexity to other applications.

Thursday, March 1, 2007

proximity and the network

Dan Chudnov gave a talk on making library resources available for sharing like itunes does on a LAN. It was hard to immediately sense the value in this. He spoke of walking into a library and having access to the whole of the library. Isn't that what we get through our digital presence on the web?

But thinking about it more, I like the idea of our computers being able to sense services and resources based on proximity. What if you met you met a group to study and when on the same wireless network, had immediate access to others' personal digital library on an application like Zotero or the like. What if when you walked through a physical library, the web presence of the library changed based on the section of the building you're in. Suppose you're studying in the East European Language Reading room late at night and you notice that somebody else has a similarly esoteric set of references on Polish intellectuals in their shared digital library...and perhaps that's her across the room. Could be a good way to get dates.

why code4lib?

Despite the fact I'm kind of burnt out on writing code, I find code4lib to be one of the most invigorating conferences I've attended in the last few years. Why? I think it's because it's where the new opportunities in the broader web world meet the digital library world.

Some interesting ideas that have come up this year:
  • the SOLR platform for indexing and faceting a library catalog or a digital library of anything really, XML based
  • The Talis platform's concept of data "stores": large bodies of xml data that can be queried and related to data in other stores in an unlimited number of ways using "web scale" infrastructure
  • the idea of hooking up openurl resolver type services as a microformat
  • using del.icio.us as a content management system for library subject guides
  • a subject recommendation engine based on crawling intellectual data associated with university departments
  • using a protocol like zeroconf so that library patrons can auto-discover library services upon entering the physical library space
It seems like most of the big players here work in larger universities or organizations that have large local data sets to work with in the form of institutional repositories or digital collections. There's a lot of concern about building large, searchable digital libraries . This is fine if you have control over a large body of data. I can tell you that in the small college library environment, most of the data we work with is generic data about books and journal articles that is living in some database that is out of our control. We're often only able to add value to that data once it's arrived in a user's search results, through an OpenURL resolver or perhaps a tweak to our catalog.

This is not to say that the what the big players are doing isn't useful or interesting to us. It's just different and makes me wish we had more opportunities to creatively manipulate the digital content to which we provide our patrons access.

bib-app

A team at my old place of employment, Wendt Library at UW-Madison, showed off a pretty cool application, bibapp, that gathers data about what faculty on their campus have published. Among other things, the data is used to find articles that our legally storable in the IR; almost cooler than that is the connections that they demonstrate between the publications. They can visualize who's publishing with who, analyze popular research subjects across discplines, etc.

Wednesday, February 28, 2007

Karen Schneider Keynote

The keynote had a few good points to it. She spent a lot of time talking about better ways to market open source projects. This seems like pretty obvious stuff.

One thing I liked was the concept of "rebuilding library artisans". The idea is that developers are the artisans of libraries. They build the systems that deliver library services. She argued that every library should have a developer--that this should be a given like a reference librarian, catalog librarian.

I tend to agree with this line of thinking. Libraries need developers to specialize their services to their local clietele. This is where they add value. At a place like L&C, we're really trying to put together a "rich" liberal arts learning environment. It's the micro-brew of higher ed (at least that's how we price it), so you really need the artisan to brew it up.

Another comment I kind of liked was related to library directors going to conferences and coming back determined to create a "learning commons." Funny how administrators attach so much importance to moving walls and furniture around when the revolution is happening online.

I got the feeling from this talk, and generally from this conference, that there's a lot of momentum building around the Evergreen ILS. The buzz is just starting over in III-land on the West Coast.