The Economist has a leader supporting the Google Books Deal, and an interview with Paul Courant, Dean of Libraries at Univ. of Michigan.
He talks some about the product that Google will be offering to libraries with this deal.
I have to wonder if this product will be the watershed moment for e books in academic libraries. If Google's library of books is big and broad enough to serve as a general library on its own, Google's platform for e books could become the place to do research in books.
Much of its success will depend on how much current content is in their index, and this is really dependent on Google doing deals with thousands of publishers. If Google's index is largely made up of older scanned books, it'll be a useful research tool, but not compelling as place for general research.
Google might become the place to do research in books, whereas recreational e book reading will happen through other vendors like Amazon.
Showing posts with label Google Book Search. Show all posts
Showing posts with label Google Book Search. Show all posts
Tuesday, September 8, 2009
Tuesday, June 17, 2008
on Google Books, yet out of reach
A long time friend of mine, Patrick Michelson, recently completed his PhD thesis in Russian history at University of Wisconsin-Madison. News of this brings me back to my days as a grad student in history at UW in the mid 1990s.
More than the academics, graduate school was really about the distractions: frequenting the smoky bars of Madison, watching old movies, acquiring a taste for bourbon and country music, hunting for prized possessions at garage sales, sailing and windsurfing on Lake Mendota, taking backpacking and canoing trips, and...working at the library.
Back then, Patrick and I both worked at the Memorial Library Information Desk under the venerable Information Librarian and Building Manager Dennis Auburn Hill. "Working at the library" usually meant doing some reading and shooting the shit with other library employees and friends while holding down the desk. Indeed, it was probably the cushiest job in the library outside of the late night study id checker. In fact, at Memorial after 5 p.m. underemployed graduate students sort of took over the library by running the various service counters throughout the building.
When we both finally overcame the distractions and procrastinations that kept us from completing our MA theses back in the summer of 1996, we had but one way of paying tribute to the friendships and pastimes that gave meaning to our humble existence at the time: the acknowledgments page. This was an opportunity, using veiled references and inside humor, to interject a little personalization into a weighty academic document. We packed it to the gills with mention of mentors, friends, and family.
I was delighted to find that my MA thesis as well as Patrick's have been digitized by Google. My first instinct was to look at that acknowledgments page. I was really bummed to find that the document is restricted to snippets. I'm not sure who to blame? The University? Google? University Microfilms? Or me for not giving permission? The whole beauty of the Google Project is to get obscure and generally insignificant works like these out in the public domain.
I felt honored to make it into Patrick's acknowledgments again in is PhD thesis. One of the things I asked Patrick about his thesis was whether he got to put a hardbound copy in the library. Having your thesis in the basement of the library along with all the "giants" who preceded you always seemed like a big part of the reward. (We also used to take pride in looking ourselves up in OCLC WorldCat, was at least as exciting as finding your own name in the phonebook.)
He said that the library still was putting a paper copy in the stacks, though now its a tinier paperback version. For future dissertators, it'll probably eventually shrivel down into a digital copy .
More than the academics, graduate school was really about the distractions: frequenting the smoky bars of Madison, watching old movies, acquiring a taste for bourbon and country music, hunting for prized possessions at garage sales, sailing and windsurfing on Lake Mendota, taking backpacking and canoing trips, and...working at the library.
Back then, Patrick and I both worked at the Memorial Library Information Desk under the venerable Information Librarian and Building Manager Dennis Auburn Hill. "Working at the library" usually meant doing some reading and shooting the shit with other library employees and friends while holding down the desk. Indeed, it was probably the cushiest job in the library outside of the late night study id checker. In fact, at Memorial after 5 p.m. underemployed graduate students sort of took over the library by running the various service counters throughout the building.
When we both finally overcame the distractions and procrastinations that kept us from completing our MA theses back in the summer of 1996, we had but one way of paying tribute to the friendships and pastimes that gave meaning to our humble existence at the time: the acknowledgments page. This was an opportunity, using veiled references and inside humor, to interject a little personalization into a weighty academic document. We packed it to the gills with mention of mentors, friends, and family.
I was delighted to find that my MA thesis as well as Patrick's have been digitized by Google. My first instinct was to look at that acknowledgments page. I was really bummed to find that the document is restricted to snippets. I'm not sure who to blame? The University? Google? University Microfilms? Or me for not giving permission? The whole beauty of the Google Project is to get obscure and generally insignificant works like these out in the public domain.
I felt honored to make it into Patrick's acknowledgments again in is PhD thesis. One of the things I asked Patrick about his thesis was whether he got to put a hardbound copy in the library. Having your thesis in the basement of the library along with all the "giants" who preceded you always seemed like a big part of the reward. (We also used to take pride in looking ourselves up in OCLC WorldCat, was at least as exciting as finding your own name in the phonebook.)
He said that the library still was putting a paper copy in the stacks, though now its a tinier paperback version. For future dissertators, it'll probably eventually shrivel down into a digital copy .
Monday, May 5, 2008
how could Google help search in academic libraries?
John Wilkin has an interesting post about various ways Google Scholar could add functionality that would help academic library patrons get to the specialized databases provided by academic libraries. Interestingly, he brings Anurag Acharya, the guy who created Google Scholar, in on the discussion. The ideas generally have to do with learning about the user's needs and then pointing them to the more specialized resources. The post really addresses the problem of metasearch, that is, finding a way to give users a simple, single search box and get them from there to some of the richer, more powerful databases produced for academic research.
But what about once a library patron is in a research database like MLA Bibliography, Historical Abstracts, or Psychinfo? Many of these resources are fairly primitive when it comes to the search functionality and content that they cover. Often you get to search the citations, abstracts, sometimes the fulltext of academic articles. Sure, sometimes more is less, but typically, they don't cover the increasing amount of scholarly material that is out there on the open web. They also certainly don't offer the fulltext of books.
If Google (or another big search vendor) offered a platform that database vendors could mount their systems on, those vendors could make so much better products. Services available to the vendor could include:
I suspect that it wouldn't be worth it to Google to design a product for the library research sector. This would need to be an infrastructure product that could span proprietary search needs of multiple industries.
When we got a Search Appliance here at Lewis & Clark, I have to admit, I was kind of disappointed playing around with the admin interface, that you couldn't easily mix in parts of Google's web index with your own proprietary stuff. Guess this is sort of what I'm asking for here.
Scirus is sort of a development in this direction, that is a hybrid of the research database and search engine. Another sort-of-related idea: Dan Cohen has called for Google Books to open up its APIs for scholarly inquiry.
Some folks will no doubt be horrified that I'm suggesting putting more of our eggs in Google's basket. But the idea really is about bringing web scale infrastructure to the service of more specialized, niche needs. Not giving ourselves over to Google, but rather using their data and software as a platform on which to accomplish bigger things.
But what about once a library patron is in a research database like MLA Bibliography, Historical Abstracts, or Psychinfo? Many of these resources are fairly primitive when it comes to the search functionality and content that they cover. Often you get to search the citations, abstracts, sometimes the fulltext of academic articles. Sure, sometimes more is less, but typically, they don't cover the increasing amount of scholarly material that is out there on the open web. They also certainly don't offer the fulltext of books.
If Google (or another big search vendor) offered a platform that database vendors could mount their systems on, those vendors could make so much better products. Services available to the vendor could include:
- access to Google search software
- ability to create an continually updated index of portions of the web alongside proprietary data
- ability to provide advanced search functionality and data analysis specific to the needs of a particular discipline
- access to Google Books index
I suspect that it wouldn't be worth it to Google to design a product for the library research sector. This would need to be an infrastructure product that could span proprietary search needs of multiple industries.
When we got a Search Appliance here at Lewis & Clark, I have to admit, I was kind of disappointed playing around with the admin interface, that you couldn't easily mix in parts of Google's web index with your own proprietary stuff. Guess this is sort of what I'm asking for here.
Scirus is sort of a development in this direction, that is a hybrid of the research database and search engine. Another sort-of-related idea: Dan Cohen has called for Google Books to open up its APIs for scholarly inquiry.
Some folks will no doubt be horrified that I'm suggesting putting more of our eggs in Google's basket. But the idea really is about bringing web scale infrastructure to the service of more specialized, niche needs. Not giving ourselves over to Google, but rather using their data and software as a platform on which to accomplish bigger things.
Monday, April 14, 2008
OCLC's competitive advantage
In a previous post, I mentioned the "cloud computing" aspects of the WorldCat.org platform when used in the form of WorldCat Local or WorldCat Group to replace an ILS based catalog. A couple weeks ago, I got some more thoughts together on this, but some other distractions pulled me away.
Last year, as part of my work with the Orbis Cascade Alliance's Catalog Committee, I got to survey the market for next generation library catalogs/discovery systems. In my mind, OCLC's WorldCat Group/Local option stands out against both the open source and commercial competition. Why? Because of network effects.
Competitors making products like ExLibris's Primo and Innovative Interface's Encore simply don't have access to the data that OCLC does, and neither does the open source community. OCLC's holdings data lets them do relevance ranking by the number of libraries that own the item. Its global database allows the potential of expanding a search beyond a single library or group of libraries to a global database with built-in ILL. WorldCat offers records that are updated and improved over time by shared cataloging, and the possiblity of enrichment of those records by web-scale social networking (reviews, tags, etc.). WorldCat is a living organism that can't simply be replicated on someone else's server.
These advantages are byproducts of the cataloging and resource sharing networks that OCLC has had in place for years. They are more about a community committed to sharing resources than about technology.
It's also about exclusive access to data, which fits in nicely with Tim O'Reilly's Web 2.0 tenant, data is the next Intel Inside.
Another way to think about this is that OCLC is in a position to be a sort of EBay for libraries. EBay is valuable precisely because of its wide user base. It is the dominant player in online auctions because it offers the widest possible marketplace for buyers and sellers. That same comprehensiveness and global scope also has value when building a search system for books and doing resource sharing.
Traditional ILS vendors have little tradition of sharing data between their customers. I really can't imagine Innovative Interfaces putting together any kind of product that involves mixing customers data (that weren't part of a consortium doing direct business with them). The whole idea of mixing customer data seems like it would run counter to traditional notions of enterprise level systems, and would really be hard for a longstanding software company to grasp (though I have to point out that Talis is very much an exception). Most customers buying enterprise software wouldn't want to share their data with peers anyway, right?
But that is really a non-issue for libraries, and precisely what this Web 2.0 world calls for, and what OCLC has been doing for a long time. OCLC has this valuable data, and great potential to develop things with it, as well as a general current towards network-level computing moving in its favor. When libraries compare OCLC's products with that of a traditional ILS vendor, they need to see that the OCLC product is more than technology. Rather, it is an extension of a community, a network. The commercial vendors just can't offer that.
OCLC doesn't really have a monopoly on bibliographic data. But they definitely are one of the largest players out there, and their data could put them way out ahead. They are in a position to create an impressive platform with their WorldCat.org line of products. As they do this, they should build it so that it's open enough that other companies and organizations can build products on top of it.
The model I'm thinking of is Flickr. Flickr is a great global platform for photo sharing, but through their API, Yahoo also lets other firms get in there and add value to it. OCLC should be build WorldCat as an information ecosystem that allows the library community to have a healthy marketplace of technology products. (They sort of have this model going with ILL, but there aren't too many competitors out there to OCLC's ILL management product, ILLIAD.) OCLC's new APIs for WorldCat are a sign that they are moving in this direction.
OCLC's competitive advantage in holdings data and shared cataloging applies largely to print materials. They really don't have any such advantage in the realm of electronic information: e journals, digital collections, etc.. They might be going after this with acquisition of companies like Openly Informatics and the ContentDM software, but even so, they are not in possession of data that gives them a competitive advantage in the same way as their cataloging and resource sharing networks. Many companies like ExLibris and Serials Solutions have e-serials holdings data. And data about library digital collections is generally open for harvesting/crawling.
Even in print material, Google Books could be a viable competitor to OCLC in the library search arena, especially because they have the advantage of full text searching. They have lots of data that OCLC doesn't. It would even be possible now with the Google AJAX search API to create a Google Book Search that linked back to your library catalog for books held in your library.
Last year, as part of my work with the Orbis Cascade Alliance's Catalog Committee, I got to survey the market for next generation library catalogs/discovery systems. In my mind, OCLC's WorldCat Group/Local option stands out against both the open source and commercial competition. Why? Because of network effects.
Competitors making products like ExLibris's Primo and Innovative Interface's Encore simply don't have access to the data that OCLC does, and neither does the open source community. OCLC's holdings data lets them do relevance ranking by the number of libraries that own the item. Its global database allows the potential of expanding a search beyond a single library or group of libraries to a global database with built-in ILL. WorldCat offers records that are updated and improved over time by shared cataloging, and the possiblity of enrichment of those records by web-scale social networking (reviews, tags, etc.). WorldCat is a living organism that can't simply be replicated on someone else's server.
These advantages are byproducts of the cataloging and resource sharing networks that OCLC has had in place for years. They are more about a community committed to sharing resources than about technology.
It's also about exclusive access to data, which fits in nicely with Tim O'Reilly's Web 2.0 tenant, data is the next Intel Inside.
Another way to think about this is that OCLC is in a position to be a sort of EBay for libraries. EBay is valuable precisely because of its wide user base. It is the dominant player in online auctions because it offers the widest possible marketplace for buyers and sellers. That same comprehensiveness and global scope also has value when building a search system for books and doing resource sharing.
Traditional ILS vendors have little tradition of sharing data between their customers. I really can't imagine Innovative Interfaces putting together any kind of product that involves mixing customers data (that weren't part of a consortium doing direct business with them). The whole idea of mixing customer data seems like it would run counter to traditional notions of enterprise level systems, and would really be hard for a longstanding software company to grasp (though I have to point out that Talis is very much an exception). Most customers buying enterprise software wouldn't want to share their data with peers anyway, right?
But that is really a non-issue for libraries, and precisely what this Web 2.0 world calls for, and what OCLC has been doing for a long time. OCLC has this valuable data, and great potential to develop things with it, as well as a general current towards network-level computing moving in its favor. When libraries compare OCLC's products with that of a traditional ILS vendor, they need to see that the OCLC product is more than technology. Rather, it is an extension of a community, a network. The commercial vendors just can't offer that.
OCLC doesn't really have a monopoly on bibliographic data. But they definitely are one of the largest players out there, and their data could put them way out ahead. They are in a position to create an impressive platform with their WorldCat.org line of products. As they do this, they should build it so that it's open enough that other companies and organizations can build products on top of it.
The model I'm thinking of is Flickr. Flickr is a great global platform for photo sharing, but through their API, Yahoo also lets other firms get in there and add value to it. OCLC should be build WorldCat as an information ecosystem that allows the library community to have a healthy marketplace of technology products. (They sort of have this model going with ILL, but there aren't too many competitors out there to OCLC's ILL management product, ILLIAD.) OCLC's new APIs for WorldCat are a sign that they are moving in this direction.
OCLC's competitive advantage in holdings data and shared cataloging applies largely to print materials. They really don't have any such advantage in the realm of electronic information: e journals, digital collections, etc.. They might be going after this with acquisition of companies like Openly Informatics and the ContentDM software, but even so, they are not in possession of data that gives them a competitive advantage in the same way as their cataloging and resource sharing networks. Many companies like ExLibris and Serials Solutions have e-serials holdings data. And data about library digital collections is generally open for harvesting/crawling.
Even in print material, Google Books could be a viable competitor to OCLC in the library search arena, especially because they have the advantage of full text searching. They have lots of data that OCLC doesn't. It would even be possible now with the Google AJAX search API to create a Google Book Search that linked back to your library catalog for books held in your library.
Labels:
Google Book Search,
III,
Innovative Interfaces,
OCLC,
WorldCat Local
Monday, July 2, 2007
Google Book Search Local
So here's an idea...
The Google AJAX search API makes it pretty darn easy to create a customized Google Book Search for your own web page. The API will return an identifier (which is typically an ISBN) for the book, among other metadata in results lists. Why not use a little JSON and JQuery to figure out whether or not your library holds the item and insert that information w/ link in the results? A database of your library and/or union catalog holdings would facilitate the task.
This would create a nicely localized version of Book Search for your library.
The Google AJAX search API makes it pretty darn easy to create a customized Google Book Search for your own web page. The API will return an identifier (which is typically an ISBN) for the book, among other metadata in results lists. Why not use a little JSON and JQuery to figure out whether or not your library holds the item and insert that information w/ link in the results? A database of your library and/or union catalog holdings would facilitate the task.
This would create a nicely localized version of Book Search for your library.
Labels:
Google,
Google Book Search,
specialize,
WorldCat Local
Subscribe to:
Posts (Atom)
