Showing posts with label cloud computing. Show all posts
Showing posts with label cloud computing. Show all posts

Tuesday, October 28, 2008

NITLE cloud computing event report

Checking in from sunny Orlando Florida. I caught a ride from the Convention center zone to Rollins College in Winter Park Florida for the NITLE cloud computing event today. Overall, Rollins has a charming campus. The on- campus food wasn't bad either.

We heard a couple Google Apps migration stories from some CTOs. In one case, Wesleyan U, the school was only planning on switching students over, whereas at Macalester they had switched the whole enterprise, students and staff. Interestingly, in the Macalester case, the switch was done in a matter of days as the old email system failed. It seems that a crisis situation really served as an important catalyst and brought the community together. Now Macalester is ahead of the curve as it takes advantage of the whole Google Apps suite.

Jerry Sanders of Macalester said that with Google Apps, IT's role had become more "consultative" and "less reactive." It was now more about discussing the possibilities with these new web 2.0 applications than troubleshooting problems. He likened this shift and the renewed sense of unknown possibilities to the introduction of personal computers and the advent of the web.

We heard from the D-Space federation, who has some plans to enable D-Space to run on cloud-based storage. I continue to think that D-Space is not the right model for digital repositories in these times. It was designed before the Web 2.0 and Cloud Computing era and remains positioned as an isolated silo of data for supporting a single institution.

David Young, CEO of Joyent, gave his view of the cloud. Joyent provides infrastructure for some huge applications on the web. He disagrees with Carr's view that the cloud will be dominated by a few small companies and sees it as a more heterogenous beast. It was kind of fun to listen to an industry insider throw around jargon like "cloud stack" and "cloud primitives." A couple quotes:
cloud computing=aggressively distributed
a cloud should abstract away all consideration except the application and its operation
Young seemed a little concerned that some cloud providers where creating a situation where they would lock users into their platform...perhaps Amazon is trying to do this with its EC2 virtual machines. He said that Joyent's philosophy was "openness is lock in," akin to Southwest Airlines' flexibility in reservations. Using open application stacks like RoR keeps users loyal...plus the more data users put on your servers, the less likely they are to move (the dirty secret of cloud providers).

Finally, we heard form Lee Dirks of Microsoft's Education division. He said that MS sees academics as "extreme information workers." Microsoft has developed a few open source applications based on their Sharepoint platform that are designed to facilitate research, including software that can do conference planning and facilitate peer review. I was a little skeptical of some of these scholarly collaboration platforms--how far beyond more generic collaboration software do they take it? I'd have to have a closer look.

The day ended with some heated dialogue about information privacy and security concerns when using SaS providers. Many of the CTOs felt like it would be a big hurdle to get their campus legal counsel to agree to putting their data on external servers, but pretty much all agreed that this was the direction things are going.

Saturday, October 25, 2008

Economist article on cloud computing

The wife and baby are in Wisconsin this week visiting relatives while I head to EDUCAUSE in Orlando. There might be more blogging as a result.

The Economist has a piece this week on cloud computing. It's a pretty good overview of the concept for those who haven't been following it closely. Overall, however, I think it overemphasizes what I would call the raw, technical aspect of the phenomenon and under-emphasizes network-effects angle.

The idea of highly flexible computing power is a pretty cool one, and the piece cites an Amazon Web Services case study demonstrating just that. Using AWS, a Washington Post Engineer built a digital library of a massive collection of potentially newsworthy government documents about Hillary Clinton in nine hours. What a contrast to the timelines we're used to in libraries!

The most powerful aspect of the cloud computing phenomenon, in my opinion is the aggregation of data and the network effects that rise as systems get larger. The key feature of a cloud application is that it's data is part of a greater organic whole, and that it's able to do things that an isolated application can't. This is where the distinction between Web 2.0 and Cloud Computing gets fuzzy. The piece starts to touch on this concept when it brings up Tim O'Reilly
A raft of start-ups is also trying to build a business by observing its users, in effect turning them into human sensors. One is Wesabe (in which Mr O’Reilly has invested). At first sight it looks much like any personal-finance site that allows users to see their bank account and credit-card information in one place. But behind the scenes the service is also sifting through its members’ anonymised data to find patterns and to offer recommendations for future transactions based, for instance, on how much a particular customer regularly spends in a supermarket. Wireless devices, too, will increasingly become sensors that feed into the cloud and adapt to new information.
We now use Mint to track our home finances...for some reason we didn't like Wesabe. It knows how to categorize purchases on our credit card statement because it picks up on the ways other users categorize purchases with similar labels. Much nicer and easier than using Quicken used to be.

The piece brings up a concept of "industry operating systems" that will arise to allow businesses to become more modular and flexible, while relying more heavily on the services of others.
Both trends could mean that in future huge clouds—which might be called “industry operating systems”—will provide basic services for a particular sector, for instance finance or logistics. On top of these systems will sit many specialised and interconnected firms, just like applications on a computing platform.
This is interesting to contemplate. You could almost argue that Flickr fits this model. It provides the basic operating system and then so many other firms jump in and provide specialized services image service: prints, calendars, cards, etc. In this case the industry is totally virtual.

I liked this quote:
Twenty years ago, he argues, 80% of the knowledge that workers required to do their jobs resided within their company. Now it is only 20% because the world is changing ever faster.
There's a parallel here with libraries. We've seen a similar flip in terms of information residing in-house vs. outside. We're preparing students for the business world where information is also in the cloud.

I've been reading the Economist for 20 years now but I've come to realize that they are a bit technologically stodgy. Their online stories have no hyperlinks within them.

Tuesday, June 24, 2008

NITLE workshop on cloud computing

NITLE is hosting a workshop on cloud computing (focusing on the EC2 platform) and having a post-EDUCAUSE meeting on changes related to supporting enterprise applications
Server virtualization, software as a service, cloud computing, and open source software systems are all key technological and business factors that are dramatically changing how campuses select, deploy, and support enterprise software systems. This event will illuminate the fiscal and operational implications of these innovations for campus computing and library units. Featured presentations will include reports from campuses that have been investigating and exploiting these innovations.
It's great to see NITLE taking these issues head on. It'll be interesting to see if some colleges come up with interesting ways of using EC2. A researcher at a small college using powerful remote computers that their institution would never be able to provide is really what cloud computing is all about.

Also, its interesting that this is using a commercial entity for "cyberinfrastructure" rather than something designed specifically for academic research (though these projects could be more administrative than academic in nature).

Thursday, June 19, 2008

moving into to the cloud with Google Analytics

My experience with web statistics applications provides a good example of the move to cloud computing. Back in 2001, I recall painstakingly configuring our $900 copy of WebTrends desktop app and having to remember to download log files once a month. Then in the mid-2000s, we switched to the open source Webalizer, which conveniently is web-based and resided on our Linux server. Still, there was plenty of monkeying around with log files and cron jobs to get it to record the right data.

About a year ago, I got turned onto Google Analytics. It's powerful, super easy to configure and customize. It resides in the cloud. And its "free".


The usage pattern for the Watzek Library web site goes in pretty consistent waves, with troughs as the weekend approaches and and crests as the week starts.

Every hit on our website now gets registered on a Google server. So many sites are using Analytics now, it's amazing how much traffic and data Google is digesting. With "utility" services like Analytics, they are truly making themselves part of the basic infrastructure of the Internet. Their offer to host popular Javascript libraries fits into this as well.

Thursday, March 6, 2008

code4lib presentation on Flickr/ceramics database

This is the code4lib lightning talk that Jeremy McWilliams and I did.

I'm including the first slide, which advertises our pretty well attended trail run up on Wildwood trail.

We're hoping that this project can demonstrate some of the advantages of using a digital asset management system "in the cloud," one that is part of a greater, participatory network.

Wednesday, February 27, 2008

broke breakout

I proposed a breakout at code4lib today but I didn't really get enough takers for it to fly. So here are some notes on it, DOA:

The topic was to be on "cloud computing and network level services" a discussion of the Nick Carr Big Switch thesis and its application to library environments; in addition, consideration of what library applications, services, and databases should be provided as network level services.

Cloud computing questions:
  • Is Nick Carr's thesis about utility style computing applicable to the library software world?
  • Do some of our open source projects (eg LibraryFind, Vufind, Scriblio, eXtensible catalog, Evergreen, Koha) miss out on the network effects of a more centralized model of data/services provision?
  • Should open source projects be approached differently from a cloud computing perspective?
  • Who in the library world is positioned to provide utility-level services? OCLC? Talis? Internet Archive? What should they provide?
  • Are there two visions of cloud computing out there, one more commercial another more open?
  • How does the cloud computing model intersect with the semantic web?
  • Are people using utility style computing as they build applications (Talis platform, OCLC web services, bibliocommons, Amazon S3, etc.)

Examples of network level services floating around the conference:
  • OCLC grid services
  • Talis platform
  • LibraryThing
  • OpenLibrary
  • BiblioCommons
  • Zotero 2.0

Wednesday, January 30, 2008

Open Source ILS for Academic Libraries

This got forwarded to my email, from the CNI listserv:
The Duke University Libraries are preparing a proposal for the Mellon Foundation to convene the academic library community to design an open source Integrated Library System (ILS). We are not focused on developing an actual system at this stage, but rather blue-skying on the elements that academic libraries need in such a system and creating a blueprint. Right now, we are trying to spread the word about this project and find out if others are interested in the idea.

We feel that software companies have not designed Integrated Library Systems that meet the needs of academic libraries, and we don’t think those companies are likely to meet libraries’ needs in the future by making incremental changes to their products. Consequently, academic libraries are devoting significant time and resources to try to overcome the inadequacies of the expensive ILS products they have purchased. Frustrated with current systems, library users are abandoning the ILS and thereby giving up access to the high quality scholarly resources libraries make available.

Our project would define an ILS centered on meeting the needs of modern academic libraries and their users in a way that is open, flexible, and modifiable as needs change. The design document would provide a template to inform open source ILS development efforts, to guide future ILS implementations, and to influence current ILS vendor products.Our goal is not to create an open-source replica of current systems, but to rethink library workflows and the way we make library resources available to our
constitutiencies. We will build on the good work and lessons learned in other open source ILS projects. This grant would fund a series of planning meetings, with broad participation in some of those meetings and a smaller, core group of schools developing the actual design requirements document.
I agree that the current ILS marketplace doesn't deliver for academic libraries.

I'm not sure if a traditional open source project is the best solution, either. Seems to me that the next generation ILS should follow more of a cloud computing model instead of many disparate systems sharing a single base of code.

In my opinion, the question of a next generation ILS should be approached first from the data side, and then the software application side. As Tim O'Reilly puts it, Web 2.0 means "data is the next Intel Inside." The next generation ILS should be all about large pools of programmable, shared data. Organizations like OCLC and Serials Solutions have some of this data, but lots of other data dispersed across the net in various silos.

If libraries want to deliver a user experience at anywhere near the level of Google, we need to be using the same techniques that they are. And their most important technique is aggregating large amounts of data.

What am I talking about, more concretely? Our digital collections of unique materials should be managed in a centralized system that can leverage network effects in search, folksonomies, and more. Our library catalogs should simply be a subset of a larger shared catalog. Organization of licensed content should be facilitated by sharing metadata about that content. Even user/patron data should be managed in a network fashion using systems like OpenID.

In some ways, the next generation ILS is already emerging in the form of data driven products like Serials Solutions 360 and WorldCat Local. Of course, there is also more work to be done.

If a consortium of universities creates their own ILS, I'm afraid that it'll be a glacially moving monstrosity of a project like Fedora or DSpace. A theoretically wonderful piece of code that doesn't amount to much when it's installed in many isolated instances.