Some of you may be aware of Google's project called Google Print (and the 
Google Print Library Project).  I don't think it's an exaggeration to say 
that their intention is to digitize ALL of the world's print literature and 
make it searchable on the web (see the beta version at print.google.com), at 
least as far as possible.  Copyright aside and given that the world's 
published literature is probably under 100 million titles (my personal guess 
based on the number of records--over 50 million in OCLC WorldCat, the 
repository of library shared cataloging records), including the very 
obscure, it may well now be technically feasible to do so! 

Scanning is one thing, viewing the search results (at least large sections 
of books) is another matter, of course, because of copyright restrictions.  
Google has partnered with the libraries at Oxford, Michigan, Stanford, NY 
Public and Harvard in this endeavor.  It seems that Michigan is the only one 
that is allowing scanning of all of its titles (with full display access 
limited to on-campus users along a promise of litigation from some 
publishers).  The otherlibraries are reported to have limited the scanning 
to those titles that are in the public domain (basically pre-1923). 

A report I read recently stated that the resulting "library" will be 10.5 
million titles, with only a little over half in English, and somewhere 
around 20% being in the public domain (I don't have the article in front of 
me, and I don't remember if these figures reflect the complete holdings of 
these libraries or only those parts that they will allow to be digitized at 
this time).  It should be noted that some publishers are participating 
voluntarily as they do with Amazon. 

After a huge outcry from publishers over Google doing this without specific 
permission from them, there has been a pause in the project along with a 
promise that publishers can "opt out" if they want to.  Whether this will 
suffice remains to be seen.  Of course, those publishers that are 
participating voluntarily are no doubt considering this free advertising 
that will result in more book sales.  Presumably some of them have had a 
successful experience with this in Amazon. 

One of the problems that the project has uncovered is what to do about the 
millions of out of print titles that are not in the public domain--in many 
cases, the copyright holder can't even be located.  I would also note that 
the web (Alibris, Amazon, etc., even eBay) have also completely changed the 
way used books are sold, as well as their availability, not that this would 
help the publishers or copyright holders. 

It's not clear to me what they're doing with periodicals, although I noticed 
Science came up in one of my searches (probably a voluntary participant).  I 
also note that in order to view some hits, one must sign in using a Google 
account.  The site also notes that in the future they will be looking for 
additional titles in special collections and other libraries that have 
titles not in the collections of the participating libraries (a considerable 
number, I would think, particularly in non-English, older titles, and titles 
in specialized areas not represented in the original 5).  In addition, Yahoo 
has announced a rival plan (using the collections of the University of 
California--not clear to me if only Berkeley or more), as has a group in 
Europe. 

I'd like to hear some comments about the implications for orchidists.  It's 
still very early in the project, but perhaps not too early to evaluate its 
potential usefulness.  One thing that I noted in my own searching is that 
often we orchidists are only looking for a relatively small piece of data 
rather than to read a whole book.  As a result, many orchid books are 
reference in nature (lots of pieces of data) rather than a narrative 
monograph as in language or history.  If this is the case, simply having the 
tremendous indexing and search capabilities of Google, plus a rich database 
of scanned titles might be of considerable interest AND use precisely 
because we won't want to read more than a page or two at a time. 

I'll note in passing that the now-defunct AOS digital orchid project also 
highlighted yet again that there is MUCH beautiful and useful information in 
books in the public domain, most of which are long out of print and may even 
be rare.  It only takes one good copy to create a digital image, though.  It 
is my hope that these will quickly find their way into projects such as 
this. 

I did a little search on Paphiopedilum venustum and found 22 pages in 9 
books, along with one hit for Paphiopedilum venustus (an Oxford University 
Press book--typo or scanning error?).  In addition I found additional 
citations to "paph. venustum" and "P. venustum," demonstrating yet again the 
nomenclatural problems of searching botanical names, not to mention 
taxonomic ones. 

One thing that I think the AOS, RHS and others could do to make the project 
more useful to orchidists would be to identify all of the titles that they'd 
like to see digitized (beginning certainly with the older titles) and 
assisting Google in getting them into the database.  Another area that would 
also be useful, I would think, is to see if there are partners willing to 
enhance the metadata (so that retrievals are more consistent) and even to 
begin translating projects for those items in the public domain and making 
them available on the web or through this project. 


Comments are welcome. 

Sincerely,
Harvey Brenneise 

BTW, is there any news on that "AOS Bulletin" project?  I've been out of the 
loop for awhile, but I recall that about 4 years ago we were promised a 
product in about a year. 




_______________________________________________
the OrchidGuide Digest (OGD)
[email protected]
http://orchidguide.com/mailman/listinfo/orchids_orchidguide.com

Reply via email to