On approximately August 9, ITSS at UMD made the rather significant and 
far-reaching decision to not allow Google and other automated bots to  
index (or crawl) "personal" web pages. Those are pages that start with ~, 
like : http://www.d.umn.edu/~tpederse and most probably your own home  
page, if you are a UMD student or faculty. 

In effect, ITSS has decided that Google and other search engines don't 
need to know about our web content. Google and other search engines 
used to "crawl" our web pages and index the content with some regularity  
(at least once a week if the pages change) and that is how they know how  
to find your pages when someone does a query. Now, ITSS has said that 
they stopped allowing crawls due to privacy concerns, but, the web is by  
definition a public space so I don't understand that at all. 

I do not recall getting any notification of this, although I was in 
England at that time, so maybe I missed it. I only noticed this when my  
searches to google for my name and terms that I know I have high page 
ranks for (like word sense disambiguation, sense tagged text, etc.)  
started to not show my pages. If you search for my name now, you get 
virtually no UMD pages back, and that was not the case even a week ago. 

The problem is not only the fact that Google does not get info on current 
pages now, it sees that it is being blocked, and effectively is dropping 
all UMD content from personal pages from it's indexes. 

So, I mention this because I know those of you that are students put your 
resumes, papers, and software on your web pages, and I think you do that 
so that people can find it. That just got less likely, I'm afraid to 
report. I know faculty do that too - I put pretty much everything I do on 
my web pages, including papers, code, data, etc. and my fondest hope is 
that some fellow in Singapore or Argentina who doesn't have the slightest 
idea of who I am will search for something like "word sense  
discrimination" and find out about me and the work we've done here. As I 
say though, that just got less likely. 

What can be done? Well, right now about the best you can do is to move 
your content to a web site that is being indexed by Google. However, ITSS 
has indicated that they are working on an opt-in mechanism for people who 
want their pages indexed by Google and other crawlers and bots. However,  
I truly don't understand why they didn't have this in place before they 
stopped allowing crawling of UMD web pages, and until I see such a 
mechanism up and running I'm a little skeptical. I think in the end the 
right thing for them to do is to simply allow Google and other bots to 
index our pages. I will keep you posted as to how this develops.

BTW, you can (maybe) tell what a site is doing with respect to allowing 
bots to crawl by checking out the robots.txt file. This is a voluntary 
convention that is used to tell bots what they can access when they crawl.

http://www.d.umn.edu/robots.txt

You'll notice here that Disallow: /~ is being used, and that tells Google
and other crawlers not to index. Now, I say maybe becasue in fact UMD has 
gone one step beyond robots.txt and also (apparently) disabled things on 
the server side to keep our personal pages from being indexed. This is 
because robots.txt is a voluntary thing, and "bad" bots don't respect it. 
But, I think you can get a good idea of what your web sites "policy" with
respect to bots is by checking out robots.txt. For example...

http://www.cs.umn.edu/robots.txt
http://www.utah.edu/robots.txt

In the event you didn't know about all this, be aware that it has gotten   
harder for the average web surfer to find your UMD web pages, and any  
content that you have put there. Obviously is somebody knows your name  
and URL already, they will still find you (we are still connected to the 
internet, at least as of Aug 20 ;) But, people searching for content (not  
your name) will have a much harder time finding you, and that is a shame.  
I wish I had an easy way around this, but I'm afraid unless you have  
access to another web server you are pretty much stuck. 

Please let any colleagues who may have significant UMD web content know 
about this, so that they will understand what is happening with search 
engine results.  It took me a little while to figure this out, and I hope  
to save others the confusion this all caused me. 

Thanks,
Ted

--
Ted Pedersen
http://www.d.umn.edu/~tpederse



 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
    http://groups.yahoo.com/group/nlpatumd/

<*> To unsubscribe from this group, send an email to:
    [EMAIL PROTECTED]

<*> Your use of Yahoo! Groups is subject to:
    http://docs.yahoo.com/info/terms/
 


Reply via email to