At NAACL (and most other conferences) most of the program is made up
of "regular" presentations,  which are based on papers accepted at the
conference that are on very specific technical issues. However, there are
also a few special "invited" presentations that try and capture the bigger
picture, and talk about  issues that are related but perhaps a bit less
typical than the normal fare at the conference. They are meant to be a
little bit speculative and motivating I think, in order to encourage folks
to look into new areas and problems.

There were two invited speakers at NAACL this year, Andrei Broder and Jill
Burstein. This summary is about the first of those (Andrei Broder), as his
talk was actually the first event of NAACL. So my postings have taken us
through the day of tutorials, and finally into Day 1 of the conference. At
this rate my report will be finished sometime in October. :)

Andrei Broder is a researcher at IBM Watson Research Lab, and he focuses
on the web.  Strangely enough, for a person who does research on the web,
Andrei doesn't have a home page, or at least doesn't have one that can
easily be found. Perhaps that should tell us something? :) But, here's a
generic page from IBM about their web research:

http://www.research.ibm.com/compsci/spotlight/web/

In any case, he started by reminding us of a bit of web history, and in
particular how web searching has evolved.

In June 1994, one of the early search engines was known as webcrawler, and
it indexed about 4000 servers, and 200k pages. In December 1995, Alta
Vista indexed about 30 million pages. Then in 1998 along came Google, and
now in 2004 they are talking in terms of an index of 3 billion pages. The
general point is that the web has developed very very quickly indeed, and
it's very hard to tell how large  it really is/isn't. He made the point
that it's very easy now for a single  user to generate lots and lots of
web pages automatically, and he said that he had a web server that ran on
a laptop that was capable of  generating 1 billion pages! (Of
course these aren't indexed, so do they really exist? :)

Some  figures he had from 2002 (I think) showed that there are about 250
million domain names, and 50 million hosts. Of course the size of the web
is hard to estimate, but he suggested that it might be around 3-10 billion
pages now, and is doubling every 8-12 months. There are some studies on
the size of the web at:

http://www.netcraft.com/survey/archive.html

and of course good general info about the state of search engines is
available at:

http://searchenginewatch.com/

He then went on to talk about how people now use the web, and he pointed
to a few studies that have been done by the Pew Foundation that seem
quite interesting (they do lots of studies, but he focused mostly on
reports 80 and 106 from the following site).

http://www.pewinternet.org/

direct links...

http://www.pewinternet.org/reports/toc.asp?Report=106
http://www.pewinternet.org/reports/toc.asp?Report=80

I won't try and summarize these, they simply provide some insight into
how people are using these internet these days - who they are and what
they are wanting to do. Broder has done some of his own studies and
classifies web use into several categories (from a 2002 SIGIR paper
"A Taxonomy of Web Search" http://www.acm.org/sigir/forum/F2002/broder.pdf)

1) informational - the traditional one, where you are trying to find some
piece of information on the web

2) navigational - where you know a web page exists, and you just want to
get there. best example is to type in someones name to a search engine,
knowing that it will point you there. This makes up a large percentage
(maybe 30%) of how people use the web, and has led to much less use of
bookmarks.

3) transactional - buy, sell, download, etc.

4) misc other things...

His own conclusions about web users are that...

1) make short ill-defined queries that reflect little effort.

2) look only at the first result and don't scroll down to see other
results.

3) are good at following links.

4) use various search engines (only 56% of web users say they use a single
search engine).

He discussed the state of web search, and divided things into three
distinct generations of technology.

The first generation of web searching technology is what we found in
1994-1998 roughly (pre-Google days), and it was much like classical IR.
You treat each page as a document, and then use standard text based
techniques like stemming, term counting, tf/idf, etc.

In the second generation of web search (as exemplified by Google) the use
of "off the page" content became popular. In other words, you use link
information about a page to characterize a page, as well as the actual
content of the page. For example, Google uses "anchor text" to search for
pages. Anchor text is that summarizing text that people put to identify a
page that they are linking to...for example...

Click here to see a page about George Washington, the First President of
the USA and Hero of the American Revolution
: http://www.whitehouse.gov/history/presidents/gw1.html

So, not only would the content of the whitehouse page be used by Google to
identify that it's about George Washington, so would my little blurb which
essentially acts as a summary of what the page is about (at least from my
perspective).

Right now the second generation represents the state of the art, and he
posed the third generation as being the day when we integrate searching
and some text analysis and natural language processing. This might include
reformating results appropriately, recognizing what the user "really"
wants, and generally doing more to help the user. He summarized the
mission of the third generation of web search as :

"Meet diverse user needs given their poorly made queries and the size and
heterogeneity of the web."

He stressed how bad user queries are at various times during his talk.
Usually they are 2-3 words, typically misspelled. :) He gave several
examples of simple services that could be provided to users that would
really help them (and that aren't currently done). One of them was name
disambiguation. In other words, if someone searches for "John Smith" help
them recognize how many distinct people who have that name they are
finding. This general problem is referred to as Name Disambiguation, and
is very much related to Word Sense Disambiguation/Discrimination, so
needless to say I find it quite interesting. I find it personally
interesting and important, and there are at least two distinct Ted
Pedersens who appear on the first page of Google search results for "Ted
Pedersen"!!!

He also talked a bit about the economics of the web, which was quite
illuminating. He said that Google gets about 97% of their revenue from
advertising! In particular, Google has gotten good at this idea of
"contextual advertising". A person can create an ad that contains certain
keywords, and when a user types those keywords they see an ad displayed.
If a user clicks on that ad, then the placer of the ad must pay. So
obviously you want to choose your keywords carefully, because if folks who
aren't really interested in being customers start clicking on your ad you
can be out of a lot of money.

For example, it appears that the "other" Ted Pedersen or his publisher has
placed an ad with Google. Whenever you search for "Ted Pedersen" in
Google, you get a helpful little ad that sends you to booksellers. The
other Ted Pedersen writes books, so this makes sense. Or if you search
for a term like "machine translation" you get a variety of ads displayed,
all of which relate to either human or automatic translation. And if you
search for "word sense disambiguation" you get an ad from Google telling
you to apply for work with them. :)

Anyway,  as long as folks don't click on your ad, you don't have to pay.
But if they click, then you pay for the ad.  This model is called the
"cost per click". It seems to me that this is a wonderful example of
where word  sense disambiguation is really really important, because you
don't want to bad ad keywords that are in some way ambiguous, and might
end up having lots of people flood your ads with clicks who are not
likely to be "real" customers.

All in all a most interesting talk! Nice way to start the conference I
thought.

Ted

--
Ted Pedersen
http://www.d.umn.edu/~tpederse


------------------------ Yahoo! Groups Sponsor ---------------------~-->
Make a clean sweep of pop-up ads. Yahoo! Companion Toolbar.
Now with Pop-Up Blocker. Get it for free!
http://us.click.yahoo.com/L5YrjA/eSIIAA/yQLSAA/x3XolB/TM
---------------------------------------------------------------------~->

 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
     http://groups.yahoo.com/group/nlpatumd/

<*> To unsubscribe from this group, send an email to:
     [EMAIL PROTECTED]

<*> Your use of Yahoo! Groups is subject to:
     http://docs.yahoo.com/info/terms/
 

Reply via email to