One of the topics that seemed to come up with some regularity at NAACL was
that of discovering relations between concepts in text. For example, when
reading a sentence like

The soldier and his whole platoon fought very bravely.

It might be good to know that soldier and platoon are related by the "is a
member of" relation (a soldier is a member of a platoon).

Or, it might also be good to be able to acquire concept  hierarchies
automatically from text, that tell us things like "a dog is a kind of
canine, etc". This can be very useful in domain specific contexts when
existing resources like WordNet might be incomplete. This becomes an
exercise in finding relations, because you essentially identify clusters
of related words, and then join them up with IS-A relations. Of course you
could do the same with other kinds of relations (like has part, etc.) and
end up with an automatically created semantic network.

So after hearing so much about these issues, I decided to do a little bit
of background research, and this is what I have learned so far...

There are various ways to categorize work on finding relations, and
probably the easiest way to do that is with respect to the kind of
relations that are found. The largest amount of work seems to be
surrounding the discovery of IS-A relations (a dog is a canine), and that
works takes various forms. In short, there have been "pattern based"
approaches, that look for certain sequences of words (usually noun
phrases) that might indicate an IS-A relation (like "Noun Phrase, Noun
Phrase, and Noun Phrase are examples of Noun Phrase". Then there are
approaches that are more based on co-occurrence data and distributional
information. There has also been some work on discovering part-whole
relations, and then work on discovering more general sorts of relations.

So, below I try and summarize a little of what I know about each of these
areas.

========================================

The earliest pattern based work on IS-A relations seems to be the following:

Marti Hearst, 1992. Automatic acquisition of hyponyms from large text
corpora. Proceedings of the 14th International Conference on Computational
Linguistics (COLING-92).
http://www.sims.berkeley.edu/~hearst/papers/coling92.pdf

There have been other "pattern based" approaches since then, two recent
examples in the context of Question Answering systems include:

Fleischman, M. and Hovy, E. and Echihabi, A. 2003. Offline strategies for
online question answering: Answering questions before they are asked. In
Proceedings of ACL-03.
http://acl.ldc.upenn.edu/acl2003/main/pdfs/Fleischman.pdf

Mann, G.S. 2002. Fine Grained Proper Noun Ontologies for Question
Answering. SemaNet'02: Building and Using Semantic Networks.
http://www.cs.ust.hk/~hltc/semanet02/pdf/mann.pdf

=================================================================

There has also been some work on discovering part-whole relations (aka
meronyms) using patterns and other techniques. A current approach that
seems to do quite well is the following:

Girju, R. and Badulescu, A. and Moldovan, D. 2003. Learning semantic
constraints for the automatic discovery of part-whole relations. In
Proceedings of HLT/NAACL-2003.
http://acl.ldc.upenn.edu/N/N03/N03-1011.pdf

An earlier example of work that discovers part-whole relations is :

Berland, M and Charniak, E. 1999. Finding parts in very large corpora.
Proceedings of ACL-99.
http://acl.ldc.upenn.edu/P/P99/P99-1008.pdf

The Berland work seems to use a pattern based approach, while the Girju
uses patterns in combination with machine learning techniques.

==================================================================

More general work in discovering relations can be found here, where a
number of different kinds of relations are discovered in text:

Turney, P.D., and Littman, M.L. (2003), Learning Analogies and Semantic
Relations, National Research Council, Institute for Information
Technology, Technical Report ERB-1103.
http://cogprints.ecs.soton.ac.uk/archive/00003084/01/NRC-46488.pdf

==================================================================

Finally, there is quite a bit of work in discovering sets of related words
or clusters, but that doesn't really get into the issue of determining the
relations between words or sets of words. Good recent examples of this
style of work might be:

Riloff, E. and Shepherd, J., (1999) "A Corpus-Based Bootstrapping
Algorithm for Semi-Automated Semantic Lexicon Construction", Journal of
Natural Language Engineering , 1999, Vol. 5, No. 2, pp. 147-156.
http://www.cs.utah.edu/~riloff/psfiles/jnle-semlex.pdf

Patrick Pantel and Dekang Lin. 2002. Discovering Word Senses from Text. In
Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data
Mining (KDD-02). pp. 613-619. Edmonton, Canada.
http://www.isi.edu/~pantel/Download/Papers/kdd02.pdf

==================================================================

Again, the work that finds clusters of related words is clearly connected
to the relation finding work, in that it seems very useful to be able to
identify various clusters of related words, and then join those clusters
up according to whatever relation might truly connect them.

I hope this makes some sense. More on all of these issues to follow!

Ted

--
Ted Pedersen
http://www.d.umn.edu/~tpederse


------------------------ Yahoo! Groups Sponsor ---------------------~-->
Yahoo! Domains - Claim yours for only $14.70
http://us.click.yahoo.com/Z1wmxD/DREIAA/yQLSAA/x3XolB/TM
---------------------------------------------------------------------~->

 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
     http://groups.yahoo.com/group/nlpatumd/

<*> To unsubscribe from this group, send an email to:
     [EMAIL PROTECTED]

<*> Your use of Yahoo! Groups is subject to:
     http://docs.yahoo.com/info/terms/
 

Reply via email to