One of the topics that seemed to come up with some regularity at NAACL was that of discovering relations between concepts in text. For example, when reading a sentence like
The soldier and his whole platoon fought very bravely. It might be good to know that soldier and platoon are related by the "is a member of" relation (a soldier is a member of a platoon). Or, it might also be good to be able to acquire concept hierarchies automatically from text, that tell us things like "a dog is a kind of canine, etc". This can be very useful in domain specific contexts when existing resources like WordNet might be incomplete. This becomes an exercise in finding relations, because you essentially identify clusters of related words, and then join them up with IS-A relations. Of course you could do the same with other kinds of relations (like has part, etc.) and end up with an automatically created semantic network. So after hearing so much about these issues, I decided to do a little bit of background research, and this is what I have learned so far... There are various ways to categorize work on finding relations, and probably the easiest way to do that is with respect to the kind of relations that are found. The largest amount of work seems to be surrounding the discovery of IS-A relations (a dog is a canine), and that works takes various forms. In short, there have been "pattern based" approaches, that look for certain sequences of words (usually noun phrases) that might indicate an IS-A relation (like "Noun Phrase, Noun Phrase, and Noun Phrase are examples of Noun Phrase". Then there are approaches that are more based on co-occurrence data and distributional information. There has also been some work on discovering part-whole relations, and then work on discovering more general sorts of relations. So, below I try and summarize a little of what I know about each of these areas. ======================================== The earliest pattern based work on IS-A relations seems to be the following: Marti Hearst, 1992. Automatic acquisition of hyponyms from large text corpora. Proceedings of the 14th International Conference on Computational Linguistics (COLING-92). http://www.sims.berkeley.edu/~hearst/papers/coling92.pdf There have been other "pattern based" approaches since then, two recent examples in the context of Question Answering systems include: Fleischman, M. and Hovy, E. and Echihabi, A. 2003. Offline strategies for online question answering: Answering questions before they are asked. In Proceedings of ACL-03. http://acl.ldc.upenn.edu/acl2003/main/pdfs/Fleischman.pdf Mann, G.S. 2002. Fine Grained Proper Noun Ontologies for Question Answering. SemaNet'02: Building and Using Semantic Networks. http://www.cs.ust.hk/~hltc/semanet02/pdf/mann.pdf ================================================================= There has also been some work on discovering part-whole relations (aka meronyms) using patterns and other techniques. A current approach that seems to do quite well is the following: Girju, R. and Badulescu, A. and Moldovan, D. 2003. Learning semantic constraints for the automatic discovery of part-whole relations. In Proceedings of HLT/NAACL-2003. http://acl.ldc.upenn.edu/N/N03/N03-1011.pdf An earlier example of work that discovers part-whole relations is : Berland, M and Charniak, E. 1999. Finding parts in very large corpora. Proceedings of ACL-99. http://acl.ldc.upenn.edu/P/P99/P99-1008.pdf The Berland work seems to use a pattern based approach, while the Girju uses patterns in combination with machine learning techniques. ================================================================== More general work in discovering relations can be found here, where a number of different kinds of relations are discovered in text: Turney, P.D., and Littman, M.L. (2003), Learning Analogies and Semantic Relations, National Research Council, Institute for Information Technology, Technical Report ERB-1103. http://cogprints.ecs.soton.ac.uk/archive/00003084/01/NRC-46488.pdf ================================================================== Finally, there is quite a bit of work in discovering sets of related words or clusters, but that doesn't really get into the issue of determining the relations between words or sets of words. Good recent examples of this style of work might be: Riloff, E. and Shepherd, J., (1999) "A Corpus-Based Bootstrapping Algorithm for Semi-Automated Semantic Lexicon Construction", Journal of Natural Language Engineering , 1999, Vol. 5, No. 2, pp. 147-156. http://www.cs.utah.edu/~riloff/psfiles/jnle-semlex.pdf Patrick Pantel and Dekang Lin. 2002. Discovering Word Senses from Text. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD-02). pp. 613-619. Edmonton, Canada. http://www.isi.edu/~pantel/Download/Papers/kdd02.pdf ================================================================== Again, the work that finds clusters of related words is clearly connected to the relation finding work, in that it seems very useful to be able to identify various clusters of related words, and then join those clusters up according to whatever relation might truly connect them. I hope this makes some sense. More on all of these issues to follow! Ted -- Ted Pedersen http://www.d.umn.edu/~tpederse ------------------------ Yahoo! Groups Sponsor ---------------------~--> Yahoo! Domains - Claim yours for only $14.70 http://us.click.yahoo.com/Z1wmxD/DREIAA/yQLSAA/x3XolB/TM ---------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

