On Thursday Sept 16 we discussed LSA, based largely on our experiences with the demos supported at http://lsa.colorado.edu.
You can see the slides that Jason, Anagha, Apurva, and Prath prepared to describe their experiments here: http://www.d.umn.edu/~tpederse/Group04/group-04f.html We had a rather interesting and far reaching discussion, but here are a few of the interesting points. We spent a good bit of time thinking about a "car bonnet" (see Anagha's slides). This is a term that we rarely use in the USA - it refers to the hood of a car apparently. I suspect that the college reading text that was used to create the semantic space doesn't have many (if any) occurrences of "car bonnest", yet it has a rather long vector (meaning that some information is surely available). This suggested to use that both "car" and "bonnet" must have occurred a reasonably large number of times, and were combined (averaged?) to create a vector for "car bonnet". Now "bonnet" was most likely used in a clothing sense when it occurs without car, so we have a case where the lsa reported term might really be representing something quite different than we expected. In general we observed that compounds like "car bonnet" or "mr. edward rochester" might in fact give us such surprises, where we are thinking of them as a compound that refers to one particular concept, whereas LSA is viewing them as separate terms with different underlying senses. For example, in the case of "Edward Rochester" there was shown to be little or no similarity with "Jane Eyre". Of course, students of literature know that this is just wrong, they are quite related in the novel. But, what's happening is that "Edward Rochester" is being represented by the combination of "Edward" and "Rochester", both reasonably common names, whereas "Edward Rochester" is a well known character in the story of "Jane Eyre." These concerns extend to sentences, where again individual words are being combined, and we might be thinking of concepts that extend across multiple words. For example, "Lake Superior" (see Apurva's examples) clearly refers to our favorite lake, but you take the two terms separately, it may not be nearly that specific. The lake is full of water. The food was superior to any other. Apurva had a number of interesting examples of coherence in sentences, which led us to see how LSA might actually work reasonably well for some aspects of essay scoring. We also saw that LSA is a competing technique to both Prath and Jason's thesis topics. Jason has started working on the problem of measuring text similarity, and showed that LSA could do a reasonably good job in recognizing newspaper text that was similar in content (but not at all in actual word choices). Prath showed that LSA can do a good job of finding sets of related words. So now we must answer the question, why aren't we doing LSA? :) Well, those are just a few of the issues that arose. It was quite interesting, and there are lots of followup questions that we can pursue. We'll have a meeting on Wednesday Sept 22 at 5:30 pm with Peter Wiemer - Hastings. We'll be talking mostly about our own projects, and will do our best to relate and compare those to LSA. Ted -- Ted Pedersen http://www.d.umn.edu/~tpederse ------------------------ Yahoo! Groups Sponsor --------------------~--> $9.95 domain names from Yahoo!. Register anything. http://us.click.yahoo.com/J8kdrA/y20IAA/yQLSAA/x3XolB/TM --------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

