There will be a CS department colloquia on Tuesday December 13 at 4pm, in Heller Hall 302. There will be two short (approx 15 minutes each, followed by questions) presentations, the first by Anagha Kulkarni and the second by Mahesh Joshi. These are intended to be practice talks for presentations they will make the following week at the 2nd Indian Conference on Artificial Intelligence (See http://www.iiconference.org/ for more information).
Please plan to attend! (If you are a CS graduate student, then you should attend unless you have a class or TA conflict). Below is information about both talks. ------------------------------------------------------------------------------ TITLE: ------ Name Discrimination and Email Clustering using Unsupervised Clustering and Labeling of Similar Contexts AUTHORS: -------- Anagha Kulkarni (with Ted Pedersen) ABSTRACT: --------- In this paper, we apply an unsupervised word sense discrimination technique based on clustering similar contexts (Purandare and Pedersen, 2004) to the problems of name discrimination and email clustering. Names of people, places, and organizations are not always unique. This can create a problem when we refer to or seek out information about such entities. When this occurs in written text, we show that we can cluster ambiguous names into unique groups by identifying which contexts are similar to each other. It has been previously shown by (Pedersen, Purandare, and Kulkarni, 2005) that this approach can be successfully used for discrimination of names with two-way ambiguity. Here we show that it can be extended to multiway distinctions as well. We adapt the cluster labeling technique introduced by (Kulkarni, 2005) for the multiway distinctions of name discrimination. On the similar lines of contextual similarity, we also observe that email messages can be treated as contexts, and that in clustering them together we are able to group them based on their underlying content rather than the occurrence of specific strings. ------------------------------------------------------------------------------ TITLE: ------ A Comparative Study of Support Vector Machines Applied to the Supervised Word Sense Disambiguation problem in the Medical Domain AUTHORS: -------- Mahesh Joshi (with Ted Pedersen and Rich Maclin) ABSTRACT: --------- We have applied five supervised learning approaches to word sense disambiguation in the medical domain. Our objective is to evaluate Support Vector Machines (SVMs) in comparison with other well known supervised learning algorithms including the naive Bayes classifier, C4.5 decision trees, decision lists and boosting approaches. Based on these results we introduce further refinements of these approaches. We have made use of unigram and bigram features selected using different frequency cut-off values and window sizes along with the statistical significance test of the log likelihood measure for bigrams. Our results show that overall, the best SVM model was most accurate in 27 of 60 cases, compared to 22, 14, 10 and 14 for the naive Bayes, C4.5 decision trees, decision list and boosting methods respectively. -- Ted Pedersen http://www.d.umn.edu/~tpederse ------------------------ Yahoo! Groups Sponsor --------------------~--> DonorsChoose.org helps at-risk students succeed. Fund a student project today! http://us.click.yahoo.com/9.ZgmA/FpQLAA/HwKMAA/x3XolB/TM --------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

