My favorite paper at ACL isn't always what I think is the best paper, but
it's simply the paper that I enjoy the most, or is the most novel, or
something that makes it stand out. This year that paper was found in the
student research workshop, and it was about clustering suras from the
Koran/Quran in order to understand thematic relationships.
In particular, here's the full reference:
@InProceedings{thabet:2005:Student,
author = {Thabet, Naglaa},
title = {Understanding the Thematic Structure of the {Q}ur'an: An
Exploratory Multivariate Approach},
booktitle = {Proceedings of the ACL Student Research Workshop},
month = {June},
year = {2005},
address = {Ann Arbor, Michigan},
publisher = {Association for Computational Linguistics},
pages = {7--12},
url = {http://www.aclweb.org/anthology/P/P05/P05-2002}
}
What I like about this work is that first of all, it's about the Koran.
That's a rather brave thing in some respects - so much of ACL is about the
Wall Street Journal and similarly dull content :), so something about the
Koran is strikingly different. If you search the ACL Anthology, you find
only about 3-4 mentions of the Koran/Quran, and one of them is a
misspelling of Koran Sparck-Jones. Going against fashion is always a
difficult thing, and that's why I say it is somewhat brave.
You can find the transliteration of the Quran used in this work here,
along with some other interesting material (index, search facility, etc.)
http://www.usc.edu/dept/MSA/quran/
In any case, here's what I learned - the Koran consists of 114 suras, each
of which is between 4 and 286 ayats/verses long. This work focuses on the
24 suras that have more than 1000 words. The suras have been processed in
the usual ways - that is to say that stop words are removed, words of
frequency 1 are removed, and then the words are stemmed. This leaves a
corpus of 3672 tokens.
Then, a matrix representation of 24 rows (suras) by 3642 columns (tokens)
is created. The values in the cells are the variances of the frequency of
each word in the suras (I'm not entirely clear on this point, but I get
the gist of it I think). Then they do a bit of dimensionality reduction
and cluster the results.
What is quite interesting about their results is that they get two
clusters, each of which corresponds to the place of revelation of the
Quran. Just a note, the suras of the Quran are categorized by scholars as
having been revealed at either Mecca or Medina, two holy places in the
Islamic world. Well, the clusters that were discovered corresponded with
these holy places. One cluster consisted of suras revealed at Medina, and
the other at Mecca. Isn't that neat? :)
In any case, my next plan is to try to cluster the Quran using SenseClusters.
So, that is my nominee for the most interesting paper at ACL! The data is
quite novel, the methods are solid, the results are intriguing, and it
caused me to learn a bit about the Quran!
Enjoy,
Ted
--
Ted Pedersen
http://www.d.umn.edu/~tpederse
Yahoo! Groups Links
<*> To visit your group on the web, go to:
http://groups.yahoo.com/group/nlpatumd/
<*> To unsubscribe from this group, send an email to:
[EMAIL PROTECTED]
<*> Your use of Yahoo! Groups is subject to:
http://docs.yahoo.com/info/terms/