The second invited speaker at NAACL was Jill Burstein, of Educational Testing Services (ETS), and I'll try and do her talk justice here. ETS is probably well known to all of us, as they are the good folks who provide tests like the GRE, SAT, etc. etc.
Jill's talk focused on some tools that ETS is developing to automatically score student essays, both in the context of classroom instruction and then in standardized testing. In particular, she talked about E-Rater and Criterion mostly, and you can see more about those (as well as get related papers, etc.) at: http://www.ets.org/research/erater.html Now, as to these essays that are being graded, Jill stressed that this is not literary writing, but rather the standard 5 paragraph essay that students in the United States tend to see quite a lot of by the time they finish high school and even into college. These essays have 5 paragraphs, and start with a thesis statement, continue with the main points and supporting evidence for those points, and then a conclusion. All in 5 paragraphs. Generally speaking when students write these they are given a topic on which to write, they are being timed, and typically they end up writing essays about 500 words long. ETS grades lots of these kinds of essays when used in tests, and normally they employ two human graders who assign scores of 1 to 6 to each essay. 6 is great, 1 is horrible. Each grader spends at most 2 minutes on an essay, and if the two graders disagree on the score then there is a third grader who comes into the picture to resolve these differences. (By third grader I don't mean elementary school student, but rather a third expert who looks over the papers and resolves the score difference. :) In any case, there would be some advantage to having a computer serve as at least one of the graders, in order to speed up the grading process, and also possibly make grading a bit more reliable and consistent. E-Rater is the ETS tool that assigns scores to essays and can be used in testing situations, whereas Criterion provides feedback to the writer and suggests how to improve their essay. Jill put it rather succinctly, E-Rater gives a score, while Criterion gives feedback. E-Rater has been used since 1999 as one of two judges in the GMAT test, which is taken by about 350,000 students each year. So this is a an example of a real live deployed NLP application. Jill went into a little bit of the history of automatic essay evaluation, which goes back quite a ways. An early method was Project Essay Grade (Page 1996), aka PEG. In browsing around the web I found the following site, which includes a demo, which doesn't seem to completely work, but it gives you a feel for what this is all about http://134.68.49.185/PEGDEMO/. She mentioned two other systems, the Writer's Workbench, which seems to be available here : http://www.emo.com/ and the Intelligent Essay Assessor (IEA), which is available from: http://www.knowledge-technologies.com/ It's interesting to note that IEA is based on Latent Semantic Analysis, which is one of the techniques that is supported in Amruta's SenseClusters package. (http://senseclusters.sourceforge.net). Now, back to E-Rater and Criterion. One of the interesting issues that was discussed was that of "anomalous essay evaluation". How can a program pick up on an essay that is horribly bad, or non-responsive. For example, what if a student simply repeats the question they are supposed to address in their essay, and then mindlessly repeats. Or as Jill put it "annoying repetition". For example, suppose the question to address in the essay is: "Do you think high school students should have a curfew?" And suppose this is the essay: Do I think high school students should have a curfew. As a high school student I think curfews are an intersting idea, and merit serious consideration. High school students might need a curfew, but let us consider what that might imply. A curfew for high school students is intriguing, but might affect family life too much. To conclude, I think the question of whether high school students need curfews is very important to answer carefully. Clearly this is a very horrible essay, in that it is completely vacant and devoid of any intelligence whatsover (I am proud to say that I wrote it by myself. :) In any case, despite the fact that the above is fairly well formed grammatically, it's not a very good essay, and that is the kind of thing that E-Rater wants to be able to identify. They also do work that prevents system "gaming". She mentioned that a colleague Tom Morton has designed a "word salad detector" to identify when a student is just randomly entering words or gibberish. My essay above isn't really gibberish, it is intelligible English text that says absolutely nothing. The technical details of how E-Rater and Criterion work are a cornucopia of modern nlp technology, and use lots of the ideas that have been developed, like pos tagging, parsing, etc. For example (and this is an interesting note for those of you who are fans of WordNet::Similarity), Leacock and Chodorow (authors of the lch measure in our package) work at ETS, and they have developed a method of grammar checking that is based on bigrams. In particular, they look for unexpected sequences of part of speech tags, to detect problems like "these pencil". I think this is a good example of the kinds of techniques that they use in these tools - individual techniques that are fairly standard stuff, but put together in a way that does something very unique, interesting, and dare I say it, important. Anyway, I guess I've gone on long enough. This was a really interesting talk though. We all have written these kinds of essays, and the use of these tools seems like an interesting way to improve classroom instruction (allowing students to get more frequent feedback) and also lower the cost and hopefully improve the consistency of testing situations. There are lots of publications on the ETS site mentioned above that describe how E-Rater and Criterion work, and I also found the following online article, which goes into some more detail about essay evaluation, and describes PEG, the Intelligent Essay Assessor, and E-Rater. Rudner, Lawrence & Phill Gagne (2001). An overview of three approaches to scoring written essays by computer. Practical Assessment, Research & Evaluation, 7(26). Retrieved May 25, 2004 from http://PAREonline.net/getvn.asp?v=7&n=26 . This seems like a pretty decent source, and it provides some references to related work. Ted -- Ted Pedersen http://www.d.umn.edu/~tpederse ------------------------ Yahoo! Groups Sponsor --------------------~--> Yahoo! Domains - Claim yours for only $14.70 http://us.click.yahoo.com/Z1wmxD/DREIAA/yQLSAA/x3XolB/TM --------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

