The second invited speaker at NAACL was Jill Burstein, of Educational
Testing Services (ETS), and I'll try and do her talk justice here. ETS is
probably well known to all of us, as they are the good folks who provide
tests like the GRE, SAT, etc. etc.

Jill's talk focused on some tools that ETS is developing to automatically
score student essays, both in the context of classroom instruction and
then in standardized testing. In particular, she talked about E-Rater and
Criterion mostly, and you can see more about those (as well as get related
papers, etc.) at:

http://www.ets.org/research/erater.html

Now, as to these essays that are being graded, Jill stressed that this is
not literary writing, but rather the standard 5 paragraph essay that
students in the United States tend to see quite a lot of by the time they
finish high school and even into college.

These essays have 5 paragraphs, and start with a thesis statement,
continue with the main points and supporting evidence for those points,
and then a conclusion. All in 5 paragraphs. Generally speaking when
students write these they are given a topic on which to write, they are
being timed, and typically they end up writing essays about 500 words
long.

ETS grades lots of these kinds of essays when used in tests, and normally
they employ two human graders who assign scores of 1 to 6 to each essay.
6 is great, 1 is horrible. Each grader spends at most 2 minutes on an
essay, and if the two graders disagree on the score then there is a third
grader who comes into the picture to resolve these differences. (By  third
grader I don't mean elementary school student, but rather a third expert
who looks over the papers and resolves the score difference. :)

In any case, there would be some advantage to having a computer serve as
at least one of the graders, in order to speed up the grading process,
and also possibly make grading a bit more reliable and consistent.

E-Rater is the ETS tool that assigns scores to essays and can be used in
testing situations, whereas Criterion provides feedback to the writer and
suggests how to improve their essay. Jill put it rather succinctly,
E-Rater gives a score, while Criterion gives feedback. E-Rater has
been used since 1999 as one of two judges in the GMAT test, which is
taken by about 350,000 students each year. So this is a an example of a
real live deployed NLP application.

Jill went into a little bit of the history of automatic essay evaluation,
which goes back quite a ways. An early method was Project Essay Grade
(Page 1996), aka PEG. In browsing around the web I found the following
site, which includes a demo, which doesn't seem to completely work, but it
gives you a feel for what this is all about http://134.68.49.185/PEGDEMO/.

She mentioned two other systems, the Writer's Workbench, which seems to
be available here : http://www.emo.com/
and the Intelligent Essay Assessor (IEA), which is available from:
http://www.knowledge-technologies.com/

It's interesting to note that IEA is based on Latent Semantic Analysis,
which is one of the techniques that is supported in Amruta's
SenseClusters package. (http://senseclusters.sourceforge.net).

Now, back to E-Rater and Criterion. One of the interesting issues that was
discussed was that of "anomalous essay evaluation". How can a program pick
up on an essay that is horribly bad, or non-responsive. For example, what
if a student simply repeats the question they are supposed to address in
their essay, and then mindlessly repeats. Or as Jill put it "annoying
repetition".

For example, suppose the question to address in the essay is: "Do you
think high school students should have a curfew?"

And suppose this is the essay:

Do I think high school students should have a curfew. As a high school
student I think curfews are an intersting idea, and merit serious
consideration. High school students might need a curfew,  but let us
consider what that might imply. A curfew for high school students is
intriguing, but might affect family life too much. To conclude, I think
the question of whether high school students need curfews is very
important to answer carefully.

Clearly this is a very horrible essay, in that it is completely vacant and
devoid of any intelligence whatsover (I am proud to say that I wrote it by
myself. :) In any case, despite the fact that the above is fairly well
formed grammatically, it's not a very good essay, and that is the kind of
thing that E-Rater wants to be able to identify.

They also do work that prevents system "gaming". She mentioned that a
colleague Tom Morton has designed a "word salad detector" to identify when
a student is just randomly entering words or gibberish. My essay above
isn't really gibberish, it is intelligible English text that says
absolutely nothing.

The technical details of how E-Rater and Criterion work are a cornucopia
of modern nlp technology, and use lots of the ideas that have been
developed, like pos tagging, parsing, etc.

For example (and this is an interesting note for those of you who are
fans of WordNet::Similarity), Leacock and Chodorow (authors of the lch
measure in our package) work at ETS, and they have developed a method of
grammar checking that is based on bigrams. In particular, they look for
unexpected sequences of part of speech tags, to detect problems like
"these pencil". I think this is a good example of the kinds of techniques
that they use in these tools - individual techniques that are fairly
standard stuff, but put together in a way that does something very
unique, interesting, and dare I say it, important.

Anyway, I guess I've gone on long enough. This was a really interesting
talk though. We all have written these kinds of essays, and the use of
these tools seems like an interesting way to improve classroom instruction
(allowing students to get more frequent feedback) and also lower the cost
and hopefully improve the consistency of testing situations.

There are lots of publications on the ETS site mentioned above that
describe how E-Rater and Criterion work, and I also found the following
online article, which goes into some more detail about essay  evaluation,
and describes PEG, the Intelligent Essay Assessor, and E-Rater.

Rudner, Lawrence & Phill Gagne (2001). An overview of three approaches to
scoring written essays by computer. Practical Assessment, Research &
Evaluation, 7(26). Retrieved May 25, 2004 from
http://PAREonline.net/getvn.asp?v=7&n=26 .

This seems like a pretty decent source, and it provides some references to
related work.

Ted

--
Ted Pedersen
http://www.d.umn.edu/~tpederse


------------------------ Yahoo! Groups Sponsor --------------------~--> 
Yahoo! Domains - Claim yours for only $14.70
http://us.click.yahoo.com/Z1wmxD/DREIAA/yQLSAA/x3XolB/TM
--------------------------------------------------------------------~-> 

 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
     http://groups.yahoo.com/group/nlpatumd/

<*> To unsubscribe from this group, send an email to:
     [EMAIL PROTECTED]

<*> Your use of Yahoo! Groups is subject to:
     http://docs.yahoo.com/info/terms/
 

Reply via email to