There is a relatively new yahoo group that is meant to be a place to
discuss "ACL information resources", including things like a possible
software repository or registry...

If you are interested in that discussion or participating, you can join
here :

http://tech.groups.yahoo.com/group/acl-information-services/

 ... believe it or not, some of you are among the leading experts
on releasing and supporting open source software in the NLP
community, so your contributions to this discussion would be valuable.

I posted the following today as a followup to some previous
discussion of a software repository ... some of you will recognize this as
very much inspired by CPAN - say what you will about Perl,
I think CPAN presents a great model for distributing code and
create a somewhat integrated community of users, authors, testers,
etc.

=================================================

http://tech.groups.yahoo.com/group/acl-information-services/message/107

I think the discussion of the merits of Wikis versus other
alternatives, as well as the possible role of the ACL Anthology is
quite interesting. I should say though that my overall observation of
the discussion so far is that it seems to focus a bit more on what is
going to be easier for the authors of code to deal with, and perhaps
not so much on what will actually help out the users of the code. In
the end what will help the authors most is if they have lots and lots
of happy users, and users will be helped if they have lots of nice
code to choose from, so thinking about both authors and users in how
this designed will surely help the whole enterprise.

I think simply allowing authors to provide a link to an existing code
page is nice, but probably will result in a ton of dead or missing
links within a year or less (assuming that anyone contributes). We
actually have had experiments like that (the NLP software registry
http://registry.dfki.de/ comes to mind) and they surely provide a very
useful service, but, on the other hand, it doesn't seem like that
model has solved the problem. So, I do think providing something like
a persistent archive where code and possibly data can be stored is
very worthwhile, assuming that it allows authors to update their code
as time goes by (ie it should be possible for an author to upload
improved versions of that code should they wish or should it be
necessary).

I also think there should be some thought given to the advantages of
being able to have a more user centered view of this, where user
downloads are tracked, and where users are able to comment on and even
rate the code, and have those comments and ratings available not only
for the authors but to the community as a whole (especially potential
users). We could want to encourage users to provide feedback, as in
"successful install on Ubuntu Linux system (Hardy), using kernel
2.6.9 SMP, ..." or provide tips like "I had to tweak the install
instructions like this to get this working on my Mac", or comments
like "I got rather different results when I ran this code", and so on
... this will help users avoid reinventing the wheel, and give a sense
of what code is working and successful, and which is not. It's a kind
of open-source model of peer-review in fact, and might turn out to be
a motivating factor in persuading authors that they want to make code
available.

Finally, I do think there should be some discussion (here or
elsewhere) of how to persuade authors that it's worth their while to
make software and data available. I think there are lots of reasons
for doing that (which was the point of the Last Word piece previously
mentioned), but I also think perhaps ACL could spur that along a bit
with some sorts of inducements. For example, to be eligible for a best
paper award at ACL, you must provide code and data that replicates
your results in your initial submission. That might be a bit much for
some folks tastes, but it's an example at least. I personally feel
like the ability to connect with users and get their feedback via a
mechanism as described above would be a powerful inducement, since
code that is useful and well done would quickly rise to the top of the
heap.

-- 
Ted Pedersen
http://www.d.umn.edu/~tpederse

Reply via email to