Russ Altman gave another of the invited talks at AAAI that I attended.

http://helix-web.stanford.edu/people/altman/

He works in the general area of Bioinformatics, but seems to have a
genuine interest in AI. In fact, he mentioned that his first conference
paper ever was at IJCAI in 1985, co-authored with Bruce Buchanan, who in
turn had worked with Edward Feigenbaum on Dendral back in the 1960's, so
you see how the world is a very small place indeed.

There was an interesting link between his talk and that of Ed
Feigenbaum's. Both seem to feel we spend too much time on inference. :)
Altman thinks that "data structures get a raw deal compared to
algorithms". He also invoked the slogan "knowledge is power", and in fact
cited Dendral as an example of a system that uses simple algorithms and
complex knowledge.

His feeling is that Bioinformatics has been dominated by complex
algorithms (inference) that deal with simple data. His belief is that we
need more knowledge, and we need to avoid purely statistical methods.

Bioinformatics is a vast area, but he suggested that some of the main
problems right now are:

1) sequence alignment
2) structure prediction
3) mico-array analysis

To explain all of these in any detail is well beyond my limited
understanding of biology, etc. You can find a good introduction to
bioinformatics here if you are interested:

http://www.macdevcenter.com/pub/a/mac/2004/06/11/bioinformatics.html

But, in *very* general terms,  bioinformatics is trying to take
computational and statistical techniques to manage and understand large
amounts of medical information. Altman described the problem generally as
having lots and lots of data that you must somehow sift through in order
to store and manipulate the knowledge therein, where knowledge takes the
form of relationships betweeen measured entities.

He then talked about a number of projects at his lab:

PharmGkB ( http://www.pharmgkb.org/ ) is a repository of information about
drugs and reactions to drugs. Here he said one of the main issues was that
they need data structures that capture the relationships between entities.

RiboWEB ( http://smi-web.stanford.edu/projects/helix/riboweb.html ) is an
ontology based system for supporting structural biological studies of the
ribosome. It is based on information extracted from 171 scientific
articles, which contained 8,000 experimental data items.

He also talked about a project where text analysis is used to try and
label genes as found in the Gene Ontology. http://www.geneontology.org/

There are a number of papers that seem to deal with this and related
issues at:

http://smi-web.stanford.edu/people/sxr/publications.html

But the one he specifically mentioned was:

Raychaudhuri S, Chang JT, Sutphin PD, and Altman RB Associating genes with
Gene Ontology codes using a maximum entropy analysis of biomedical
literature. Genome Research 12:203-214, 2002.
http://smi-web.stanford.edu/people/sxr/pdfs/NLP-GO.pdf

In looking at these publications, I couldn't help but noticing this one
too...

Raychaudhuri S, Schutze H, and Altman RB Using text analysis to identify
functionally coherent gene groups. Genome Research 10:1582-1590, 2002.
http://smi-web.stanford.edu/people/sxr/pdfs/geneGroups.pdf

This is yet another example of the small world phenomena again, because
of course the second author is our good friend (intellectually at least)
Hinrich Schutze, a main inspiration behind our SenseClusters package.

Then he moved on to talk about some of the infracstructure available in
the Bioinformations community.

1) Terminologies - these are lists of terms as used in the medical
profession - there are many terminology lists available - it would be nice
if they used the same terms for the same concepts, but they don't...

2) Ontologies - these organize concepts into structured relationships
(often is-a) - something like WordNet for medical terms. There are lots
and lots of different ontologies available, and again you have the same
problem of different ways of organizing the same types of information
between ontologies.

He then introduced UMLS and Mesh, both products of the National Library of
Medicine - UMLS attemps to resolve different terminologies in a common
way, by creating an ontology of sorts. Mesh organizes about 20,000
concepts into a hierarchy. I know these are very sketchy descriptions, so
if you want to know more than you ever wanted to know about UMLS and
Mesh, you can go here:

http://www.nlm.nih.gov/research/umls/
http://www.nlm.nih.gov/mesh/

He also talked a bit about the previously mentioned Gene Ontology.

The message I took away from this talk was that there is lots of
information available in Bioinformatics, and that getting a handle on it
and organizing it more effectively is the key to progress. There is a
problem both with the volume of data available, and the diversity of that
data.

Ted

--
Ted Pedersen
http://www.d.umn.edu/~tpederse


------------------------ Yahoo! Groups Sponsor --------------------~--> 
$9.95 domain names from Yahoo!. Register anything.
http://us.click.yahoo.com/J8kdrA/y20IAA/yQLSAA/x3XolB/TM
--------------------------------------------------------------------~-> 

 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
    http://groups.yahoo.com/group/nlpatumd/

<*> To unsubscribe from this group, send an email to:
    [EMAIL PROTECTED]

<*> Your use of Yahoo! Groups is subject to:
    http://docs.yahoo.com/info/terms/
 

Reply via email to