Russ Altman gave another of the invited talks at AAAI that I attended. http://helix-web.stanford.edu/people/altman/
He works in the general area of Bioinformatics, but seems to have a genuine interest in AI. In fact, he mentioned that his first conference paper ever was at IJCAI in 1985, co-authored with Bruce Buchanan, who in turn had worked with Edward Feigenbaum on Dendral back in the 1960's, so you see how the world is a very small place indeed. There was an interesting link between his talk and that of Ed Feigenbaum's. Both seem to feel we spend too much time on inference. :) Altman thinks that "data structures get a raw deal compared to algorithms". He also invoked the slogan "knowledge is power", and in fact cited Dendral as an example of a system that uses simple algorithms and complex knowledge. His feeling is that Bioinformatics has been dominated by complex algorithms (inference) that deal with simple data. His belief is that we need more knowledge, and we need to avoid purely statistical methods. Bioinformatics is a vast area, but he suggested that some of the main problems right now are: 1) sequence alignment 2) structure prediction 3) mico-array analysis To explain all of these in any detail is well beyond my limited understanding of biology, etc. You can find a good introduction to bioinformatics here if you are interested: http://www.macdevcenter.com/pub/a/mac/2004/06/11/bioinformatics.html But, in *very* general terms, bioinformatics is trying to take computational and statistical techniques to manage and understand large amounts of medical information. Altman described the problem generally as having lots and lots of data that you must somehow sift through in order to store and manipulate the knowledge therein, where knowledge takes the form of relationships betweeen measured entities. He then talked about a number of projects at his lab: PharmGkB ( http://www.pharmgkb.org/ ) is a repository of information about drugs and reactions to drugs. Here he said one of the main issues was that they need data structures that capture the relationships between entities. RiboWEB ( http://smi-web.stanford.edu/projects/helix/riboweb.html ) is an ontology based system for supporting structural biological studies of the ribosome. It is based on information extracted from 171 scientific articles, which contained 8,000 experimental data items. He also talked about a project where text analysis is used to try and label genes as found in the Gene Ontology. http://www.geneontology.org/ There are a number of papers that seem to deal with this and related issues at: http://smi-web.stanford.edu/people/sxr/publications.html But the one he specifically mentioned was: Raychaudhuri S, Chang JT, Sutphin PD, and Altman RB Associating genes with Gene Ontology codes using a maximum entropy analysis of biomedical literature. Genome Research 12:203-214, 2002. http://smi-web.stanford.edu/people/sxr/pdfs/NLP-GO.pdf In looking at these publications, I couldn't help but noticing this one too... Raychaudhuri S, Schutze H, and Altman RB Using text analysis to identify functionally coherent gene groups. Genome Research 10:1582-1590, 2002. http://smi-web.stanford.edu/people/sxr/pdfs/geneGroups.pdf This is yet another example of the small world phenomena again, because of course the second author is our good friend (intellectually at least) Hinrich Schutze, a main inspiration behind our SenseClusters package. Then he moved on to talk about some of the infracstructure available in the Bioinformations community. 1) Terminologies - these are lists of terms as used in the medical profession - there are many terminology lists available - it would be nice if they used the same terms for the same concepts, but they don't... 2) Ontologies - these organize concepts into structured relationships (often is-a) - something like WordNet for medical terms. There are lots and lots of different ontologies available, and again you have the same problem of different ways of organizing the same types of information between ontologies. He then introduced UMLS and Mesh, both products of the National Library of Medicine - UMLS attemps to resolve different terminologies in a common way, by creating an ontology of sorts. Mesh organizes about 20,000 concepts into a hierarchy. I know these are very sketchy descriptions, so if you want to know more than you ever wanted to know about UMLS and Mesh, you can go here: http://www.nlm.nih.gov/research/umls/ http://www.nlm.nih.gov/mesh/ He also talked a bit about the previously mentioned Gene Ontology. The message I took away from this talk was that there is lots of information available in Bioinformatics, and that getting a handle on it and organizing it more effectively is the key to progress. There is a problem both with the volume of data available, and the diversity of that data. Ted -- Ted Pedersen http://www.d.umn.edu/~tpederse ------------------------ Yahoo! Groups Sponsor --------------------~--> $9.95 domain names from Yahoo!. Register anything. http://us.click.yahoo.com/J8kdrA/y20IAA/yQLSAA/x3XolB/TM --------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

