Peter,

>>But the *updated* specification is dated October 25, 1996. I would think
>>that > 7.5 years would be long enough for them to act on it.
>
> Without knowing I suspect this was part of the backlog before the
> addition. Part of what I was involved in was seeing whether
> the conversion could be automated.

In my opinion, correcting existing entries should be given much higher
priority by the PDB.

It seems inexcusable to me that after 7.5 years they cannot even follow
their own spec. In 1996 there were only ~5K entries in the database.


>>But, from my perspective, I am already doing crazy mapping in order to
>>turn CA into alpha carbon.
>
> So long as the ***column*** is correct in a PDB file there isn't a
> problem.
> There are exactly 2 fields for the atom symbol
> |  C|A   is carbon
> |CA|   is calcium

This works for alpha carbon because they distinguish between the columns,
as you have shown.

It does not work for AC? in 1EBL because columns 13 & 14 (the critical
columns) contain AC ... presumably the name they have given to some carbon
atom ... which maps to another element.

> But I suspect your problem is that in the scripting language there are no
> columns and so CA and CA look pretty similar :-(. In CML I allow for this
> by having non-space delimiters in arrays, e.g.

That is not the issue in this case.

It is a separate problem that only affects the scripting language. The
scripting language can generally be solved by using the set 'alpha' or
saying 'carbon & *.CA'


>>As a newcomer on the scene, it is clear to me that the PDB has brought
>>most of these problems on themselves by not enforcing the file format and
>>by continuing to publish files that do not meet their own format
>>specification.
>
> Be gentle... It was a pioneering effort in the early 1970's when
> scientific databases were very rare.

I am sure that it is true that it was pioneering in its time.

As I said before, I believe that they have a responsibility to put more
resources into cleaning up existing entries.


> I think if it comes from PDB it represents "best endeavour with limited
> resources"
[snip]
>>And, I fear, they will continue to reply to my data-quality
>> bug reports by saying "No Plan To Fix".
>
> The PDB is fully aware of the problem. In fact the protein community has
> put a lot of effort into cleaning up protein structure files. So I
> wouldn't spend time reporting.

Again, I question the allocation of their resources.

> I guess the likely response would be "we are migrating to mmCIF which
> should remove many of these problems"

That is basically their response. They also made reference to the database
(which they are still building) that is going to (somehow) eliminate all
the problems associated with fixed file formats.

I believe that their thinking is flawed, and that they are doing a
disservice to their users. I believe that they need to distribute more of
their existing resources to cleaning up existing entries in their current
(relatively popular) file format.

1EBL.cif defines the element type for this atom to be 'C'. I can think of
no justifiable reason why this change (editorial or automated) wasn't
propagated back to the .pdb entry into the element symbol field in order
to comply with their own file format specification.


Miguel



-------------------------------------------------------
This SF.Net email is sponsored by: GNOME Foundation
Hackers Unite!  GUADEC: The world's #1 Open Source Desktop Event.
GNOME Users and Developers European Conference, 28-30th June in Norway
http://2004/guadec.org
_______________________________________________
Jmol-developers mailing list
[EMAIL PROTECTED]
https://lists.sourceforge.net/lists/listinfo/jmol-developers

Reply via email to