Hi Andrew,

Apologies for sidetracking your thread. I went off on a tangent because in fact 
text is not quite so easily extracted from a pdf as it might first appear, and 
if you cant extract the text then the question of the content format is 
irrelevant. For example, you find that you may get page headers and footers 
appearing in the middle of the text. With Pubmed documents you may even get 
some text from their logo in the middle of a sentence. So it is necessary to 
consider first how you extract the text, and that leads me back to file 
formats, sorry. I am no apologist for Microsoft but the point about the docx 
format is that it is plain old xml, as open a format as you can get. I think it 
should be applauded and thereby encouraged, whatever their motives.

Moving swiftly on to your actual question :) My experience to date with trying 
to extract metadata automatically has not been good. Partly due to the problems 
of extracting from pdf but also because of the variety of content in our 
documents. That having been said, I think publishers often dictate a format, 
right down to how references should appear, and authors will jump through hoops 
when preparing documents for their preferred publishers. So its is not beyond 
the bounds of possibility that some industry standards could be developed.

Cheers, Robin.   

  

Robin Taylor
Main Library
University of Edinburgh
Tel. 0131 6515208  

> -----Original Message-----
> From: Andrew Marlow [mailto:[email protected]] 
> Sent: 15 December 2008 21:41
> To: [email protected]
> Subject: Re: [Dspace-tech] standards to facilitate metadata 
> extractionduringtext extraction
> 
> On Mon, Dec 15, 2008 at 9:14 PM, Mark H. Wood <[email protected]> wrote:
> 
> 
>       Most common formats other than plain text have some 
> sort of tagging
>       feature.  In some cases, few know about them so they aren't much
>       used.  That could be fixed easily.
>       
> 
>       > Microsoft docx documents looks like a step in the 
> right direction 
> 
> 
>       The older Office formats are readable programmatically too.  
> 
> 
>       But then that only works for MS Office documents.  Not 
> for OpenOffice
>       or Symphony.  Not for Acrobat.  We have tens of 
> thousands of PDFs.  
> 
> 
> As the OP I would like to chip in again. I think people have 
> missed the point of what I was trying to say. I probably 
> wasn't very clear. What I meant was conventions for the 
> placement/position of title, authors, IISN, abstract, etc so 
> it could be easily extracted. I assumption was that the 
> document is a PDF. Text can easily be extracted from PDFs but 
> extraction on its own it not enough. We need to be able to 
> determine what the values are for title, authors etc. If 
> these were layed out in a std way it would help. That's all 
> I'm saying. This was not supposed to be a discussion about 
> file formats!
> 
> 
> --
> Regards,
> 
> Andrew M.
> 
>


-- 
The University of Edinburgh is a charitable body, registered in
Scotland, with registration number SC005336.


------------------------------------------------------------------------------
SF.Net email is Sponsored by MIX09, March 18-20, 2009 in Las Vegas, Nevada.
The future of the web can't happen without you.  Join us at MIX09 to help
pave the way to the Next Web now. Learn more and register at
http://ad.doubleclick.net/clk;208669438;13503038;i?http://2009.visitmix.com/
_______________________________________________
DSpace-tech mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/dspace-tech

Reply via email to