Date: Sun, 19 Mar 2000 14:23:56 -0500
From: Walter Scott <[EMAIL PROTECTED]>
Subject: Re: adobe and DOS? a tad off topic
Karen:
Generally, the E-Mail conversion services may not do well in
respect to formatted documents -- especially those with extensive
tables in them. After learning of XPDF here, I downloaded and tested
the software on the PDFs provided by the FBI on crime statistics. The
tables in such documents were preserved in their original format. That
was good -- although some might not think so because the original
format called for a table of information that could stretch as much as
255 columns from left to right margin. But, that's what you'd have to
deal with in the Acrobat reader as much as you'd have to deal with it
in a text editor after the PDF had been converted to text. Based on
what I've seen thusfar -- admittedly, not the full course of .PDF
files one might encounter -- I'm recommending this software over
E-Mail conversion. The only qualifiers are the size of your hard disk
and your CPU. You'll need lots of disk space to convert .PDF files
because these documents are ordinarily quite huge. XPDF itself is also
relatively large for DOS software.
The DOS version will run on any 80386 CPU or better with a DPMI
driver installed. But, it'll run REAL slow on older and slower
equipment. Attempt to convert a large PDF on an old 386 machine
running at less than 33 MHZ, and you'll probably need some patience.
If you want to try the DOS software, you can FTP it from
ftp.foolabs.com/pub/xpdf/xpdf-0.90-dos.zip and extract it with
PKUNZIP. Bear in mind, though, that files inside the archive are
maintained in UNIX format. Thus, PKUNZIP prior to version 2.50 will
not successfully extract files individually where such files don't
adhere to the DOS 8x3 standard for filenames. Just unzip the entire
archive into a temporary directory. The program you'll need is
PDFTOTEX.EXE. (It's PDFTOTEXT.EXE in the archive) Running the software
is simple. If the file you wish to convert is F1040EZ.PDF, you'd type
"PDFTOTEX F1040EZ.PDF" at the DOS command line. After PDFTOTEX is done,
you'll have a converted F1040EZ.TXT file.
After downloading several forms from the IRS website, and after
running them through PDFTOTEX.EXE, I found that the raw text was
correctly translated to ASCII text. Even the 1040 tax tables were
preserved in their original format. *HOWEVER*, and this is a big
HOWEVER, the raw text does not a valid tax form make. There are boxes
drawn in a normal form where you write or type in information. Since
PDFTOTEX.EXE only extracts pure text and a rather limited set of 8-bit
ASCII graphics characters, you only get the stark text of a tax form
without any of the boxes, lines that separate columns of text from one
another, and etc. XPDF provides an image extractor in the archive noted
above. I tried to find images to extract in, for example, the 1040EZ
form. There was none. So, it appears as though the IRS uses special
non-text fonts for drawing boxes and separator lines. You'll only be
able to generate a valid form with the boxes, lines, and other
necessary non-text information, under an Acrobat or Acrobat-compatible
reader.
One test I did not attempt to conduct was to see what might
happen if I used the utility (in the archive noted above) to convert
a PDF file to a Postscript file. That might preserve some of the
special features of an IRS tax form that are missing in the pure text
version. I have no way to verify that possibility, though.
Walter Scott
Quoting Karen Lewellen <[EMAIL PROTECTED]> to Walter Scott
f<[EMAIL PROTECTED]> and the Net-Tamer List Group on
18-Mar-00 at 19:48 PST:
KL> has anyone had success having government pdf files translated via the
KL> e-mail service?
********************************************************
To unsubscribe from this list,
send a message to [EMAIL PROTECTED] with the single word
Unsubscribe
as the subject.
You MUST use the same address with which you subscribed!
********************************************************