Well, looks like hocr2pdf needs the boundingbox informations per
character where cuneiforms 1.0.0 output has changed to boundingboxes per
line. This is an allowed format as described by hocr specification. So
its not cuneiforms fault, but it would be nice to get a switch that
forces cuneiform to give boundingboxes per character instead of per
line. That would be a feature request so and no issue.

-- 
Font size not correct in merged sandvich PDF
https://bugs.launchpad.net/bugs/623438
You received this bug notification because you are a member of Cuneiform
Linux, which is the registrant for Cuneiform for Linux.

Status in Linux port of Cuneiform: Invalid

Bug description:
After processing with Cuneiform for Linux 1.0.0 and hOCR to PDF converter, 
version 0.7.4 (should be the most current version) I get a sandvich pdf that 
looks nice until I select text.

See the sample 5AADFEE1-0000.* files in the attachment and the result.pdf.
The effect is shown in screen087.png

For another file (Test10pages.pdf) the effect is either worse - basically I 
cannot really select any more text to copy because I only can guess where to 
move with the mouse.

It looks like that the font size in the HTML is somehow not correct - I am not 
an expert, but this link might help you:
http://www.emdpi.com/fontsize.html



_______________________________________________
Mailing list: https://launchpad.net/~cuneiform
Post to     : cuneiform@lists.launchpad.net
Unsubscribe : https://launchpad.net/~cuneiform
More help   : https://help.launchpad.net/ListHelp

Reply via email to