Y'know, I was just wondering if deskewing was a good idea or if I was being 
overly persnickety. Of course, that should have happened *before* the OCR stuff 
I was doing. Really, it ought to have been the first thing, with the conversion 
to JBIG2 afterward. 

Are you doing the deskewing automatically (unpaper) or by hand? 

I agree it is important to not have too many versions and to be clear about 
what is the primary source. 

I see my version as more of a test bed, something small enough that I can 
experiment with processing it, but always making sure anything I do could be 
applied easily by Brian to the master copy. Also, since Brian will be making 
archival quality scans of more Model T documents, I hope what I'm doing will 
result in a useful recipe for how to post-process archaic documentation in a 
way that is most useful to computer archaeologists like us.

As you say, deskew is always lossy,¹ so I am not sure Brian's primary source 
should fold back in that change. However, it would be good if derivations like 
mine could easily apply your work. Do you think it'd be possible to make a 
deskew script that could be run on any PDF to apply the rotations you measure? 

--b9


____
¹ Always lossy, unless you do the deskew by editing the PDF by hand to add a 
matrix transform to each page so the original image is untouched in the file. 


On July 20, 2026 9:19:17 AM PDT, Joshua O'Keefe <[email protected]> 
wrote:
>I've been experimenting with doing a bit of de-skew on B9's JBIG2-based "daily 
>driver" style copy, just to remove the distraction it creates. Clearly there's 
>no reasonable way to create an archival-adjacent copy while doing that — the 
>rotation is by definition lossy  — however I thought that for something I'd 
>just keep open next to a project it'd be good to have something as close as 
>possible to a typeset document, with usability in mind, rather than an 
>archival version.
>
>If there's interest, I'm happy to put a little more time into it and share. As 
>always, keeping clear pointers to the source document to reduce the chance any 
>derived versions would be assumed authoritative original scans.
>
>

Reply via email to