Y'know, I was just wondering if deskewing was a good idea or if I was being overly persnickety. Of course, that should have happened *before* the OCR stuff I was doing. Really, it ought to have been the first thing, with the conversion to JBIG2 afterward.
Are you doing the deskewing automatically (unpaper) or by hand? I agree it is important to not have too many versions and to be clear about what is the primary source. I see my version as more of a test bed, something small enough that I can experiment with processing it, but always making sure anything I do could be applied easily by Brian to the master copy. Also, since Brian will be making archival quality scans of more Model T documents, I hope what I'm doing will result in a useful recipe for how to post-process archaic documentation in a way that is most useful to computer archaeologists like us. As you say, deskew is always lossy,¹ so I am not sure Brian's primary source should fold back in that change. However, it would be good if derivations like mine could easily apply your work. Do you think it'd be possible to make a deskew script that could be run on any PDF to apply the rotations you measure? --b9 ____ ¹ Always lossy, unless you do the deskew by editing the PDF by hand to add a matrix transform to each page so the original image is untouched in the file. On July 20, 2026 9:19:17 AM PDT, Joshua O'Keefe <[email protected]> wrote: >I've been experimenting with doing a bit of de-skew on B9's JBIG2-based "daily >driver" style copy, just to remove the distraction it creates. Clearly there's >no reasonable way to create an archival-adjacent copy while doing that — the >rotation is by definition lossy — however I thought that for something I'd >just keep open next to a project it'd be good to have something as close as >possible to a typeset document, with usability in mind, rather than an >archival version. > >If there's interest, I'm happy to put a little more time into it and share. As >always, keeping clear pointers to the source document to reduce the chance any >derived versions would be assumed authoritative original scans. > >
