[
https://issues.apache.org/jira/browse/TIKA-4856?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18110874#comment-18110874
]
ASF GitHub Bot commented on TIKA-4856:
--------------------------------------
dschmidt commented on PR #3118:
URL: https://github.com/apache/tika/pull/3118#issuecomment-5516265156
Presets are the better mechanism, and I would drop my configuration layer
entirely: `ThumbnailDefaults`, the `thumbnail-defaults` config block and the
`renderThumbnails` switch all go, and a preset carries that instead. Resolving
the content worker-side at config trust, without `allowPerRequestConfig`, and
giving a proxy an addressable path segment are all things my version could not
do.
Your `thumbnail-unpack-selector` is the right home for the rest. The
selection rule is Tika knowledge rather than client knowledge, and as a
component it stops being a server endpoint question: `/unpack/preset/thumbnail`
returns the one image and its metadata, no new route, no base64 JSON. I am
happy to give up the endpoint for that.
One thing to settle before I build it. `UnpackSelector.select(Metadata)` is
a per-document boolean with no lookahead, but the rule is a ranking: a raster
THUMBNAIL directly below the document, else the rendering under a vector
THUMBNAIL, else a depth-1 RENDERING. Streaming, that becomes "the first
acceptable candidate wins", which differs from the ranking whenever a
lower-ranked candidate is emitted before a higher-ranked one, e.g. a rendered
PDF page ahead of a stored raster thumbnail. The two look mutually exclusive in
practice, since a PDF has no stored thumbnail, but that is an assumption about
emission order rather than something the interface guarantees. I can verify it
against the files I checked the endpoint with (doc, docx, ppt, pptx, xls, xlsx,
odt, epub, ggs, pages, numbers, key, mp3, m4a, flac, ogg, pdf, nef, pef), and
if it holds I would keep the selector streaming rather than widening the
interface.
One correction to the analysis in the table on #3096:
`standard-unpack-selector`'s `includeEmbeddedResourceTypes` is not an
equivalent of the selection, and my endpoint already uses it for what it does
do. It decides which embedded documents are unpacked at all; it cannot say
which of them is the document's thumbnail. Filtering to THUMBNAIL and RENDERING
still leaves the archive case (a DOCX in a ZIP contributes its own THUMBNAIL at
depth 2), the OLE2 case (the THUMBNAIL is a WMF and the usable image is the
rendering below it), and the OLE-object case (with rendering on, an embedded
object's picture is a RENDERING too). That is the part clients keep getting
wrong.
`maxRenderedPages` stays either way, and I will keep it in this PR.
> /unpack/thumbnail: return the document thumbnail with its metadata
> ------------------------------------------------------------------
>
> Key: TIKA-4856
> URL: https://issues.apache.org/jira/browse/TIKA-4856
> Project: Tika
> Issue Type: New Feature
> Reporter: Dominik Schmidt
> Priority: Major
>
> With TIKA-4850 through TIKA-4855 every container format that carries a
> thumbnail emits it as a THUMBNAIL embedded document, the PDF parser renders
> pages as RENDERING documents, and the EMF/WMF renderer turns the vector
> thumbnails of Office documents into raster ones. Getting "the thumbnail of
> this file" out of that still takes format knowledge on the client: the
> THUMBNAIL of a Word or Excel file is an EMF/WMF whose usable form is the
> RENDERING underneath it, a PDF has no THUMBNAIL but a page RENDERING, the
> THUMBNAIL of a DOCX inside a ZIP is not the ZIP's, and with rendering enabled
> the picture of an embedded OLE object is a RENDERING too. Plus the request
> config that switches the renderers on.
> Proposal: POST /unpack/thumbnail next to /unpack and /unpack/all, multipart
> like them. It runs the usual forked parse in unpack mode with a fixed parse
> context (PDF page 1 rendered, EMF/WMF rendered) and picks, in this order: the
> raster THUMBNAIL at depth 1; the rendering of that thumbnail; the depth-1
> RENDERING of PDF page 1. The endpoint extracts what the document carries; it
> does not resize, convert or generate previews.
> The response is JSON: the /rmeta metadata object of the selected embedded
> document, and the image as base64. Thumbnails are small, so the encoding
> overhead does not matter, and the caller gets type, dimensions, origin
> (stored thumbnail or rendering, tk:rendering:rendered-by) and path in one
> round trip without unpacking a zip. 204 when the document has no thumbnail.
> {
> "metadata": {
> "Content-Type": "image/png",
> "Content-Length": "8459",
> "tiff:ImageWidth": "800",
> "tiff:ImageLength": "1131",
> "tk:embedded-resource-type": "RENDERING",
> "tk:embedded-resource-path": "/thumbnail.emf/thumbnail.png",
> "tk:embedded-depth": "2",
> "tk:rendering:rendered-by": "poi-metafile-renderer",
> "tk:resource-name": "thumbnail.png"
> },
> "image": "iVBORw0KGgoAAAANSUhEUgAA..."
> }
> To keep the selection rule short, the metafile renderer could give the
> rendering of a THUMBNAIL the THUMBNAIL type as well (its
> tk:rendering:rendered-by tells it apart), so a raster thumbnail is a
> THUMBNAIL regardless of whether the document stored it as PNG or as EMF.
> What do you think?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)