[
https://issues.apache.org/jira/browse/TIKA-4856?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18110776#comment-18110776
]
ASF GitHub Bot commented on TIKA-4856:
--------------------------------------
tballison commented on PR #3096:
URL: https://github.com/apache/tika/pull/3096#issuecomment-5513615469
Argh. I'm sorry. I thought I posted a review yesterday. Something went wrong.
I'm hesitant to add a special endpoint/parameter for thumbnails.
I asked :robot: to come up with a plan to get you what you want largely with
what exists and some of what you've added.
This is what it came up with:
```
What the PR adds vs. what exists
┌───────────────────────────────────────────────────────────┬──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ PR piece │
Existing equivalent
│
├───────────────────────────────────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ ThumbnailDefaults (built-in → server block → request, │
parse-context section in tika-config.json (baseline, loaded by the forked
worker from the same file) + the │
│ per-component merge) │ multipart
config part (runtime delta, overlaid in the worker)
│
├───────────────────────────────────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ ?renderThumbnails=true query param │ sending the
same parser blocks in the request's config part
│
├───────────────────────────────────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ /unpack/thumbnail + ThumbnailSelector │
standard-unpack-selector already has includeEmbeddedResourceTypes — filter to
THUMBNAIL/RENDERING on plain │
│ │ /unpack
│
├───────────────────────────────────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ thumbnail-defaults top-level key + TikaJsonConfig change │ not needed
│
├───────────────────────────────────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ PDFParserConfig.maxRenderedPages │ no
equivalent — this is the one genuinely missing piece (without it, rendering
only page 1 forces maxPages: 1, │
│ │ which cuts
the text too)
│
└───────────────────────────────────────────────────────────┴──────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
So my recommendation: keep maxRenderedPages (+ its test,
PDF2XHTML/PDFParser wiring — it's pure parser config, exactly your preferred
vehicle), drop the endpoint, the query
params, ThumbnailDefaults, and ThumbnailSelector, and turn the built-in
defaults JSON into a documented recipe (server docs + the ticket). The recipe
is the deliverable — it's
what the contributor actually needed a central place for.
Literal JSON
Initialization time — tika-config.json on a server that should render
thumbnails on every parse (the forked worker loads this same file, so it
applies to all endpoints with no
request cooperation):
{
"parse-context": {
"pdf-parser": {
"imageStrategy": "RENDER_PAGES_AT_PAGE_END",
"maxRenderedPages": 1,
"ocr": { "dpi": 96, "imageType": "RGB" }
},
"emf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] },
"wmf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] }
}
}
Parse time, scenario A — the contributor's search-indexing case (metadata
+ text + thumbnail in one parse). POST /rmeta/config, multipart parts file +
config, config part:
{
"parse-context": {
"pdf-parser": {
"imageStrategy": "RENDER_PAGES_AT_PAGE_END",
"maxRenderedPages": 1,
"ocr": { "dpi": 96, "imageType": "RGB" }
},
"emf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] },
"wmf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] }
}
}
Parse time, scenario B — "just give me the thumbnail" (replaces
/unpack/thumbnail). POST /unpack, config part:
{
"parse-context": {
"pdf-parser": {
"imageStrategy": "RENDER_PAGES_AT_PAGE_END",
"maxRenderedPages": 1,
"ocr": { "dpi": 96, "imageType": "RGB" }
},
"emf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] },
"wmf-parser": { "renderImage": true,
"renderOnlyEmbeddedResourceTypes": ["THUMBNAIL"] },
"standard-unpack-selector": { "includeEmbeddedResourceTypes":
["THUMBNAIL", "RENDERING"] },
"unpack-config": { "suffixStrategy": "DETECTED", "outputFormat":
"FRICTIONLESS",
"outputMode": "ZIPPED", "includeFullMetadata": true
},
"embedded-limits": { "maxDepth": 3, "maxCount": 20 }
}
}
```
> /unpack/thumbnail: return the document thumbnail with its metadata
> ------------------------------------------------------------------
>
> Key: TIKA-4856
> URL: https://issues.apache.org/jira/browse/TIKA-4856
> Project: Tika
> Issue Type: New Feature
> Reporter: Dominik Schmidt
> Priority: Major
>
> With TIKA-4850 through TIKA-4855 every container format that carries a
> thumbnail emits it as a THUMBNAIL embedded document, the PDF parser renders
> pages as RENDERING documents, and the EMF/WMF renderer turns the vector
> thumbnails of Office documents into raster ones. Getting "the thumbnail of
> this file" out of that still takes format knowledge on the client: the
> THUMBNAIL of a Word or Excel file is an EMF/WMF whose usable form is the
> RENDERING underneath it, a PDF has no THUMBNAIL but a page RENDERING, the
> THUMBNAIL of a DOCX inside a ZIP is not the ZIP's, and with rendering enabled
> the picture of an embedded OLE object is a RENDERING too. Plus the request
> config that switches the renderers on.
> Proposal: POST /unpack/thumbnail next to /unpack and /unpack/all, multipart
> like them. It runs the usual forked parse in unpack mode with a fixed parse
> context (PDF page 1 rendered, EMF/WMF rendered) and picks, in this order: the
> raster THUMBNAIL at depth 1; the rendering of that thumbnail; the depth-1
> RENDERING of PDF page 1. The endpoint extracts what the document carries; it
> does not resize, convert or generate previews.
> The response is JSON: the /rmeta metadata object of the selected embedded
> document, and the image as base64. Thumbnails are small, so the encoding
> overhead does not matter, and the caller gets type, dimensions, origin
> (stored thumbnail or rendering, tk:rendering:rendered-by) and path in one
> round trip without unpacking a zip. 204 when the document has no thumbnail.
> {
> "metadata": {
> "Content-Type": "image/png",
> "Content-Length": "8459",
> "tiff:ImageWidth": "800",
> "tiff:ImageLength": "1131",
> "tk:embedded-resource-type": "RENDERING",
> "tk:embedded-resource-path": "/thumbnail.emf/thumbnail.png",
> "tk:embedded-depth": "2",
> "tk:rendering:rendered-by": "poi-metafile-renderer",
> "tk:resource-name": "thumbnail.png"
> },
> "image": "iVBORw0KGgoAAAANSUhEUgAA..."
> }
> To keep the selection rule short, the metafile renderer could give the
> rendering of a THUMBNAIL the THUMBNAIL type as well (its
> tk:rendering:rendered-by tells it apart), so a raster thumbnail is a
> THUMBNAIL regardless of whether the document stored it as PNG or as EMF.
> What do you think?
--
This message was sent by Atlassian Jira
(v8.20.10#820010)