[ 
https://issues.apache.org/jira/browse/TIKA-4856?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18113032#comment-18113032
 ] 

ASF GitHub Bot commented on TIKA-4856:
--------------------------------------

dschmidt commented on PR #3148:
URL: https://github.com/apache/tika/pull/3148#issuecomment-5592445393

    High level this fits both of my use cases, but it ships only one of them.
   
   The "just give me the thumbnail" case is covered by this PR as is: the 
selector keeps exactly one file in the zip, so my client doesn't need any 
picking logic anymore, and NO_OCR + skipOcr + maxPages 1 keep the parse cheap. 
PUT /unpack/preset/thumbnails, done. That fully replaces my endpoint from #3096.
   
   My other use case is indexing: metadata, full text and all embedded files in 
one request, with OCR left on, but previews still rendered (PDF page 1, 
EMF/WMF) so I get dimensions etc. That needs a second catalog preset: no 
selector (I want all embedded files), no NO_OCR, and crucially maxRenderedPages 
instead of maxPages, because cutting the text off at page 1 is exactly what I 
can't have there. Something like:
   
   ```json
    "render-thumbnails": {
       "pdf-parser": { "imageStrategy": "RENDER_PAGES_AT_PAGE_END", 
"maxRenderedPages": 1,
                      "ocr": { "dpi": 96, "imageType": "RGB" } },
      "emf-parser": { "renderImage": true, "renderOnlyEmbeddedResourceTypes": 
["THUMBNAIL"] },
      "wmf-parser": { "renderImage": true, "renderOnlyEmbeddedResourceTypes": 
["THUMBNAIL"] }
    }
   ```
   
   That's also the reason for #3146: the thumbnails preset here can arguably 
keep maxPages 1 (cheaper, and /unpack discards the text anyway), but the 
indexing preset can't work without maxRenderedPages.




> /unpack/thumbnail: return the document thumbnail with its metadata
> ------------------------------------------------------------------
>
>                 Key: TIKA-4856
>                 URL: https://issues.apache.org/jira/browse/TIKA-4856
>             Project: Tika
>          Issue Type: New Feature
>            Reporter: Dominik Schmidt
>            Priority: Major
>
> With TIKA-4850 through TIKA-4855 every container format that carries a 
> thumbnail emits it as a THUMBNAIL embedded document, the PDF parser renders 
> pages as RENDERING documents, and the EMF/WMF renderer turns the vector 
> thumbnails of Office documents into raster ones. Getting "the thumbnail of 
> this file" out of that still takes format knowledge on the client: the 
> THUMBNAIL of a Word or Excel file is an EMF/WMF whose usable form is the 
> RENDERING underneath it, a PDF has no THUMBNAIL but a page RENDERING, the 
> THUMBNAIL of a DOCX inside a ZIP is not the ZIP's, and with rendering enabled 
> the picture of an embedded OLE object is a RENDERING too. Plus the request 
> config that switches the renderers on.
> Proposal: POST /unpack/thumbnail next to /unpack and /unpack/all, multipart 
> like them. It runs the usual forked parse in unpack mode with a fixed parse 
> context (PDF page 1 rendered, EMF/WMF rendered) and picks, in this order: the 
> raster THUMBNAIL at depth 1; the rendering of that thumbnail; the depth-1 
> RENDERING of PDF page 1. The endpoint extracts what the document carries; it 
> does not resize, convert or generate previews.
> The response is JSON: the /rmeta metadata object of the selected embedded 
> document, and the image as base64. Thumbnails are small, so the encoding 
> overhead does not matter, and the caller gets type, dimensions, origin 
> (stored thumbnail or rendering, tk:rendering:rendered-by) and path in one 
> round trip without unpacking a zip. 204 when the document has no thumbnail.
> {
>   "metadata": {
>     "Content-Type": "image/png",
>     "Content-Length": "8459",
>     "tiff:ImageWidth": "800",
>     "tiff:ImageLength": "1131",
>     "tk:embedded-resource-type": "RENDERING",
>     "tk:embedded-resource-path": "/thumbnail.emf/thumbnail.png",
>     "tk:embedded-depth": "2",
>     "tk:rendering:rendered-by": "poi-metafile-renderer",
>     "tk:resource-name": "thumbnail.png"
>   },
>   "image": "iVBORw0KGgoAAAANSUhEUgAA..."
> }
> To keep the selection rule short, the metafile renderer could give the 
> rendering of a THUMBNAIL the THUMBNAIL type as well (its 
> tk:rendering:rendered-by tells it apart), so a raster thumbnail is a 
> THUMBNAIL regardless of whether the document stored it as PNG or as EMF.
> What do you think? 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to