[
https://issues.apache.org/jira/browse/TIKA-4795?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18110609#comment-18110609
]
Davide Polato commented on TIKA-4795:
-------------------------------------
Implemented on
[TIKA-4795-parseBytes|https://github.com/ai-pipestream/tika/tree/TIKA-4795-parseBytes]
: bounded unary on TikaV2, bytes routed like a tika-server body (inline up to
pipes.maxInlineBytes, spooled above it), content cap and inbound limit as
separate knobs, provenance echoed, fields 7-10 still reserved. Description
updated with what got decided.
Tested end to end from StormCrawler:
[parsebytes-probe|https://github.com/dpol1/parsebytes-probe], 15 checks on 15
URLs, archived run against that commit.
> ParseBytes: parse-only entrypoint for callers that already hold the document
> bytes
> ----------------------------------------------------------------------------------
>
> Key: TIKA-4795
> URL: https://issues.apache.org/jira/browse/TIKA-4795
> Project: Tika
> Issue Type: New Feature
> Components: tika-pipes
> Affects Versions: 4.0.0
> Reporter: Davide Polato
> Priority: Major
> Labels: grpc, pipes, protobuf
>
> h4. Goal
> Add a parse-only tika-grpc entrypoint that parses the exact bytes supplied by
> the caller, without asking Tika to fetch or re-fetch the resource.
> This was split from the typed {{Document}} work in
> [TIKA-4766|https://issues.apache.org/jira/browse/TIKA-4766] / [PR
> #2961|https://github.com/apache/tika/pull/2961], whose description states
> that "{{ParseBytes}} is deferred to its own JIRA."
> h4. Motivation
> {{FetchAndParse}} obtains content through a registered fetcher. Callers such
> as web crawlers have already acquired the resource bytes. Requiring another
> acquisition path can duplicate the transfer or parse a different
> representation from the one the caller observed because of redirects,
> cookies, authentication, robots and rate-limiting decisions, or transient
> content.
> The parse result should correspond to the exact representation captured by
> the caller.
> Apache StormCrawler is a concrete consumer. Its parsing bolts already receive
> the acquired bytes together with URL and protocol metadata. StormCrawler can
> contribute crawler requirements and input fixtures including truncated
> payloads, declared-vs-actual charset mismatches, malformed archives,
> documents containing embedded resources, HTML with {{<base>}}, and distinct
> URLs with identical content.
> h4. Proposed observable behavior
> This issue defines observable behavior, not the internal buffering, fetcher,
> or Pipes implementation.
> The request carries the exact byte sequence. It may also carry:
> * an optional opaque caller correlation id, echoed but never interpreted
> * source URI and effective URI when available, plus a base URI for link
> resolution; these are provenance metadata and must never be dereferenced by
> the service
> * resource name
> * declared media type and charset hints
> * optional declared content length and digest
> * a truncation flag
> Input must be bounded, and deadline or cancellation must terminate the
> associated parse work.
> The reply reuses the typed {{Document}} contract introduced by #2961 rather
> than defining a second parse-result model.
> The current direction discussed in #2961 is to place this on the experimental
> v2 surface so that it does not gate the 4.0.0 release.
> h4. Decided in the branch (2026-09-02)
> * bounded unary {{TikaV2.ParseBytes}} on the experimental v2 surface; the
> reply is the typed {{Document}} of #2961
> * provenance carried and echoed: correlation id (as {{Document.id}}, the
> caller's key), source/effective/base URI, resource name, truncation flag.
> Declared media type, charset, length and digest keep field numbers 7-10
> reserved until they have verifiable semantics
> * payload bounded by a content cap ({{grpc.parseBytesMaxContentBytes}}, 64
> MiB by default), separate from the gRPC inbound limit; bytes routed like a
> tika-server request body, inline up to {{pipes.maxInlineBytes}}, spooled
> above it
> * branch
> [TIKA-4795-parseBytes|https://github.com/ai-pipestream/tika/tree/TIKA-4795-parseBytes],
> stacked on #2961; PR once #2961 lands
> h4. Non-goals
> This issue does not define the structured content tree or its Markdown
> projection, the embedded-document representation, extension payloads such as
> {{google.protobuf.Any}}, or downstream NLP and embedding enrichment. Those
> remain separate TIKA-4766 stages and follow-up issues.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)