JustinBinber opened a new pull request, #10041:
URL: https://github.com/apache/paimon/pull/10041

   Part of #10021.
   
   ## Purpose
   
   Spark currently embeds every Paimon `Split` in an `InputPartition`. For 
scans with many file descriptors, the serialized task payload can become large 
enough to put pressure on driver memory, RPC, retries, and speculation.
   
   This PR adds an opt-in transport path for oversized input partition 
metadata. Small partitions keep the existing inline behavior. Large partitions 
are written to a shared seekable `FileIO` container, while Spark tasks carry 
only a compact descriptor and read their own byte range on the executor.
   
   ## Changes
   
   - Add session options for the external metadata path and inline threshold. 
The feature is disabled unless a path is configured; the default threshold is 
128 MiB.
   - Add a versioned, framed container format with per-split lengths, CRC32 
validation, and footer statistics.
   - Share one metadata container per scan and preserve bucket information in 
the external descriptor.
   - Read external metadata through range reads on executors.
   - Keep row count, split count, and estimated data bytes in the descriptor so 
metrics do not need to decode the payload.
   - Clean up staged metadata on planning failure and at Spark application 
shutdown.
   - Document the configuration and its trust requirements.
   
   ## Scope
   
   This PR only bounds the metadata carried by Spark tasks. It does not yet 
make driver-side split planning streaming, stream split decoding on executors, 
or split one exceptionally large `DataSplit`. Those are intentionally left as 
separate follow-up changes discussed in #10021.
   
   ## Validation
   
   - Spark 3 / JDK 8: focused unit tests and full `SparkReadITCase`
   - Spark 4.1.2 / Scala 2.13 / JDK 17: the same test matrix
   - Spotless, Checkstyle, Apache RAT, and `git diff --check`
   
   A metadata-only probe with 350 `DataSplit`s and 1,000 file descriptors per 
split reduced the serialized `InputPartition` payload from about 136.5 MB to a 
small descriptor. The full metadata remains in the external container; this 
result demonstrates task-payload reduction, not an end-to-end scan speedup.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to