JingsongLi commented on PR #10041: URL: https://github.com/apache/paimon/pull/10041#issuecomment-5756657576
Requirement fit: PIVOT (triage: NO-GO for this PR as scoped) Issue #10021 reports Driver OOM or restart around split discovery, task construction, and serialization. This PR only shrinks the serialized Spark `InputPartition` sent to tasks. `PaimonBatch` still starts with fully materialized `inputPartitions` and each partition's `Seq[Split]`; the externalized partition retains `plannedSplits` on the Driver; and the Executor still decodes the full split list. The 136.5 MB probe measures task payload rather than a completed scan. The integration test forces the path with a 1-byte threshold on a tiny table, so it does not show that a workload which currently fails can finish. For this roughly 1,700-line change, the new shared staging path, serialization format, extra encoding pass, and file/broadcast cleanup lifecycle need a demonstrated end-to-end capacity benefit. Please provide a reproducible SQL workload that completes planning but fails specifically at task transport, then show it completes with this option enabled, including peak Driver/Executor memory and runtime at a representative file count. Alternatively, pivot to the planning or decode boundary that blocks the reported workload. I am closing this PR for now; it can be reconsidered with that evidence and a narrower implementation. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
