JustinBinber opened a new pull request, #10388: URL: https://github.com/apache/paimon/pull/10388
### Purpose Large append-only tables can run out of Driver memory before Spark starts any task because the scan first materializes every manifest entry and then builds the complete split list. This PR adds a Core-level, single-pass planning API for that boundary. It is intentionally narrower than #10041: it does not add Spark metadata staging, executor decoding, or commit-side changes. ### What changed - add a closeable `StreamingPlan` over effective manifest entries; - add a two-phase `SplitPlan` API so snapshot metadata is fixed before split iteration starts; - expose one-file, raw-convertible splits only for append-only scans; - preserve the existing eager path for limits, Top-N, auth conversion, deletion vectors, partition sorting, deleted manifests, and unsupported table types; - create the batch read-protection tag before the first split can be consumed; - preserve snapshot ID, watermark, filters, sidecar pruning, scan metrics, and resource cleanup on fallback paths. The existing `plan()` API is unchanged. Other scan implementations use the old materialized path through default methods. ### Why fine-grained splits The Core iterator deliberately does not reproduce the old task grouping. It yields independently readable append-only files and leaves bounded bin-packing to the connector. This avoids rebuilding a complete split list inside Core while keeping connector-specific task sizing out of the storage layer. ### Validation - 225 focused Core tests passed, including add-only streaming, deleted-manifest fallback, direct limit fallback, read-protection ordering, and scan-duration semantics. - full Core suite: 5,925 tests, 0 failures, 0 errors, 36 skipped; - dependency modules in the same reactor: 14,463 tests, 0 failures, 0 errors, 6 skipped; - Spotless, Checkstyle, RAT, JDK 8 clean compile/testCompile, and `git diff --check` passed. I also tested the complete follow-up stack on a frozen one-million-file table under a 1 GiB Driver and 512 MiB Executor. The default path failed while serializing the real Spark task; the integrated single-pass planning + bounded packing + external transport path completed twice and returned the expected 2,000,000 rows. That experiment demonstrates the end-to-end direction, not the isolated contribution of this Core-only PR. The Spark consumer and commit-side work will be proposed separately. This is a pivot to the planning boundary requested in #10041 and is part of #10021. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
