JustinBinber commented on issue #10021: URL: https://github.com/apache/paimon/issues/10021#issuecomment-5750295090
I’ve started looking into the first part of this problem and built a small proof of concept. My current idea is to keep small Spark `InputPartition`s unchanged, but move oversized split metadata to a shared, seekable `FileIO` path. The task would only carry a small descriptor and read its own byte range on the executor. This would be opt-in, so the existing behavior would not change unless a staging path is configured. For a first PR, I’d like to keep the scope limited to this transport path. It would include: - inline fallback for small partitions; - one shared metadata container per scan; - range reads on executors; - bucket information preservation; - cleanup when planning fails or the Spark application ends; - a versioned frame format with length and checksum validation. I’m also planning to make the writer accept an iterator, so later work on streaming split planning will not require changing the storage format again. The current local probe uses 350 `DataSplit`s with 1,000 `DataFileMeta` entries each. The serialized task payload is reduced from 136,489,062 bytes to about 1.4 KB, while the full metadata remains in the external container. This only demonstrates the reduction in task payload; it does not yet solve driver-side planning memory or executor-side reconstruction of a very large split list. I’d prefer to handle those two problems separately: 1. stream or page split planning on the driver; 2. stream split decoding on the executor; 3. deal with a single exceptionally large `DataSplit`. Does this sound like a reasonable scope for the first PR? If so, I’d be happy to continue with it. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
