zjw1111 commented on PR #10101: URL: https://github.com/apache/paimon/pull/10101#issuecomment-5810317738
> This has clear end-to-end value: the current V2 write and rewrite paths still assemble the whole container in memory, so a container over 2 GiB cannot be produced. The V2 footer and spill-to-file path address that production limit, with V1 remaining the default. > > **Blocking compatibility issue (P1):** `FileIndexer.createReader(SeekableInputStream, int, int)` is replaced by the `(long, long)` descriptor. Existing third-party `FileIndexer` implementations loaded through `ServiceLoader` can still be discovered, but the first read now calls a method their binary does not implement and fails with `AbstractMethodError`. This affects existing V1 indexes even when the new V2 option is never enabled. Akash3121 already raised this on `FileIndexer.java:35`; I agree it needs a compatibility bridge and a test using a plugin compiled against the old API before merge. The rollout should also keep the documented reader-first order when V2 is enabled. > > The V1/V2 container tests passed locally (10/10). `DataFileIndexWriterTest` passed locally (5/5), including the over-2-GiB streamed payload and failed-write cleanup cases. The `FileIndexProcessorTest` map-key rewrite cases could not initialize in this local workspace due to the existing `CodeGenerator` service-loading failure; `FileIndexPredicateCloseTest` failed before its logic ran because Mockito could not attach to this JVM. The PR head's JDK 8/11, Flink 1/2, Spark and E2E CI checks are green. No other confirmed correctness regression was found in the format and write/rewrite paths. Done -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
