jackylee-ch opened a new pull request, #10001: URL: https://github.com/apache/paimon/pull/10001
### Purpose The footer and the block index describe the same blocks twice. `blockCount` is written into every row file and never read: Java bounds its block loop by the index array length, the Python reader by the footer field. Rewriting `blockCount` to 99 in a 1000-row file changes nothing in Java — it opens, and all 1000 rows come back. Cross-check them where both are in hand. `blockCompressedSizes` must sum to `indexOffset`, since blocks are written contiguously from position 0 and the index follows the last one; the three arrays must agree in length; that length must equal `blockCount`; row starts must not go backwards. The spec asserted the prefix-sum rule in prose, so its lookup algorithm gains the consistency step. Footer offsets are bounded against the file size too, which is what makes `new byte[indexLength]` safe — an `indexOffset` outside the file was an `ArrayIndexOutOfBoundsException` from `arraycopy`. Under `ignore-corrupt-files` such a file is skipped either way, so the named error only helps readers that did not opt into lenience. `RowFileFooter.readFrom(SeekableInputStream, long)` had no callers and goes with it. ### Tests `RowFileIndexConsistencyTest`. `paimon-format`: 710 run, 0 failures. Written with Claude Code; verification is mine. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
