jackylee-ch opened a new pull request, #10069: URL: https://github.com/apache/paimon/pull/10069
### Purpose Follow-up to #10001, which had the two readers enforce seven block-index invariants while the spec stated three. A writer built against the published spec could emit a file both readers reject, and there is a third implementation — `paimon-rust`, exercised in the Python CI lane. The Block Index section now states all of them: no negative size (the sum does not imply it, two sizes can cancel), `blockRowStarts[0] == 0` with strictly increasing starts, the last start below `totalRowCount`, and an empty index only for an empty file. The lookup algorithm points at that list instead of restating a subset. Two fixes came out of writing it down. `blockUncompressedSizes` was the one array with no value constraint — the sum bounds the compressed sizes transitively, nothing bounded these, and they size the decompression buffer. And the Python docstring described Java's failure mode: this reader bisects the row starts rather than skipping a block, so a first start past 0 gives `block_idx == -1`, a negative local row, and a row decoded from the wrong place. ### Tests Both sides gain the negative-uncompressed-size case; Python also gains the single-block/zero-rows case Java already had. `paimon-format`: 746 run, 0 failures. `pypaimon`: 25 passed, flake8 clean. Written with Claude Code; verification is mine. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
