jackylee-ch opened a new pull request, #10069:
URL: https://github.com/apache/paimon/pull/10069

   ### Purpose
   
   Follow-up to #10001, which had the two readers enforce seven block-index 
invariants while the spec stated three. A writer built against the published 
spec could emit a file both readers reject, and there is a third implementation 
— `paimon-rust`, exercised in the Python CI lane.
   
   The Block Index section now states all of them: no negative size (the sum 
does not imply it, two sizes can cancel), `blockRowStarts[0] == 0` with 
strictly increasing starts, the last start below `totalRowCount`, and an empty 
index only for an empty file. The lookup algorithm points at that list instead 
of restating a subset.
   
   Two fixes came out of writing it down. `blockUncompressedSizes` was the one 
array with no value constraint — the sum bounds the compressed sizes 
transitively, nothing bounded these, and they size the decompression buffer. 
And the Python docstring described Java's failure mode: this reader bisects the 
row starts rather than skipping a block, so a first start past 0 gives 
`block_idx == -1`, a negative local row, and a row decoded from the wrong place.
   
   ### Tests
   
   Both sides gain the negative-uncompressed-size case; Python also gains the 
single-block/zero-rows case Java already had.
   
   `paimon-format`: 746 run, 0 failures. `pypaimon`: 25 passed, flake8 clean.
   
   Written with Claude Code; verification is mine.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to