JingsongLi commented on code in PR #10001: URL: https://github.com/apache/paimon/pull/10001#discussion_r4056471363
########## docs/docs/concepts/spec/rowformat.md: ########## @@ -182,11 +182,12 @@ To read a row by its zero-based row number within the file: 1. **Read Footer**: Seek to file end - 32 bytes, read the 32-byte footer. Validate magic number. 2. **Read Block Index**: Seek to `indexOffset`, read `indexLength` bytes, decode the three arrays. Compute block offsets by prefix sum of `blockCompressedSizes[]`. -3. **Select Block**: Find block `b` where `blockRowStarts[b] <= rowNum < blockEnd`. For the last block, `blockEnd` is `totalRowCount`; otherwise it is `blockRowStarts[b + 1]`. -4. **Read Block**: Seek to `blockOffset(b)`, read `blockCompressedSizes[b]` bytes. -5. **Decompress**: ZSTD decompress into a buffer of size `blockUncompressedSizes[b]`. -6. **Locate Row**: Compute `localIdx = rowNum - blockRowStarts[b]`. Read `offsets[localIdx]` from the offset array at the end of the decompressed block. -7. **Deserialize**: Read the row starting at the computed offset using the row serialization format. +3. **Check Consistency**: The three arrays must have the same length, that length must equal `blockCount`, and `blockCompressedSizes[]` must sum to `indexOffset`, because the blocks are written contiguously from position 0 and the index follows the last one. A reader that bounds its block loop by one of the two — the footer's `blockCount` or the index array length — must reject a file where they disagree rather than silently reading fewer blocks. Review Comment: [P1] Apply the new consistency contract to the Python row reader too\n\nThis now defines rejection as a format-reader requirement, and the PR description specifically calls out that Python bounds iteration by the footer count, but still trusts , , and and never compares the three decoded array lengths or their compressed-size sum. For example, changing to 0 still makes Python return an empty result for a non-empty file rather than reject it. Please implement the same validation and regression cases in Python so Java and Python do not retain the cross-language behavior this change is meant to remove. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
