JingsongLi commented on code in PR #10001:
URL: https://github.com/apache/paimon/pull/10001#discussion_r4056471363


##########
docs/docs/concepts/spec/rowformat.md:
##########
@@ -182,11 +182,12 @@ To read a row by its zero-based row number within the 
file:
 
 1. **Read Footer**: Seek to file end - 32 bytes, read the 32-byte footer. 
Validate magic number.
 2. **Read Block Index**: Seek to `indexOffset`, read `indexLength` bytes, 
decode the three arrays. Compute block offsets by prefix sum of 
`blockCompressedSizes[]`.
-3. **Select Block**: Find block `b` where `blockRowStarts[b] <= rowNum < 
blockEnd`. For the last block, `blockEnd` is `totalRowCount`; otherwise it is 
`blockRowStarts[b + 1]`.
-4. **Read Block**: Seek to `blockOffset(b)`, read `blockCompressedSizes[b]` 
bytes.
-5. **Decompress**: ZSTD decompress into a buffer of size 
`blockUncompressedSizes[b]`.
-6. **Locate Row**: Compute `localIdx = rowNum - blockRowStarts[b]`. Read 
`offsets[localIdx]` from the offset array at the end of the decompressed block.
-7. **Deserialize**: Read the row starting at the computed offset using the row 
serialization format.
+3. **Check Consistency**: The three arrays must have the same length, that 
length must equal `blockCount`, and `blockCompressedSizes[]` must sum to 
`indexOffset`, because the blocks are written contiguously from position 0 and 
the index follows the last one. A reader that bounds its block loop by one of 
the two — the footer's `blockCount` or the index array length — must reject a 
file where they disagree rather than silently reading fewer blocks.

Review Comment:
   [P1] Apply the new consistency contract to the Python row reader too\n\nThis 
now defines rejection as a format-reader requirement, and the PR description 
specifically calls out that Python bounds iteration by the footer count, but  
still trusts , , and  and never compares the three decoded array lengths or 
their compressed-size sum. For example, changing  to 0 still makes Python 
return an empty result for a non-empty file rather than reject it. Please 
implement the same validation and regression cases in Python so Java and Python 
do not retain the cross-language behavior this change is meant to remove.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to