JingsongLi commented on code in PR #10001: URL: https://github.com/apache/paimon/pull/10001#discussion_r4058927408
########## docs/docs/concepts/spec/rowformat.md: ########## @@ -182,11 +182,12 @@ To read a row by its zero-based row number within the file: 1. **Read Footer**: Seek to file end - 32 bytes, read the 32-byte footer. Validate magic number. 2. **Read Block Index**: Seek to `indexOffset`, read `indexLength` bytes, decode the three arrays. Compute block offsets by prefix sum of `blockCompressedSizes[]`. -3. **Select Block**: Find block `b` where `blockRowStarts[b] <= rowNum < blockEnd`. For the last block, `blockEnd` is `totalRowCount`; otherwise it is `blockRowStarts[b + 1]`. -4. **Read Block**: Seek to `blockOffset(b)`, read `blockCompressedSizes[b]` bytes. -5. **Decompress**: ZSTD decompress into a buffer of size `blockUncompressedSizes[b]`. -6. **Locate Row**: Compute `localIdx = rowNum - blockRowStarts[b]`. Read `offsets[localIdx]` from the offset array at the end of the decompressed block. -7. **Deserialize**: Read the row starting at the computed offset using the row serialization format. +3. **Check Consistency**: The three arrays must have the same length, that length must equal `blockCount`, and `blockCompressedSizes[]` must sum to `indexOffset`, because the blocks are written contiguously from position 0 and the index follows the last one. A reader that bounds its block loop by one of the two — the footer's `blockCount` or the index array length — must reject a file where they disagree rather than silently reading fewer blocks. Review Comment: Thanks for adding the Python array-count and row-start checks. One part of the original P1 remains unresolved: FormatRowReader reads index_offset only as a local in _read_metadata, and _validate_block_index still does not compare it with sum(_block_compressed_sizes), despite its docstring and the updated format spec requiring that invariant. Java RowBlockIndex.validate now rejects such a file, while Python accepts it and computes block offsets from the inconsistent sizes, so the cross-language behavior still diverges. Please retain or pass index_offset, validate the exact sum before the empty-index return, and add the Python counterpart of the Java compressed-size/index-offset regression. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
