JingsongLi commented on PR #9743:
URL: https://github.com/apache/paimon/pull/9743#issuecomment-5645815306

   Here is a refined version of the block-oriented layout, keeping the 
file-level partition dictionary and making each block's two payloads 
independently extensible.
   
   ```text
   Header
     magic
     formatVersion
     manifest identity (name hash, file length, entry count)
     avroHeaderLength : int
     original Avro header : bytes
   
   Partition Dictionary
     partitionCount : int
     entries[]                           // position is the partition ID
       partitionByteLength : int
       partitionBytes : bytes
   
   blockCount : int
   BlockIndexRecord[]                    // original manifest order
     offset      : long                 // byte offset in the manifest
     length      : long                 // complete Avro block length
     recordCount : long                 // number of manifest entries
   
     partitionEncoding      : byte
     partitionPayloadLength : int
     partitionPayload       : bytes
   
     rowIdEncoding          : byte
     rowIdPayloadLength     : int
     rowIdPayload           : bytes
   
   Checksum of all preceding bytes
   ```
   
   The dictionary stores each complete partition tuple once, using the existing 
manifest partition serialization. This preserves tuple values and nulls; the 
scan's existing `partitionType` supplies their interpretation. Blocks reference 
dictionary IDs. The block ID is implicit in its position, and `firstRecord` is 
derived from preceding entry counts.
   
   The encoding bytes identify **how to decode the corresponding payload**, 
with separate ID namespaces for partition and row-id payloads. They replace the 
availability flags:
   
   | Field | Encoding | Meaning and payload |
   | --- | --- | --- |
   | `partitionEncoding` | `0` | Partition coverage is unavailable. Payload 
length must be zero. |
   | `partitionEncoding` | `1` | Complete partition ID set: `partitionIdCount: 
int`, followed by that many sorted, unique `partitionId: int` values. Every ID 
references the file-level dictionary. |
   | `rowIdEncoding` | `0` | Row-id coverage is unavailable. Payload length 
must be zero. |
   | `rowIdEncoding` | `1` | Conservative interval coverage: `rangeCount: int`, 
followed by that many inclusive `(start: long, end: long)` pairs, sorted and 
disjoint. |
   
   The container's integers and the encoding-1 payload integers use fixed-width 
big-endian representation; partition bytes retain their existing serialization. 
Encoding bytes are interpreted as unsigned IDs. Each payload length counts only 
its payload bytes, excluding the encoding and length fields.
   
   Other nonzero encoding IDs are reserved for future representations. If a 
reader does not recognize one, it skips exactly that payload length and treats 
that dimension as unavailable, while still being able to use the other 
dimension. Lengths must be bounded and validated. The outer `formatVersion` 
governs the container and dictionary framing; unsupported container versions or 
malformed metadata/payloads fall back to the normal manifest read.
   
   For example, `rowIdEncoding=1` with `rangeCount=2` and ranges `[100,109]`, 
`[300,309]` has a 36-byte payload: `4 + 2 * 16`.
   
   There are several important correctness and budget rules:
   
   - Encoding 0 means “cannot prune using this information,” never “no 
matches.” An available payload must cover all relevant entries in the block, 
including ADD, DELETE and all column groups.
   - Row-id coverage may be a conservative superset. If exact interval unions 
exceed the budget, merge intervals; the coarsest representation is 
`rangeCount=1, [min,max]`, still using encoding 1. Continue processing the 
entire block to extend the bounds and detect unknown row IDs. If complete 
coverage cannot be established, use encoding 0.
   - Partition information can independently become unavailable when its budget 
is exceeded. Consequently, the global dictionary is not necessarily a complete 
list of partitions touched by the manifest. A dictionary miss must not 
eliminate blocks with unavailable partition coverage.
   - The physical block directory must always cover the entire manifest. Budget 
exhaustion may omit optional index payloads, but must never omit block 
descriptors. Validate byte coverage and entry counts, and verify the whole-file 
checksum before making pruning decisions.
   
   For conjunctive partition and row-id filters, select each block using:
   
   ```text
   keepBlock =
       (partition coverage unavailable || partition predicate matches)
       &&
       (row-id coverage unavailable || query intersects indexed ranges)
   ```
   
   Only an empty candidate block set permits skipping the manifest. Selected 
blocks still pass through the existing entry filtering and ADD/DELETE merge.
   
   This keeps one sidecar and one record per block. It still assumes a bounded 
whole-sidecar read: payload lengths allow skipping decoding and unknown 
encodings, but do not by themselves save storage I/O. Index size, block 
selectivity and planning latency should determine whether selective physical 
reads are worthwhile later.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to