JingsongLi commented on PR #9743:
URL: https://github.com/apache/paimon/pull/9743#issuecomment-5645815306
Here is a refined version of the block-oriented layout, keeping the
file-level partition dictionary and making each block's two payloads
independently extensible.
```text
Header
magic
formatVersion
manifest identity (name hash, file length, entry count)
avroHeaderLength : int
original Avro header : bytes
Partition Dictionary
partitionCount : int
entries[] // position is the partition ID
partitionByteLength : int
partitionBytes : bytes
blockCount : int
BlockIndexRecord[] // original manifest order
offset : long // byte offset in the manifest
length : long // complete Avro block length
recordCount : long // number of manifest entries
partitionEncoding : byte
partitionPayloadLength : int
partitionPayload : bytes
rowIdEncoding : byte
rowIdPayloadLength : int
rowIdPayload : bytes
Checksum of all preceding bytes
```
The dictionary stores each complete partition tuple once, using the existing
manifest partition serialization. This preserves tuple values and nulls; the
scan's existing `partitionType` supplies their interpretation. Blocks reference
dictionary IDs. The block ID is implicit in its position, and `firstRecord` is
derived from preceding entry counts.
The encoding bytes identify **how to decode the corresponding payload**,
with separate ID namespaces for partition and row-id payloads. They replace the
availability flags:
| Field | Encoding | Meaning and payload |
| --- | --- | --- |
| `partitionEncoding` | `0` | Partition coverage is unavailable. Payload
length must be zero. |
| `partitionEncoding` | `1` | Complete partition ID set: `partitionIdCount:
int`, followed by that many sorted, unique `partitionId: int` values. Every ID
references the file-level dictionary. |
| `rowIdEncoding` | `0` | Row-id coverage is unavailable. Payload length
must be zero. |
| `rowIdEncoding` | `1` | Conservative interval coverage: `rangeCount: int`,
followed by that many inclusive `(start: long, end: long)` pairs, sorted and
disjoint. |
The container's integers and the encoding-1 payload integers use fixed-width
big-endian representation; partition bytes retain their existing serialization.
Encoding bytes are interpreted as unsigned IDs. Each payload length counts only
its payload bytes, excluding the encoding and length fields.
Other nonzero encoding IDs are reserved for future representations. If a
reader does not recognize one, it skips exactly that payload length and treats
that dimension as unavailable, while still being able to use the other
dimension. Lengths must be bounded and validated. The outer `formatVersion`
governs the container and dictionary framing; unsupported container versions or
malformed metadata/payloads fall back to the normal manifest read.
For example, `rowIdEncoding=1` with `rangeCount=2` and ranges `[100,109]`,
`[300,309]` has a 36-byte payload: `4 + 2 * 16`.
There are several important correctness and budget rules:
- Encoding 0 means “cannot prune using this information,” never “no
matches.” An available payload must cover all relevant entries in the block,
including ADD, DELETE and all column groups.
- Row-id coverage may be a conservative superset. If exact interval unions
exceed the budget, merge intervals; the coarsest representation is
`rangeCount=1, [min,max]`, still using encoding 1. Continue processing the
entire block to extend the bounds and detect unknown row IDs. If complete
coverage cannot be established, use encoding 0.
- Partition information can independently become unavailable when its budget
is exceeded. Consequently, the global dictionary is not necessarily a complete
list of partitions touched by the manifest. A dictionary miss must not
eliminate blocks with unavailable partition coverage.
- The physical block directory must always cover the entire manifest. Budget
exhaustion may omit optional index payloads, but must never omit block
descriptors. Validate byte coverage and entry counts, and verify the whole-file
checksum before making pruning decisions.
For conjunctive partition and row-id filters, select each block using:
```text
keepBlock =
(partition coverage unavailable || partition predicate matches)
&&
(row-id coverage unavailable || query intersects indexed ranges)
```
Only an empty candidate block set permits skipping the manifest. Selected
blocks still pass through the existing entry filtering and ADD/DELETE merge.
This keeps one sidecar and one record per block. It still assumes a bounded
whole-sidecar read: payload lengths allow skipping decoding and unknown
encodings, but do not by themselves save storage I/O. Index size, block
selectivity and planning latency should determine whether selective physical
reads are worthwhile later.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]