JingsongLi commented on PR #9743:
URL: https://github.com/apache/paimon/pull/9743#issuecomment-5634788446
I suggest using a single manifest index sidecar organized by Avro block.
Partition information would support partition predicate pushdown during
planning. Since both partition and row-id information describe the same blocks,
they can live in the same block record and share its physical location.
A possible layout is:
```text
Header
formatVersion
manifest identity (name hash, file length, entry count)
original Avro header
Partition dictionary
partitionId -> complete partition tuple
blockCount : int
BlockIndexRecord[] // original manifest order
offset : long // byte offset in the manifest
length : long // complete Avro block length
recordCount : long // number of manifest entries
flags : byte // independent availability bits
[if ROW_ID_AVAILABLE]
rangeCount : int
ranges : (start: long, end: long)[] // inclusive interval unions
[if PARTITION_AVAILABLE]
partitionIdCount : int
partitionIds : int[] // sorted and deduplicated
Checksum of all preceding bytes
```
The partition dictionary is shared across the file and can reuse the
existing manifest partition encoding, preserving full tuples, types and nulls.
Each block only stores dictionary IDs. The block ID is implicit in its
position; `firstRecord` can be derived from preceding `recordCount` values.
The two indexes should remain independently usable within each block:
- An availability bit means that the corresponding information completely
covers the block's entries, including both ADD and DELETE entries and all
column groups.
- If row-id coverage is unknown or exceeds its budget, omit that block's
row-id payload while retaining its partition information. Apply the same rule
independently to partition information.
- An unavailable index means “cannot prune using this index,” rather than an
empty result. Invalid file metadata or a checksum failure should fall back to
the normal manifest read.
During planning, evaluate the partition predicate against the dictionary
once, then check each block's partition IDs and row-id intervals. For
conjunctive filters, intersect their candidate block sets. Read the selected
blocks and retain the existing entry filtering and ADD/DELETE merge, since
block-level matches do not guarantee that the same entry satisfies both
predicates.
This layout assumes reading the whole sidecar, as the current implementation
does. A partition-only query would also read the row-id index bytes. I would
start with this simpler layout and consider separate physical sections if
measurements show that selective index reads materially improve planning time.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]