JingsongLi commented on PR #9743:
URL: https://github.com/apache/paimon/pull/9743#issuecomment-5634788446

   I suggest using a single manifest index sidecar organized by Avro block. 
Partition information would support partition predicate pushdown during 
planning. Since both partition and row-id information describe the same blocks, 
they can live in the same block record and share its physical location.
   
   A possible layout is:
   
   ```text
   Header
     formatVersion
     manifest identity (name hash, file length, entry count)
     original Avro header
   
   Partition dictionary
     partitionId -> complete partition tuple
   
   blockCount : int
   BlockIndexRecord[]                 // original manifest order
     offset      : long               // byte offset in the manifest
     length      : long               // complete Avro block length
     recordCount : long               // number of manifest entries
     flags       : byte               // independent availability bits
     [if ROW_ID_AVAILABLE]
       rangeCount : int
       ranges     : (start: long, end: long)[]   // inclusive interval unions
     [if PARTITION_AVAILABLE]
       partitionIdCount : int
       partitionIds     : int[]        // sorted and deduplicated
   
   Checksum of all preceding bytes
   ```
   
   The partition dictionary is shared across the file and can reuse the 
existing manifest partition encoding, preserving full tuples, types and nulls. 
Each block only stores dictionary IDs. The block ID is implicit in its 
position; `firstRecord` can be derived from preceding `recordCount` values.
   
   The two indexes should remain independently usable within each block:
   
   - An availability bit means that the corresponding information completely 
covers the block's entries, including both ADD and DELETE entries and all 
column groups.
   - If row-id coverage is unknown or exceeds its budget, omit that block's 
row-id payload while retaining its partition information. Apply the same rule 
independently to partition information.
   - An unavailable index means “cannot prune using this index,” rather than an 
empty result. Invalid file metadata or a checksum failure should fall back to 
the normal manifest read.
   
   During planning, evaluate the partition predicate against the dictionary 
once, then check each block's partition IDs and row-id intervals. For 
conjunctive filters, intersect their candidate block sets. Read the selected 
blocks and retain the existing entry filtering and ADD/DELETE merge, since 
block-level matches do not guarantee that the same entry satisfies both 
predicates.
   
   This layout assumes reading the whole sidecar, as the current implementation 
does. A partition-only query would also read the row-id index bytes. I would 
start with this simpler layout and consider separate physical sections if 
measurements show that selective index reads materially improve planning time.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to