Thanks for the comments! Don't we also want to know whether the row groups align during scan planning? If they align, we can assume that a task would read significantly less bytes than if they don't.
On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]> wrote: > > I don't see any obvious advantage of keeping track of row group alignment in > the column file metadata. As Russell pointed out, the readers must read the > Parquet footers anyway before any decoding can happen. At that point, you can > derive whether the row groups are aligned by simple integer comparison. Once > the reader figures out the alignment, then the optimization Leonid mentioned > can still work? > > The closest precedent is where we keep track of the `split_offsets` in the > metadata. But that is warranted because we would like to know the offsets at > scan planning time before any Parquet footers are actually opened. That too > is an optional optimization - if the split_offsets are absent, we fall back > to fixed block split computation. > > Best, > Anoop > > On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer <[email protected]> > wrote: >> >> Before adding format-specific performance optimizations like this, I think >> we should ensure we are solving a real problem and that it is only solvable >> at the metadata layer. >> >> In this case, can’t the reader tell at read time whether row groups are >> aligned by comparing footers? The reader must be given both the Base File >> (BF) and Column File (CF) regardless of alignment, and must open both >> footers before decoding. At that point it already knows whether BF and CF >> share the same row-group row boundaries, and can plan further I/O >> accordingly (1:1 chunk reads when aligned; overlapping CF row groups when >> not). >> >> Do we get a material benefit from knowing alignment ahead of time in Iceberg >> metadata, versus deriving it from the footers we have to read anyway? >> >> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry <[email protected]> >> wrote: >>> >>> Hi Leonid, >>> >>> Thanks for the proposal and for continuing the discussion. >>> >>> I have a couple of questions: >>> >>> Do you have any suggestions for how writers could reliably produce aligned >>> row groups? In particular, how would an update writer obtain the base >>> file's row-group boundaries in practice? Would all writers be expected to >>> read the base file footer? Also, is there an easy way to implement this >>> using the current Java Parquet writer APIs? This sounds like a nice feature. >>> Do readers actually need to know in advance that row groups are aligned? It >>> seems a reader could use essentially the same algorithm for both aligned >>> and unaligned files: seek to the nearest row group, discard any rows prior >>> to the desired starting position, and continue reading from there. If >>> page-level skipping is available, it may even avoid reading some of those >>> unnecessary pages. When the row group already begins at the required >>> offset, the discard step simply becomes a no-op. If that is the case, do we >>> actually need an alignment flag at all? >>> >>> Thanks, Peter >>> >>> Leonid Lygin via dev <[email protected]> ezt írta (időpont: 2026. aug. >>> 12., Sze, 17:06): >>>> >>>> Hi all! >>>> >>>> Following up on the Column File representation thread >>>> <https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, I'd >>>> like to >>>> fire off a discussion about a possible optimization for Column Updates, >>>> where >>>> supporting writers might decide to write out the Column File aligning all >>>> row >>>> groups with the Base File, enabling supporting readers to do simple >>>> zero-copy >>>> reads. >>>> >>>> *Context* >>>> >>>> The original thread settled on a dense representation (i.e. Column Files >>>> contain >>>> exactly the same amount of rows as Base Files), while allowing >>>> unsynchronized >>>> row-group boundaries. >>>> >>>> I'm proposing adding a separate metadata flag (e.g. >>>> `ColumnFile.containsAlignedRowGroups`) which, when set, indicates that the >>>> row >>>> groups contained in the associated Column File are aligned with the Base >>>> File, >>>> and allows readers to directly swap a Base File column chunk with the >>>> Column >>>> File one on byte level. >>>> >>>> By "aligned" here I mean that if a Base File contains two row groups with >>>> rows >>>> e.g. 0-1000 and 1001-2000, the Column File contains two row groups with >>>> exactly >>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned" allows the >>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000. >>>> >>>> *Performance benefits* >>>> >>>> Assuming a somewhat uniform distrbution of row group byte sizes (say, RG >>>> is the >>>> average row group size in bytes), reading a single Column File incurs >>>> anywhere >>>> from 0 (row groups already match by chance) to 2*RG (one row before, one >>>> row >>>> after) bytes of overhead I/O and memory that is being discarded by the >>>> reader. >>>> >>>> Allowing always aligning row group boundaries allows readers to lower that >>>> overhead cost to exactly zero. >>>> >>>> *Pros:* >>>> >>>> - Writers are not required to implement -- keeping the flag false is >>>> still >>>> valid for any data shape. >>>> - Readers are not required to implement -- aligning row groups doesn't >>>> break >>>> readers that still want to do stitching on row level. >>>> - One less copy of data along the way from parquet to executor. >>>> - Less reader I/O. >>>> >>>> *Cons:* >>>> >>>> - Implementations are optional -- leads to both diverging features in >>>> different implementations on one hand, and diverging codepaths >>>> within a >>>> single Iceberg implementation as support for the unaligned case is >>>> still >>>> required. >>>> - One more field in the Column File struct -- more support work. >>>> >>>> Thanks! >>>> Leonid.
