> Would this be worth mentioning in the design doc?
I would not add this to avoid unnecessary complication in the design doc.
Readers should automatically apply the optimization by reading only the row
groups in the column files that are needed for their tasks. It would be
surprising and wasteful for a reader not to filter row groups by the rows
that it is expected to produce.

On Thu, Aug 13, 2026 at 9:45 AM Anurag Mantripragada <
[email protected]> wrote:

> Hi everyone,
>
> I'm a bit late to this discussion, but it looks like we agree that a
> separate row group alignment flag isn't necessary since we have
> splitOffsets. I plan to update the design doc soon to reflect this and the
> other recent decisions, as things have been moving quickly.
>
> Regarding Peter's questions about how writers would produce aligned row
> groups, I don't have the answers yet. However, I think individual
> implementations can explore that if needed; it shouldn't be part of the
> spec itself.
>
> Thanks,
> Anurag
>
> On Thu, Aug 13, 2026 at 12:43 AM Leonid Lygin via dev <
> [email protected]> wrote:
>
>> Ah, I missed `splitOffsets`, thanks! That really does solve both the
>> planning and the execution problem.
>>
>> Would this be worth mentioning in the design doc?
>>
>> On Thu, Aug 13, 2026 at 7:22 AM Péter Váry <[email protected]>
>> wrote:
>> >
>> > As I wrote earlier, I also see the flag as unnecessary, but having the
>> possibility of writing rowgroup aligned files would be nice to have in the
>> reference implementation. That said, I would prefer to keep this optional,
>> as in some cases it would be beneficial not to align the files.
>> >
>> > On Wed, Aug 12, 2026, 22:27 Russell Spitzer <[email protected]>
>> wrote:
>> >>
>> >> I think Anoop addressed that. We already (optionally) store split
>> offsets, which tell scan planning how to break a file into tasks so
>> different tasks can read different parts of the same file. For Parquet,
>> those are the row-group offsets in the file, so a task is effectively “read
>> bytes X–Y” and can cover a single row group.
>> >>
>> >> At the reader, that is not a different codepath from reading the whole
>> file. The reader still opens the footer, maps the task’s byte range to row
>> group(s), and only reads those groups. With a column file, the extra step
>> is: from the base file row group, take the row indices, open the
>> column-file footer, and select whichever CF row groups cover those rows.
>> >>
>> >> So alignment is still decided after both footers are available. A
>> metadata flag does not change how we plan splits or which CF ranges a task
>> needs.
>> >>
>> >> On Wed, Aug 12, 2026 at 3:08 PM Leonid Lygin via dev <
>> [email protected]> wrote:
>> >>>
>> >>> Thanks for the comments!
>> >>>
>> >>> Don't we also want to know whether the row groups align during scan
>> >>> planning? If they align, we can assume that a task would read
>> >>> significantly less bytes than if they don't.
>> >>>
>> >>> On Wed, Aug 12, 2026 at 7:58 PM Anoop Johnson <[email protected]>
>> wrote:
>> >>> >
>> >>> > I don't see any obvious advantage of keeping track of row group
>> alignment in the column file metadata. As Russell pointed out, the readers
>> must read the Parquet footers anyway before any decoding can happen. At
>> that point, you can derive whether the row groups are aligned by simple
>> integer comparison. Once the reader figures out the alignment, then the
>> optimization Leonid mentioned can still work?
>> >>> >
>> >>> > The closest precedent is where we keep track of the `split_offsets`
>> in the metadata. But that is warranted because we would like to know the
>> offsets at scan planning time before any Parquet footers are actually
>> opened. That too is an optional optimization - if the split_offsets are
>> absent, we fall back to fixed block split computation.
>> >>> >
>> >>> > Best,
>> >>> > Anoop
>> >>> >
>> >>> > On Wed, Aug 12, 2026 at 9:06 AM Russell Spitzer <
>> [email protected]> wrote:
>> >>> >>
>> >>> >> Before adding format-specific performance optimizations like this,
>> I think we should ensure we are solving a real problem and that it is only
>> solvable at the metadata layer.
>> >>> >>
>> >>> >> In this case, can’t the reader tell at read time whether row
>> groups are aligned by comparing footers? The reader must be given both the
>> Base File (BF) and Column File (CF) regardless of alignment, and must open
>> both footers before decoding. At that point it already knows whether BF and
>> CF share the same row-group row boundaries, and can plan further I/O
>> accordingly (1:1 chunk reads when aligned; overlapping CF row groups when
>> not).
>> >>> >>
>> >>> >> Do we get a material benefit from knowing alignment ahead of time
>> in Iceberg metadata, versus deriving it from the footers we have to read
>> anyway?
>> >>> >>
>> >>> >> On Wed, Aug 12, 2026 at 10:44 AM Péter Váry <
>> [email protected]> wrote:
>> >>> >>>
>> >>> >>> Hi Leonid,
>> >>> >>>
>> >>> >>> Thanks for the proposal and for continuing the discussion.
>> >>> >>>
>> >>> >>> I have a couple of questions:
>> >>> >>>
>> >>> >>> Do you have any suggestions for how writers could reliably
>> produce aligned row groups? In particular, how would an update writer
>> obtain the base file's row-group boundaries in practice? Would all writers
>> be expected to read the base file footer? Also, is there an easy way to
>> implement this using the current Java Parquet writer APIs? This sounds like
>> a nice feature.
>> >>> >>> Do readers actually need to know in advance that row groups are
>> aligned? It seems a reader could use essentially the same algorithm for
>> both aligned and unaligned files: seek to the nearest row group, discard
>> any rows prior to the desired starting position, and continue reading from
>> there. If page-level skipping is available, it may even avoid reading some
>> of those unnecessary pages. When the row group already begins at the
>> required offset, the discard step simply becomes a no-op. If that is the
>> case, do we actually need an alignment flag at all?
>> >>> >>>
>> >>> >>> Thanks, Peter
>> >>> >>>
>> >>> >>> Leonid Lygin via dev <[email protected]> ezt írta (időpont:
>> 2026. aug. 12., Sze, 17:06):
>> >>> >>>>
>> >>> >>>> Hi all!
>> >>> >>>>
>> >>> >>>> Following up on the Column File representation thread
>> >>> >>>> <
>> https://lists.apache.org/thread/jbh1gbrso5h6l4by9rh9poy2cjjtb8j0>, I'd
>> like to
>> >>> >>>> fire off a discussion about a possible optimization for Column
>> Updates, where
>> >>> >>>> supporting writers might decide to write out the Column File
>> aligning all row
>> >>> >>>> groups with the Base File, enabling supporting readers to do
>> simple zero-copy
>> >>> >>>> reads.
>> >>> >>>>
>> >>> >>>> *Context*
>> >>> >>>>
>> >>> >>>> The original thread settled on a dense representation (i.e.
>> Column Files contain
>> >>> >>>> exactly the same amount of rows as Base Files), while allowing
>> unsynchronized
>> >>> >>>> row-group boundaries.
>> >>> >>>>
>> >>> >>>> I'm proposing adding a separate metadata flag (e.g.
>> >>> >>>> `ColumnFile.containsAlignedRowGroups`) which, when set,
>> indicates that the row
>> >>> >>>> groups contained in the associated Column File are aligned with
>> the Base File,
>> >>> >>>> and allows readers to directly swap a Base File column chunk
>> with the Column
>> >>> >>>> File one on byte level.
>> >>> >>>>
>> >>> >>>> By "aligned" here I mean that if a Base File contains two row
>> groups with rows
>> >>> >>>> e.g. 0-1000 and 1001-2000, the Column File contains two row
>> groups with exactly
>> >>> >>>> the same row boundaries 0-1000 and 1001-2000; while "unaligned"
>> allows the
>> >>> >>>> Column File to contain row boundaries e.g. 0-1500 and 1501-2000.
>> >>> >>>>
>> >>> >>>> *Performance benefits*
>> >>> >>>>
>> >>> >>>> Assuming a somewhat uniform distrbution of row group byte sizes
>> (say, RG is the
>> >>> >>>> average row group size in bytes), reading a single Column File
>> incurs anywhere
>> >>> >>>> from 0 (row groups already match by chance) to 2*RG (one row
>> before, one row
>> >>> >>>> after) bytes of overhead I/O and memory that is being discarded
>> by the reader.
>> >>> >>>>
>> >>> >>>> Allowing always aligning row group boundaries allows readers to
>> lower that
>> >>> >>>> overhead cost to exactly zero.
>> >>> >>>>
>> >>> >>>> *Pros:*
>> >>> >>>>
>> >>> >>>>     - Writers are not required to implement -- keeping the flag
>> false is still
>> >>> >>>>       valid for any data shape.
>> >>> >>>>     - Readers are not required to implement -- aligning row
>> groups doesn't break
>> >>> >>>>       readers that still want to do stitching on row level.
>> >>> >>>>     - One less copy of data along the way from parquet to
>> executor.
>> >>> >>>>     - Less reader I/O.
>> >>> >>>>
>> >>> >>>> *Cons:*
>> >>> >>>>
>> >>> >>>>     - Implementations are optional -- leads to both diverging
>> features in
>> >>> >>>>       different implementations on one hand, and diverging
>> codepaths within a
>> >>> >>>>       single Iceberg implementation as support for the unaligned
>> case is still
>> >>> >>>>       required.
>> >>> >>>>     - One more field in the Column File struct -- more support
>> work.
>> >>> >>>>
>> >>> >>>> Thanks!
>> >>> >>>> Leonid.
>>
>

Reply via email to