Hi Xander,

You can follow the V4 Single File Commits Sync
<https://www.youtube.com/watch?v=G5m039mXgG4> and subscribe to community
sync calendar
<https://calendar.google.com/calendar/u/0?cid=MzkwNWQ0OTJmMWI0NTBiYTA3MTJmMmFlNmFmYTc2ZWI3NTdmMTNkODUyMjBjYzAzYWE0NTI3ODg1YWRjNTYyOUBncm91cC5jYWxlbmRhci5nb29nbGUuY29t>
.

Regards,
Manu

On Mon, Jul 27, 2026 at 7:09 PM Alexander Bailey <[email protected]>
wrote:

> Thanks for the reply, is there anywhere I can follow that work? Which
> community sync is best to attend for that etc?
>
> Best,
> Xander
>
> On Fri, 24 Jul 2026 at 18:01, Ryan Blue <[email protected]> wrote:
>
>> Quick note: we are planning changes in v4 that tighten requirements for
>> change detection across versions. That way the spec will support all types
>> of incremental scans without relying on conventions in a particular
>> implementation.
>>
>> On Thu, Jul 23, 2026 at 4:52 PM Xiening Dai <[email protected]> wrote:
>>
>>> Based on the current design of the IRC incremental scan API (which
>>> returns a set of files), it can only support append only case. But to me,
>>> incremental scan is really scanning the delta between two snapshots, so it
>>> should also include deleted and updated tuples. This would become very
>>> powerful building block for scenarios like MV refresh. But to support that
>>> it would have to return tuple set, instead of file set, and that would
>>> depart from the norms of IRC APIs.
>>>
>>> So I think at this point, we would need to clarify that `plan` API only
>>> support append-only case for incremental scan. For true data scan, it needs
>>> a good discussion on how to support that from the API/spec level.
>>>
>>> On 2026/07/23 22:06:59 Alexander Bailey wrote:
>>> > Hi all,
>>> >
>>> > I'd like to open a discussion about whether incremental scan semantics
>>> > (specifically
>>> > incremental append scans) should be codified in the table spec, and if
>>> so,
>>> > what the contract should be.
>>> >
>>> > The gap:
>>> >
>>> > Today the spec is effectively silent on incremental scans. In
>>> > practice, the Java
>>> > IncrementalAppendScan is the de facto reference implementation, and
>>> everyone
>>> > else is following it:
>>> >
>>> > - PyIceberg -- implemented by following Java
>>> > github.com/apache/iceberg-python/pull/3512
>>> >
>>> > - iceberg-cpp -- incremental scan API following Java:
>>> > github.com/apache/iceberg-cpp/pull/559
>>> >
>>> > - iceberg-rust -- I have a PR up that, likewise, just mirrors Java
>>> >
>>> >   github.com/apache/iceberg-rust/pull/2337
>>> >
>>> >
>>> > So we have four implementations converging on behavior that lives only
>>> in one
>>> > codebase. That's fragile: any of Java's incidental choices become the
>>> unwritten
>>> > contract, and there's no spec to point to when implementations
>>> disagree.
>>> >
>>> > There's a good precedent for what the fix could look like. The spec
>>> already has
>>> > a "Point-in-Time Reads (Time Travel)" section (
>>> iceberg.apache.org/spec#…
>>> > <https://iceberg.apache.org/spec/#point-in-time-reads-time-travel>)
>>> > that specifies
>>> > read semantics in prose: which metadata to use to resolve the snapshot
>>> > (snapshot-log), how the two histories can diverge, and what to do when
>>> the
>>> > metadata is missing (raise an informative error). Codifying them would
>>> fill
>>> > a gap right next to something the spec already covers.
>>> >
>>> > Why this is now also a catalog concern, not just an engine concern?
>>> >
>>> > This used to be purely an engine-side matter. But the REST
>>> scan-planning
>>> > API has changed that. REST spec already exposes start-snapshot-id
>>> > ("Starting snapshot ID for an incremental scan (exclusive)") and
>>> > end-snapshot-id ("Ending snapshot ID for an incremental scan
>>> (inclusive).
>>> > Required when start-snapshot-id is specified."), and the planTableScan
>>> > description states that incremental scans are planned by passing both,
>>> and
>>> > that mixing incremental with point-in-time parameters is invalid.
>>> >
>>> > In other words, the REST spec already obligates a catalog to produce
>>> > an incremental
>>> > scan plan, but "correct" is defined only by whatever Java happens to
>>> do.
>>> > Two conformant catalogs could return different plans for the same
>>> request
>>> > and both claim REST compliance. It also means boundary conditions
>>> > (exclusive start / inclusive end) currently live only in the REST
>>> > documentation, not in the spec, so a non-REST engine has no shared
>>> source
>>> > of truth.
>>> >
>>> > The core principle for me is that the result of an incremental scan
>>> should not
>>> > depend on where it was planned. Whether an engine plans the scan
>>> > locally against
>>> > table metadata, or delegates to a catalog via planTableScan, it should
>>> get
>>> > the same set of files back. Today that's only true by coincidence
>>> because
>>> > every planner happens to copy Java. If the semantics aren't in the
>>> spec,
>>> > then server-side planning and client-side planning are free to
>>> diverge, and
>>> > the same query returns different data purely based on whether the
>>> catalog
>>> > implements the scan-plan API. That's a correctness property that
>>> shouldn't
>>> > be left to chance I think.
>>> >
>>> >
>>> > Some of the questions I think would need answering:
>>> >
>>> > If we do put this in the spec, these are the ambiguities I'd want
>>> pinned
>>> > down (drawn from Java's current behavior):
>>> >
>>> > - Range / traversal: is the scan the linear ancestry from
>>> end-snapshot-id back
>>> > to but excluding start-snapshot-id? What's required if start is not an
>>> ancestor
>>> > of end? (Java validates ancestry and errors if the start is not
>>> anancestor
>>> > / parent-ancestor of the end.)
>>> >
>>> > - Non-append operations in the range: Java today skips non-append
>>> snapshots in
>>> > the range -- appendsBetween keeps only snapshots whose operation is
>>> "append"
>>> > and ignores the rest (replace/overwrite/delete).
>>> >
>>> > - Boundary semantics: exclusive-start / inclusive-end is currently
>>> > only asserted
>>> > in the REST spec should it be normative in the spec more generally?
>>> >
>>> > - Which schema? Java projects the incremental scan result into the
>>> > current table
>>> > schema. But the range spans multiple snapshots, so "the snapshot
>>> schema"
>>> > isn't well-defined the way it is for a point-in-time scan.
>>> >
>>> > I don't have a strong predetermined answer on whether this belongs in
>>> the spec,
>>> > I mainly want to surface that it's currently
>>> implementation-defined,that
>>> > multiple implementations are already depending on it, and that the
>>> > REST scan-plan
>>> > API has quietly turned it into a conformance question for catalogs too.
>>> >
>>> >
>>> > Interested to hear what folks think.
>>> >
>>> > Thanks,
>>> >
>>> > Xander
>>> >
>>>
>>

Reply via email to