Thanks for the reply, is there anywhere I can follow that work? Which community sync is best to attend for that etc?
Best, Xander On Fri, 24 Jul 2026 at 18:01, Ryan Blue <[email protected]> wrote: > Quick note: we are planning changes in v4 that tighten requirements for > change detection across versions. That way the spec will support all types > of incremental scans without relying on conventions in a particular > implementation. > > On Thu, Jul 23, 2026 at 4:52 PM Xiening Dai <[email protected]> wrote: > >> Based on the current design of the IRC incremental scan API (which >> returns a set of files), it can only support append only case. But to me, >> incremental scan is really scanning the delta between two snapshots, so it >> should also include deleted and updated tuples. This would become very >> powerful building block for scenarios like MV refresh. But to support that >> it would have to return tuple set, instead of file set, and that would >> depart from the norms of IRC APIs. >> >> So I think at this point, we would need to clarify that `plan` API only >> support append-only case for incremental scan. For true data scan, it needs >> a good discussion on how to support that from the API/spec level. >> >> On 2026/07/23 22:06:59 Alexander Bailey wrote: >> > Hi all, >> > >> > I'd like to open a discussion about whether incremental scan semantics >> > (specifically >> > incremental append scans) should be codified in the table spec, and if >> so, >> > what the contract should be. >> > >> > The gap: >> > >> > Today the spec is effectively silent on incremental scans. In >> > practice, the Java >> > IncrementalAppendScan is the de facto reference implementation, and >> everyone >> > else is following it: >> > >> > - PyIceberg -- implemented by following Java >> > github.com/apache/iceberg-python/pull/3512 >> > >> > - iceberg-cpp -- incremental scan API following Java: >> > github.com/apache/iceberg-cpp/pull/559 >> > >> > - iceberg-rust -- I have a PR up that, likewise, just mirrors Java >> > >> > github.com/apache/iceberg-rust/pull/2337 >> > >> > >> > So we have four implementations converging on behavior that lives only >> in one >> > codebase. That's fragile: any of Java's incidental choices become the >> unwritten >> > contract, and there's no spec to point to when implementations disagree. >> > >> > There's a good precedent for what the fix could look like. The spec >> already has >> > a "Point-in-Time Reads (Time Travel)" section (iceberg.apache.org/spec# >> … >> > <https://iceberg.apache.org/spec/#point-in-time-reads-time-travel>) >> > that specifies >> > read semantics in prose: which metadata to use to resolve the snapshot >> > (snapshot-log), how the two histories can diverge, and what to do when >> the >> > metadata is missing (raise an informative error). Codifying them would >> fill >> > a gap right next to something the spec already covers. >> > >> > Why this is now also a catalog concern, not just an engine concern? >> > >> > This used to be purely an engine-side matter. But the REST scan-planning >> > API has changed that. REST spec already exposes start-snapshot-id >> > ("Starting snapshot ID for an incremental scan (exclusive)") and >> > end-snapshot-id ("Ending snapshot ID for an incremental scan >> (inclusive). >> > Required when start-snapshot-id is specified."), and the planTableScan >> > description states that incremental scans are planned by passing both, >> and >> > that mixing incremental with point-in-time parameters is invalid. >> > >> > In other words, the REST spec already obligates a catalog to produce >> > an incremental >> > scan plan, but "correct" is defined only by whatever Java happens to do. >> > Two conformant catalogs could return different plans for the same >> request >> > and both claim REST compliance. It also means boundary conditions >> > (exclusive start / inclusive end) currently live only in the REST >> > documentation, not in the spec, so a non-REST engine has no shared >> source >> > of truth. >> > >> > The core principle for me is that the result of an incremental scan >> should not >> > depend on where it was planned. Whether an engine plans the scan >> > locally against >> > table metadata, or delegates to a catalog via planTableScan, it should >> get >> > the same set of files back. Today that's only true by coincidence >> because >> > every planner happens to copy Java. If the semantics aren't in the spec, >> > then server-side planning and client-side planning are free to diverge, >> and >> > the same query returns different data purely based on whether the >> catalog >> > implements the scan-plan API. That's a correctness property that >> shouldn't >> > be left to chance I think. >> > >> > >> > Some of the questions I think would need answering: >> > >> > If we do put this in the spec, these are the ambiguities I'd want pinned >> > down (drawn from Java's current behavior): >> > >> > - Range / traversal: is the scan the linear ancestry from >> end-snapshot-id back >> > to but excluding start-snapshot-id? What's required if start is not an >> ancestor >> > of end? (Java validates ancestry and errors if the start is not >> anancestor >> > / parent-ancestor of the end.) >> > >> > - Non-append operations in the range: Java today skips non-append >> snapshots in >> > the range -- appendsBetween keeps only snapshots whose operation is >> "append" >> > and ignores the rest (replace/overwrite/delete). >> > >> > - Boundary semantics: exclusive-start / inclusive-end is currently >> > only asserted >> > in the REST spec should it be normative in the spec more generally? >> > >> > - Which schema? Java projects the incremental scan result into the >> > current table >> > schema. But the range spans multiple snapshots, so "the snapshot schema" >> > isn't well-defined the way it is for a point-in-time scan. >> > >> > I don't have a strong predetermined answer on whether this belongs in >> the spec, >> > I mainly want to surface that it's currently implementation-defined,that >> > multiple implementations are already depending on it, and that the >> > REST scan-plan >> > API has quietly turned it into a conformance question for catalogs too. >> > >> > >> > Interested to hear what folks think. >> > >> > Thanks, >> > >> > Xander >> > >> >
