Thanks, Russell and Daniel. I will wait for the authors of the original PR to be back from the vacation and we can consolidate there. I'll leave a comment on the FileRef proposal as well.
On Tue, Jul 28, 2026 at 2:38 PM Daniel Weeks <[email protected]> wrote: > I agree with Russell here. There's already a proposal (though it needs to > be updated), so I think we should consolidate around updating that as > opposed to starting a new effort around this. > > Original thread is here: > https://lists.apache.org/thread/kz09b2rj7c00j6z2vlqg8v5myh94bgl5 > > -Dan > > On Tue, Jul 28, 2026 at 11:41 AM Russell Spitzer < > [email protected]> wrote: > >> Hopefully bumping this out of Spam. I know Talat and company have a >> proposal as well and it would be great to see some consensus amongst these >> proposals but it seems like everyone is on vacation at the moment. >> >> >> On Sun, Jul 26, 2026 at 10:21 AM Nitya Kumar Sharma < >> [email protected]> wrote: >> >>> Hi all, >>> >>> >>> Now that the Parquet FILE logical type has merged (parquet-format#585 >>> [1]), I'd like to share a proposal for a corresponding file data type >>> in Iceberg — >>> https://docs.google.com/document/d/1W25ztRqm2xXBxOzGtUHK_mv-FLviQeIcKnUcmvCKOuo/edit >>> — and start a discussion focused on the type-level semantics. >>> >>> >>> A note on scope and framing: this document deliberately covers only the >>> data type (the value) — how file lives in the type system, schema >>> serialization/evolution, statistics, and behavior under >>> transforms/deletes/time-travel. It leaves credential vending and brokered >>> access to the catalog. >>> >>> >>> I want to be explicit that this is not a counter-proposal to the >>> existing document from Talat Uyarer, Alex Stephen, and Sung Yun [2], which >>> covers both the data type and catalog-brokered credential vending. I've >>> scoped this one narrowly on purpose, so we can converge on the type >>> semantics independently of the access-control layer — which can sit on top >>> unchanged, including their REST/credential-vending design. My intent is to >>> keep the discussion moving and focused on the file data type, and I'm >>> happy to collaborate and fold this into a shared direction. >>> >>> >>> A few Iceberg-specific decisions in the doc, for discussion: >>> >>> >>> >>> - >>> >>> Representation: file is a "named struct" — it serializes with a >>> "file" discriminator and a fixed, closed set of six leaves (uri, >>> offset, size, content_type, checksum, inline), each assigned a >>> normal field ID. Field semantics, byte resolution, and checksum format >>> defer to #585. >>> - >>> >>> Two forms only (inline and external); no in-file/self-reference: >>> #585 permits packing bytes within the same data file and addressing them >>> by >>> offset/size with no uri. This proposal does not adopt that in-file >>> form — only inline bytes and external uri references are allowed — >>> keeping the type to two well-understood forms while leaving room to add >>> it >>> later. >>> - >>> >>> Relative URIs resolve against the table's base location (v4 >>> relative-location rules), so tables stay relocatable; #585 permits >>> relative >>> URIs but leaves "relative to what" undefined. >>> - >>> >>> Statistics use ordinary field-ID-keyed bounds per leaf (bounds on uri >>> enable object-prefix file skipping) — no new bounds structure, in >>> contrast >>> to variant. >>> - >>> >>> Not comparable as a whole: void transform only, and file can't be a >>> partition/sort or equality-delete key. >>> - >>> >>> Credentials/access are out of scope for the type (a catalog concern). >>> >>> >>> Feedback very welcome — especially on the representation (the "named >>> struct" question). >>> >>> >>> Thanks, Nitya >>> >>> >>> [1] https://github.com/apache/parquet-format/pull/585 >>> >>> [2] >>> https://docs.google.com/document/d/1BdA4G-38GoVcpX6UY4UI-QA1MvrDnH6uSKeCJBadILQ/edit >>> >>> >>>
