*1. Can a URI be relative in an Iceberg Table?* *2. On read, should Iceberg mandate converting that persisted relative value to an absolute URI, or should it represent the value exactly as read?* *3. How should maintenance operations operate by default when a file column is present?* I have a very similar opinion and reasoning as Dan's for question 1-3. So I'll avoid repeating it
Before answering 4, I think it would help to define the terms more precisely to avoid confusion: - Access control: whether a client can obtain credentials that allow it to access a referenced file - Ownership: whether the referenced file participates in the lifecycle of the Iceberg table, e.g. whether table-managed operations may move, rewrite, or delete it. Imo, access control should be controlled by the catalog. inferring credentials from the uri/column/table level will leave too many corner cases uncovered. *For ownership, I think it has value to have a field owned to express the ownership in the Iceberg file type due for the following reasons:* 1) Ownership is a lifecycle intent and should travel with the file reference. A persisted ownership intent in a field gives a consistent intent to engines and catalogs. 2) owned=true only expresses lifecycle intent. The engine still decides how an operation is executed, and the catalog decides whether access is allowed. 3) Looking forward, having an ownership field enables generic table operations like compacting smaller objects into a larger one and referencing each one via (uri, offset, size) Whether *owned* should be a simple boolean flag or something more extensible is a different topic and can be discussed later. Best, Shawn On Thu, Sep 3, 2026 at 8:55 AM Daniel Weeks <[email protected]> wrote: > I have a slightly different view on this now that I've thought about it > for a little while. > > I think Iceberg's position should be that the library faithfully respects > whatever is provided to it. If a user/engine provides a value to the core > library, it should not be manipulated as that is considered data. I'm not > aware of any other place where we would manipulate or normalize values > coming from a user. This also aligns with Parquet in that we're not > changing/narrowing the definition from the Parquet spec. > > From that, the answers to the above questions follow: > 1. Yes, they can be relative > 2. Iceberg should produce the value it was provided. We can provide > guidance on how engines should expand relative paths. > 3. I believe the safest option is to default to not cleaning up File > references. Changing behavior based on the location or path will always > leave some scenario where data is improperly deleted. That's not a > recoverable state, so it should be avoided. Referenced data cleanup should > be an explicit opt-in (though this is often handled by the catalog). > 4. No, we should leave ownership and enforcement up to the Catalog. Since > most permissions are delegated from the catalog via vended > credentials/signing, the catalog should play a role in ensuring data that > should not be deleted is protected. > > I do think strong guidance for engines on how to relativize or resolve > paths will be necessary, but I'm pretty convinced right now that File info > is data and we shouldn't be changing what is provided to the core library. > I also don't want to close the door on relative paths since that would be > incompatible with the objectives we set forth in allowing relative paths in > metadata. > > -Dan > > On Wed, Sep 2, 2026 at 6:50 PM <[email protected]> wrote: > >> >> Hi Russell, >> >> Thank you for summarizing and starting the thread. My initial thinking >> after listening to the discussion this morning, is that these references >> should be user defined and Iceberg should manage the references themselves >> only not the data they point to. >> >> For point 1. I think the references should be treated, at least >> initially, as absolute in terms of how Iceberg treats them. If a user wants >> to pass relative or something like that “s3://bucket/some-prefix-*” it’s >> their data. The spec should be written where there’s no expectation of >> expanding that URI to an actual object or objects path. >> >> For point 2. I think Iceberg should return the URIs exactly the same as >> they were written. It opens Iceberg to potential correctness bugs if there >> was resolution being done at Read time. >> >> For point 3. I am in favor of the model that Iceberg owns the column’s >> data, but not where the reference points. So for example if the URI pointed >> at an image, maintenance can run on the URI just never the image. Off the >> top of my head this allows duplicate URIs, objects in different buckets or >> even providers and it affords Iceberg the full suite of Maintenance >> activities without special handling. >> >> Otherwise if Iceberg owned the object being pointed to, then we’ll have >> to define different types of objects the FILE type can point to and how to >> handle each. Which at a future date we may want, but I’m initially hesitant >> to agree with because it increases the scope of the feature vs just owning >> the URIs and metadata as defined in Parquet. >> >> 4. I’m currently in favor of pushing this to a later iteration and just >> have a single mode of ownership that I tried to describe above. >> >> Best Regards, >> Kurtis >> >> On Sep 2, 2026, at 12:11, Russell Spitzer <[email protected]> >> wrote: >> >> >> Hi y'all! >> >> Obviously we just had a long discussion on the behavior of a URI within >> an Iceberg's File Type and I wanted to make sure we continue the discussion >> and try to continue making progress before we have another follow up >> meeting. >> >> The TLDR for those who weren't attending: we are debating the File Type >> Proposal >> <https://docs.google.com/document/d/1BdA4G-38GoVcpX6UY4UI-QA1MvrDnH6uSKeCJBadILQ/edit?tab=t.pnp3scifr79p#heading=h.v4a92wm7mz0o> >> and >> the corresponding spec PR <https://github.com/apache/iceberg/pull/17918> >> >> >> The key issue we want to continue discussing UDI subfield of the File >> type. >> >> In parquet this field can be persisted as either a relative or absolute >> URI. Some community members are worried about having both representations >> which led to today's discussion. >> >> I want to open up the floor here for folks who have opinions on the >> behavior here. >> >> So the major questions (and my responses) are: >> >> >> *1. Can a URI be relative in an Iceberg Table?* >> My response to this is yes. I don't think we gain much by forbidding it, >> and implementation wise I don't think we can reasonably stop it. Users can >> apply a constraint if they need this to be true. >> >> >> *2. On read, should Iceberg mandate converting that persisted relative >> value to an absolute URI, or should it represent the value exactly as read?* >> I think every reader should return the value as is. I can't stop an >> engine from showing the user a different representation, but the bytes >> should be carried through as is in the file imho. >> >> >> *3. How should maintenance operations operate by default when a file >> column is present?* >> By default, if we lack a mechanism for indicating ownership, we should >> assume all path references are not owned by the table. >> >> >> *4. Should the Spec or the implementation notes specify how tables should >> express ownership of the files they reference?* >> I think this is probably something useful that we should lay out, but I >> don't think it needs to be decided for this proposal. >> >> >> Please follow on if you have any other questions for debate or comments >> on these topics. >> >>
