Hi all, Prashant and I synced up offline and decided to converge the Spec proposal onto this PR [1].
As suggested in the review comments, I ported over the optional request and response parameters from the other PR. I'd appreciate a review. Thank you! Sung [1] https://github.com/apache/iceberg/pull/18080 On 2026/09/12 19:33:59 Sung Yun wrote: > Hi all, > > At the 9/2 catalog sync it was mentioned that William has completed his > internship. I'm invested in continuing to look into pre-signed URL support in > Iceberg, since I think it unlocks further access delegation use cases (mainly > around FILE type references), so I gave it a go at seeing what the > implementation would look like if we followed the notes from the sync. > > Based on the sync, I made a few assumptions about the scope of an initial > implementation and put together [1], which only handles pre-signed URLs when > the server returns them in place of native locations in a plan response. I > agree with the point made on this thread that this workflow doesn't need a > spec change, and I think it's worth reviewing this scope as the candidate for > initial support. > > As I brought up during the sync, looking a bit further forward, I think it > would be helpful to start thinking about the use case where the client asks > the server to sign a list of URIs, whether that is a table's own files or > external file references. [2] introduces a minor change to the existing > `/sign` endpoint for this, adding `presigned-urls` as a value of > `X-Iceberg-Access-Delegation`, following Dan's suggestion to reuse the > header. I agree too that batching is a good idea, and I think following that > up after we settle on how we want to handle a single request would make sense. > > On my earlier point about a TTL in the response and refresh handling - I set > it aside for now in favor of simplicity. This PR assumes expiry is > communicated out of band by the catalog administrator, and the client uses a > URL until it fails. > > There are a few additional assumptions I made based on the sync discussion, > listed in the PR description. I would appreciate a review, especially on > whether those assumptions are the right scope for an initial implementation. > Prashant, since you mentioned at the sync that you were planning to look into > this as well, happy to coordinate however works best. > > [1] https://github.com/apache/iceberg/pull/18079 > [2] https://github.com/apache/iceberg/pull/18080 > [3] William's initial PR: https://github.com/apache/iceberg/pull/17457 > > Sung > > On 2026/09/01 18:01:09 Sung Yun wrote: > > Thanks William for working on this proposal. > > > > I actually see three scenarios, one being a variant of scenario (1) that > > Dan framed above: > > 1a. Where a scan/plan request returns presigned URLs > > 1b. Where scan/plan request returns unsigned URLs, and the FileIO contacts > > the signing service to get a presigned URL (variant of scenario 1) > > 2. Normal client-side planning mode, where the FileIO contacts the signing > > service to get a presigned URL > > > > On scenario (1a), I agree with Dan that the client-side work needs no spec > > change. > > > > In all of the scenarios above, I think accounting for a refresh workflow > > would be a good idea. I see pre-signed URLs slightly differently in the > > protocol from remote signing. In remote signing the signature is minted per > > request when FileIO is about to access the byte range and used immediately, > > whereas a pre-signed URL is reused across its lifetime and a very long scan > > could last longer than its TTL. In William's original proposal, there's a > > TTL included in the response of the endpoint which would make the proactive > > refresh on the client side simpler. The alternative to introducing a > > dedicated TTL/expiry field would be to parse the presigned URL and have > > each provider infer the TTL from the query parameter. > > > > Both could work, though parsing would mean each provider implementation > > handling a different encoding. If we were to use a dedicated expiry > > attribute, I hesitate with 1a given that it doesn't go through /sign, so > > expiry would need to be defined twice, once on the plan response and once > > on the signing response. That seems like a reason to treat 1b and 2 as the > > golden paths for generating presigned urls. Curious what others think. > > > > Regarding batching: agreed it's worth doing, and I'd sequence it after we > > settle on the single-object shape, since batching would simply reuse the > > same object shape. > > > > Sung > > > > On 2026/08/28 20:57:18 William Hyun wrote: > > > Hey Dan, > > > > > > The two scenarios you mentioned make sense to me and PR [1] has been > > > focused > > > on implementing the first scenario. Please take a look when you have the > > > chance. > > > Also I would like to second Prashant's point of scenario 1 not requiring > > > a spec change, > > > let us know your thoughts on this. > > > > > > For scenario 2 of using normal planning mode, I agree that we should > > > reuse the > > > general `/sign` endpoint with an additional spec change for adding a new > > > option to > > > access delegation like `X-Iceberg-Access-Delegation: pre-signed-urls`. > > > Additionally, as Prashant mentions above, considering the per-file nature > > > of the current > > > `/sign` endpoint, I believe this would be a good opportunity to discuss > > > the option of a > > > bulk-sign endpoint. WDYT? I can also provide a POC for this spec change. > > > > > > Thank you again and please take a look at [1]. > > > > > > Best, > > > William > > > > > > [1] https://github.com/apache/iceberg/pull/17457 > > > > > > On 2026/08/25 03:03:01 Prashant Singh wrote: > > > > Hey Dan, > > > > The 2 scenarios you mentioned make sense to me specially from the pov > > > > that > > > > we need an HTTPInputFile / HTTPInputStream and fileIO should use the > > > > relevant input file/stream based on detecting if the url is presigned, > > > > we > > > > have trying to achieve the same via [1], where these two ly in the > > > > *core* > > > > and is > > > > available to all the fileIO and are generic enough to be repurposed. We > > > > kept your feedback of integrating metrics especially in mind while > > > > implementing this [2], please do have a look when you get some time ! > > > > I feel for scenario 1, you mentioned, we would not require spec change, > > > > please let us know wdyt ? > > > > > > > > For scenario 2 i.e client reaching out to server to get pre-signed url > > > > on > > > > per file basis (assuming this because that what we have for remote > > > > signing), based on the feedbacks we got during this discussion, i > > > > wonder if > > > > it's the right time to think about the bulk sign api (which could return > > > > pre-signed / remote signed response based on access delegation > > > > requested) > > > > Since one of the concerns of the current remote signing was that it was > > > > requested was per file and it can cause server overwhelm for very huge > > > > tables, i know we have been parallely working on an effort to return > > > > back > > > > server issue signer configs in signer requests to potentially optimize > > > > authZ per request > > > > but a bulk signing api would be really helpful to mitigate those > > > > concerns. > > > > > > > > Looking forward to your feedback ! > > > > > > > > [1] https://github.com/apache/iceberg/pull/17457 > > > > [2] https://github.com/apache/iceberg/pull/17457#discussion_r3696073044 > > > > > > > > Thanks, > > > > Prashant Singh > > > > > > > > On Mon, Aug 17, 2026 at 1:35 PM Daniel Weeks <[email protected]> wrote: > > > > > > > > > Hey William, > > > > > > > > > > I had a chance to sync with Russell on this topic and I agree that > > > > > Azure > > > > > has some unique differences that make the current remote signing > > > > > difficult. > > > > > > > > > > In terms of how to best support the pre-signing models, I believe > > > > > that we > > > > > can actually accommodate that implementation under the current remote > > > > > signing APIs. > > > > > > > > > > Back when we designed the remote signing APIs, we considered the case > > > > > where the signer wanted to return a different path than what was > > > > > originally > > > > > requested (primarily for cases like routing/redirection). However, > > > > > that > > > > > also allows for the catalog to return a presigned URI instead of the > > > > > originally requested path. The endpoint itself is general (just > > > > > `/sign`) > > > > > and adding a new option to access delegation would be a good signal > > > > > for > > > > > what a client is requesting. > > > > > > > > > > In terms of use cases, there are two different scenarios: > > > > > 1. Where a scan/plan request returns presigned URLs: in this case the > > > > > FileIO should detect the presigned url and execute it directly > > > > > instead of > > > > > using the SDK client. > > > > > 2. The normal planning mode, where the FileIO should contact the > > > > > signing > > > > > service instead of going directly through the SDK to get a presigned > > > > > URL. > > > > > > > > > > I believe both of these scenarios can be achieved by introducing a > > > > > common > > > > > BaseInputFile/BaseInputStream that shares the identification or > > > > > redirection > > > > > logic for presigning. > > > > > > > > > > Does this make sense? I'll also follow up on the PR, but I think > > > > > there's > > > > > very little we need to do in order to accommodate both paths. > > > > > > > > > > -Dan > > > > > > > > > > > > > > > On Fri, Aug 14, 2026 at 1:43 PM William Hyun <[email protected]> > > > > > wrote: > > > > > > > > > >> Here is a POC for following along with this claim about Azure not > > > > >> supporting S3-style remote signing natively. > > > > >> - https://github.com/williamhyun/Azure-SAS-Header-POC > > > > >> > > > > >> On 2026/07/21 20:01:16 William Hyun wrote: > > > > >> > Hi All, > > > > >> > > > > > >> > Following up on the ongoing discussion thread regarding the > > > > >> > file-level > > > > >> access delegation modes in the Iceberg REST spec, I wanted to share > > > > >> some > > > > >> findings from a cross-cloud perspective of how the major Cloud > > > > >> Service > > > > >> Providers handle these protocols. > > > > >> > > > > > >> > When evaluating remote signing as an access mode, two major > > > > >> > operational > > > > >> hurdles emerge: strict 15-minute Time-To-Live constraints and a lack > > > > >> of > > > > >> native support across all CSPs, particularly Azure. > > > > >> > > > > > >> > 1. Non-Configurable TTLs > > > > >> > AWS: In the AWS SigV4 protocol, the calculated signature with a > > > > >> timestamp inherently expires 15 minutes after. [1] > > > > >> > GCS: GCS enforces an identical constraint for remote header > > > > >> > signing. A > > > > >> signature passed via request headers is valid only from 15 minutes > > > > >> prior to > > > > >> the active datetime until exactly 15 minutes after. [2] > > > > >> > > > > > >> > 2. Lack of Native Remote Signing Support in Azure > > > > >> > While AWS (SigV4) and GCS (V4) natively accept short-lived, > > > > >> > granular > > > > >> signatures via the Authorization header, Microsoft Azure Blob > > > > >> Storage and > > > > >> ADLS Gen2 fundamentally reject this paradigm. > > > > >> > > > > > >> > In Azure Storage, a Shared Access Signature (SAS) is strictly a > > > > >> > token > > > > >> designed to be appended to the resource URI as query parameters. [3] > > > > >> > To execute true remote header signing for Azure Storage, a catalog > > > > >> would be forced to use the legacy SharedKey authorization scheme, > > > > >> formatting the header as Authorization: SharedKey > > > > >> <AccountName>:<Signature>. Constructing this signature mathematically > > > > >> requires the catalog to hold the storage account's master access > > > > >> key. [4] > > > > >> [5] > > > > >> > Because of this, Microsoft explicitly recommends disallowing > > > > >> > SharedKey > > > > >> authorization entirely, advising modern architectures to rely > > > > >> instead on > > > > >> User Delegation SAS tokens (via the URI) or Microsoft Entra ID. [5] > > > > >> > > > > > >> > On the other hand, pre-signed URLs are natively and uniformly > > > > >> > supported > > > > >> across all three major CSPs. While the underlying cryptographic > > > > >> signatures > > > > >> and generation methods differ (AWS SigV4, GCS V4, Azure SAS), the > > > > >> resulting > > > > >> access artifact is universally a standard, fetchable HTTPS URL. > > > > >> Furthermore, pre-signed URLs allow for highly customizable expiration > > > > >> times. The maximum TTL is upper bound by the validity of the signing > > > > >> credential itself (e.g. temporary STS credentials which are default > > > > >> 1-hour > > > > >> sessions but can be configured up to 12 hours). This provides > > > > >> significantly > > > > >> more operational flexibility than the rigid 15-minute window > > > > >> enforced on > > > > >> header-based signatures. > > > > >> > > > > > >> > Given the findings above, I would still like to discuss the > > > > >> > possibility > > > > >> of pre-signed URLs. > > > > >> > During our last sync, a valid concern was raised: our current SDKs > > > > >> > and > > > > >> FileIO implementations don't support processing raw, pre-signed > > > > >> HTTPS URLs. > > > > >> Thus requiring a new, separate read path to be built. To help us > > > > >> evaluate > > > > >> exactly what this effort entails and to discuss the trade-offs, I’ve > > > > >> put > > > > >> together a POC for what this pre-signed URL read path might look > > > > >> like: > > > > >> > https://github.com/williamhyun/iceberg/pull/206 > > > > >> > > > > > >> > Looking forward to hearing your thoughts! > > > > >> > > > > > >> > Best, > > > > >> > William > > > > >> > > > > > >> > [1] > > > > >> https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-api-request-authentication.html#sqs-api-basic-authentication-process > > > > >> > [2] > > > > >> https://docs.cloud.google.com/storage/docs/authentication/signatures#active_datetime > > > > >> > [3] > > > > >> https://learn.microsoft.com/en-us/azure/storage/common/storage-sas-overview#how-a-shared-access-signature-works > > > > >> > [4] > > > > >> https://learn.microsoft.com/en-us/rest/api/storageservices/authorize-with-shared-key#specifying-the-authorization-header > > > > >> > [5] > > > > >> https://learn.microsoft.com/en-us/azure/storage/common/authorize-data-access?tabs=blobs#authorization-for-data-operations > > > > >> > > > > > >> > > > > > >> > On 2026/07/08 01:16:35 William Hyun wrote: > > > > >> > > Hi all, > > > > >> > > > > > > >> > > Following up on the earlier discussion around file-level access > > > > >> delegation > > > > >> > > in the REST Catalog spec, I put together a small POC PR against > > > > >> > > my > > > > >> personal > > > > >> > > Iceberg fork to explore the remote signing + caching direction we > > > > >> discussed: > > > > >> > > > > > > >> > > https://github.com/williamhyun/iceberg/pull/198 > > > > >> > > > > > > >> > > The goal is to make the tradeoffs more concrete and evaluate > > > > >> > > whether > > > > >> this > > > > >> > > could be a practical approach for finer-grained delegated access. > > > > >> > > Feedback is very welcome. > > > > >> > > > > > > >> > > Best, > > > > >> > > William > > > > >> > > > > > > >> > > > > > > >> > > > > > > >> > > On Thu, Jul 2, 2026 at 5:29 PM William Hyun <[email protected]> > > > > >> wrote: > > > > >> > > > > > > >> > > > Hi all, > > > > >> > > > > > > > >> > > > Thank you again for the discussion during the 7/1 catalog > > > > >> > > > community > > > > >> sync. > > > > >> > > > I wanted to summarize two alternative directions that came up > > > > >> > > > to > > > > >> > > > pre-signed URLs: > > > > >> > > > 1. scoped credential vending > > > > >> > > > 2. remote signing with caching and bulk-signing > > > > >> > > > > > > > >> > > > I think scoped credential vending is worth considering in cases > > > > >> where the > > > > >> > > > authorization boundary maps cleanly to a small number of stable > > > > >> storage > > > > >> > > > prefixes. > > > > >> > > > Scoped credential vending seems less practical when access > > > > >> > > > needs to > > > > >> be > > > > >> > > > granted for many specific files rather than a small number of > > > > >> prefixes. > > > > >> > > > > > > > >> > > > In AWS, session policies are passed as parameters during > > > > >> AssumeRole-style > > > > >> > > > issuance when the temporary session is created. [1] > > > > >> > > > This means that scoping down an already-issued credential > > > > >> > > > cannot be > > > > >> done > > > > >> > > > locally by simply altering the existing token. > > > > >> > > > Instead, it requires minting another STS session whose policy > > > > >> encodes the > > > > >> > > > new scope. > > > > >> > > > This turns authorization into an STS issuance problem with > > > > >> limitations on > > > > >> > > > STS request frequency [2] and policy size. [3] > > > > >> > > > > > > > >> > > > By contrast, pre-signed URLs and remote signing can authorize > > > > >> > > > exact > > > > >> object > > > > >> > > > requests using already-held credentials, without creating a > > > > >> > > > new STS > > > > >> session > > > > >> > > > for each file set. > > > > >> > > > Based on that discussion, I am exploring a POC around bulk > > > > >> > > > remote > > > > >> > > > signing + caching to evaluate whether it could be a practical > > > > >> alternative > > > > >> > > > to pre-signed URLs for these finer-grained access patterns. > > > > >> > > > > > > > >> > > > Best, > > > > >> > > > William > > > > >> > > > > > > > >> > > > [1] > > > > >> https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies.html > > > > >> > > > [2] > > > > >> > > > > > > > >> https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_iam-quotas.html#reference_iam-quotas-sts-requests > > > > >> > > > [3] > > > > >> > > > > > > > >> https://docs.aws.amazon.com/STS/latest/APIReference/API_AssumeRole.html > > > > >> > > > > > > > >> > > > On Tue, Jun 30, 2026 at 7:43 PM William Hyun > > > > >> > > > <[email protected]> > > > > >> wrote: > > > > >> > > > > > > > >> > > >> Hi everyone, > > > > >> > > >> > > > > >> > > >> I'd like to open a discussion on an extension to the REST > > > > >> > > >> Catalog > > > > >> spec to > > > > >> > > >> enable file-level access delegation during scan planning. > > > > >> > > >> Today, delegated access is table-scoped, which can force > > > > >> > > >> over-provisioning or table fragmentation when consumers should > > > > >> only see a > > > > >> > > >> subset of partitions or records. > > > > >> > > >> This proposal adds a path using pre-signed URLs to make > > > > >> partition-scoped > > > > >> > > >> sharing practical in the REST Catalog model. > > > > >> > > >> > > > > >> > > >> Here is my proposal: https://s.apache.org/n16st > > > > >> > > >> > > > > >> > > >> I look forward to hearing your thoughts. > > > > >> > > >> > > > > >> > > >> Relevant discussion: > > > > >> > > >> - > > > > >> > > >> https://lists.apache.org/thread/ko9kp0gvzhx85n7cvoxqnpw4vwnhmdg6 > > > > >> > > >> > > > > >> > > >> Best, > > > > >> > > >> William > > > > >> > > >> > > > > >> > > >> > > > > >> > > > > > > >> > > > > > >> > > > > > > > > > > > > > > >
