I think these arguments are compelling arguments from my perspective.

Did you get as far as creating a draft implementation for this change on
the Java side? Would be keen to see how this might look for rust also.

Best,
Xander

On Thu, 24 Sept 2026 at 13:39, Gábor Kaszab <[email protected]> wrote:

> I get your point, and agree that whichever design we go with, we can
> implement the 1-1 mapping between encryption keys and encrypted files.
>
> If we look for the long term, keeping the existing 'encryption-keys' based
> approach might cause some headaches:
> - When adding a new file type that is encrypted this way, then we have to
> remember to implement the encryption key cleanup for that file type too.
> E.g. for indexes in the future when we dereference an index file we have to
> implement a trigger to also clean up the relevant encryption key. With the
> 'in-place encryption key' it's not an issue.
> - The length of the 'encryption-keys' list matters. We have to find stuff
> there by ID, so even if we build a map from the list internally, there is a
> cost somewhere that grows with the size of the list. Now V3 manifest list
> encryption keys are there, we can't change that, but I would be careful
> adding keys for stat files, V4 root manifests, etc. into the list.
>
> As I understood, the motivation for the 'encryption-keys' list was to
> implement a way to reuse encryption keys across files. If reusal is not an
> option, the list lost its purpose.
>
> Any feedback is appreciated!
> Gabor
>
> Xander Bailey <[email protected]> ezt írta (időpont: 2026. szept. 24.,
> Cs, 12:56):
>
>> Hi Gabor,
>>
>> I understand the cleanup benefit of storing the encrypted key metadata
>> alongside the file reference. I’m still weighing that against following the
>> existing manifest-list pattern.
>>
>>
>> With a fresh DEK and a dedicated encryption-keys entry for each
>> statistics file, the existing pattern would (does?) also support the
>> one-file-per-DEK model. The additional responsibility would be removing
>> that entry when the statistics metadata is removed or replaced, but this is
>> a lifecycle problem we already need to handle for manifest lists. We would
>> still need to manage shared KEK lifetimes in either design.
>>
>>
>> I agree on generating fresh DEKs per file, and your proposal makes
>> ownership of the key metadata more explicit. My hesitation is that
>> statistics and manifest lists would then use different representations for
>> essentially the same operation. Following the existing pattern would let
>> implementations extend the same key lookup and lifecycle handling to
>> statistics.
>>
>>
>> Do you think extending that handling to statistics would be significantly
>> more complicated than supporting the inline representation alongside the
>> existing manifest-list approach? I can see the cleanup advantage in
>> isolation, but I’m less convinced that it simplifies implementations
>> overall.
>>
>>
>> Thanks,
>> Xander
>>
>> On Thu, 24 Sept 2026 at 10:28, Gábor Kaszab <[email protected]>
>> wrote:
>>
>>> Thanks for the answers Gidon and Xander!
>>>
>>> This is how I think, a new design could look like for statistics files
>>> in terms of the spec: https://github.com/apache/iceberg/pull/17533
>>> Here the encryption key is stored directly in the metadata for the file
>>> it encrypts, meaning there is 1-1 mapping between the encryption key and
>>> the encrypted files. KEKs remain in the `encryption-keys` list and we still
>>> refer to them by `key-id`.
>>> The same design could be used for V4 root manifests, but let's conclude
>>> on this first. Let me know if this makes sense!
>>>
>>> About "retaining the existing key-id / encryption-keys pattern": I think
>>> this is unwanted complexity, especially when it comes to cleaning up
>>> encryption keys. With the proposed design cleanup is coming out of the box:
>>> when a snapshot is gone, the root manifest's encryption key is also gone,
>>> no need to look elsewhere. Same for other file types like stats.
>>>
>>> Please share your opinions!
>>> Gabor
>>>
>>>
>>> Xander Bailey <[email protected]> ezt írta (időpont: 2026. szept.
>>> 23., Sze, 16:01):
>>>
>>>> Hi Gidon, Gabor,
>>>>
>>>> I agree that we should generate a fresh DEK for each file - in fact I
>>>> would personally want this to be made explicit. This came up when reviewing
>>>> https://github.com/apache/iceberg/pull/16353
>>>>
>>>> My initial preference reading this was to retain the existing key-id /
>>>> encryption-keys pattern so implementations could use the same approach
>>>> across these file types. The point about the metadata being specific to the
>>>> file, and simplifying its cleanup, makes sense though.
>>>>
>>>> I’d like us to preserve a common encrypted key metadata structure and
>>>> wrapping/unwrapping model as we make this change. We should be able to
>>>> share that implementation whether the metadata is resolved through a V3 key
>>>> ID or obtained directly from a V4 per-file structure.
>>>>
>>>> With that, I can see the benefit of keeping the encrypted DEK metadata
>>>> alongside the file reference and retaining shared KEKs in
>>>> encryption-keys. The per-file DEK entry would no longer need its own
>>>> key-id, but would still need to identify the KEK used to wrap it.
>>>>
>>>> Thanks,
>>>> Xander
>>>>
>>>>
>>>> On Wed, 23 Sept 2026 at 14:42, Gidon Gershinsky <[email protected]>
>>>> wrote:
>>>>
>>>>> Hi Gabor,
>>>>>
>>>>> I agree that in V4, the DEK metadata is best moved into per-file
>>>>> structures (such as snapshots for manifest list files). This will make it
>>>>> easier to manage/clean-up these keys. The `key-d` field indeed
>>>>> becomes obsolete then.
>>>>>
>>>>> The key encryption keys would stay as they are in the shared
>>>>> `encryption keys` list in TableMetadata, since these keys are shared 
>>>>> across
>>>>> the table.
>>>>>
>>>>> Technically, it is possible to re-use DEKs for multiple files, but
>>>>> this requires careful management to prevent breaking the AES GCM cipher.
>>>>> The simple and practical solution is to generate a random DEK per file.
>>>>> Moreover, Iceberg key metadata includes a unique file ID and, in some
>>>>> cases, the file length - so even if a key can be reused, the key metadata
>>>>> cannot.
>>>>>
>>>>> (as for the other technical details - the nonces/IVs are not visible
>>>>> at this level, they are an internal part of PME
>>>>> <https://parquet.apache.org/docs/file-format/data-pages/encryption/>
>>>>> and GCM Stream <https://iceberg.apache.org/gcm-stream-spec/>
>>>>> mechanisms)
>>>>>
>>>>> Cheers, Gidon
>>>>>
>>>>>
>>>>> On Wed, Sep 23, 2026 at 3:18 PM Gábor Kaszab <[email protected]>
>>>>> wrote:
>>>>>
>>>>>> Hey Iceberg Community,
>>>>>>
>>>>>>
>>>>>>
>>>>>> I've been recently involved in conversations around encryption keys
>>>>>> and the current design of how we encrypt manifest list files with the
>>>>>> purpose of extending the design to additional file types like statistics
>>>>>> files and V4 root manifest files.
>>>>>>
>>>>>> I think there are some assumptions with the current design that we
>>>>>> should revisit now.
>>>>>>
>>>>>>
>>>>>>
>>>>>> *Context*
>>>>>>
>>>>>>    1. Manifest list file encryption
>>>>>>
>>>>>> The `encryption-keys` list in table metadata (spec
>>>>>> <https://iceberg.apache.org/spec/#table-metadata-fields>) contains:
>>>>>>
>>>>>>    - Encryption keys used for encrypting manifest list files (DEK)
>>>>>>    - Key encryption keys (KEK) used for encrypting DEKs.
>>>>>>
>>>>>> The format for them is described here
>>>>>> <https://iceberg.apache.org/spec/#encryption-keys>. Each of them has
>>>>>> a `key-id` and for manifest list files we store a `key-id` in the 
>>>>>> snapshot
>>>>>> to refer to a DEK in the list. The DEK in turn refers to the KEK used for
>>>>>> encrypting that particular DEK.
>>>>>>
>>>>>>    2. Manifest, data and delete file encryption
>>>>>>
>>>>>> The raw encryption keys are stored directly for these files as
>>>>>> `key-metadata`. See this
>>>>>> <https://iceberg.apache.org/spec/#manifest-lists> or this
>>>>>> <https://iceberg.apache.org/spec/#data-file-fields>.
>>>>>>
>>>>>>
>>>>>>
>>>>>> We'll focus on 1) now.
>>>>>>
>>>>>>
>>>>>>
>>>>>> *Assumption*
>>>>>>
>>>>>> The assumption is that a particular DEK in the list can be reused for
>>>>>> encrypting multiple files (manifest list files now, other files also 
>>>>>> later
>>>>>> on) reducing the space required for storing the DEKs.
>>>>>>
>>>>>>
>>>>>>
>>>>>> *Reality check*
>>>>>>
>>>>>> While I'm not an expert of this area, I did some research, and I
>>>>>> think while it's theoretically feasible, in practice it's overly
>>>>>> complicated to implement such a DEK sharing approach. Here is what I 
>>>>>> found:
>>>>>>
>>>>>>    - It's not just a DEK what we need for encryption but other
>>>>>>    auxiliary, generated information like AAD prefix, nonce and other 
>>>>>> auxiliary
>>>>>>    information baked into `key-metadata`
>>>>>>    - It's a cryptographical requirement that if DEK is reused then
>>>>>>    the some of the generated auxiliary information MUST differ. More
>>>>>>    particular:
>>>>>>    - Some source says reusing the same DEK + nonce pair for multiple
>>>>>>    files is "Catastrophic" Link
>>>>>>    
>>>>>> <https://neilmadden.blog/2024/05/23/galois-counter-mode-and-random-nonces/>
>>>>>>    - Even if nonce is generated, there is a theoretical chance of
>>>>>>    reusing the same for the same DEK
>>>>>>    - Implementation-wise a tracking information is required to keep
>>>>>>    which nones are used for each DEK. This information have to be kept 
>>>>>> in a
>>>>>>    persisted way that we don't loose after a restart. Multi-writer 
>>>>>> scenarios
>>>>>>    would make keeping this information even harder
>>>>>>    - It's common practice to use a different DEK for different files
>>>>>>
>>>>>>
>>>>>>
>>>>>> Based on the above, I'd be against reusing such keys across multiple
>>>>>> files.
>>>>>>
>>>>>>
>>>>>>
>>>>>> *Way forward*
>>>>>>
>>>>>> For V3 manifest lists we already use the approach with a `key-id` in
>>>>>> snapshot referring to an encrypted key metadata in `encryption-keys` 
>>>>>> list.
>>>>>> This has been out there and I don’t think we can change this now.
>>>>>>
>>>>>>
>>>>>>
>>>>>> For future file types, like V4 root manifests and statistics files,
>>>>>> I think we can consider using a different approach. If we conclude on not
>>>>>> reusing key metadata across files, then I don’t see any point of having 
>>>>>> the
>>>>>> indirection of storing a `key-id` that refers to an item in the list.
>>>>>>
>>>>>> Instead we can follow a more direct approach, and simply store the
>>>>>> encrypted key metadata of such a file instead of a `key-id`. For instance
>>>>>> for V4 root manifest files we can store `encryption-key` (spec
>>>>>> <https://iceberg.apache.org/spec/#encryption-keys>) directly instead
>>>>>> of `key-id`.
>>>>>>
>>>>>>
>>>>>> Note, this is the structure of an `encryption-key`:
>>>>>>
>>>>>>    - Key-id
>>>>>>    - Encrypted-key-metadata
>>>>>>    - Encrypted-by-id
>>>>>>    - Properties
>>>>>>
>>>>>> In this design `key-id` in the structure is unnecessary. (might be
>>>>>> optional?)
>>>>>>
>>>>>>
>>>>>> We could still keep the KEKs in the `encryption-keys` list.
>>>>>>
>>>>>>
>>>>>> What do you think?
>>>>>>
>>>>>> Gabor
>>>>>>
>>>>>

Reply via email to