Thank you, Daniel and Shawn, for explaining your thoughts on this.

I will see what we can do in our side to evolve our platform tooling to be able 
to pull the Iceberg metadata instead.

Javier

From: Daniel Weeks <[email protected]>
Date: Monday, 5 October 2026 at 22:06
To: [email protected] <[email protected]>
Cc: EGDL <[email protected]>; Javier Sanchez Beltran 
<[email protected]>
Subject: [External] Re: GlueCatalog and what it writes to the Glue table on 
commit

Javier,

I think issue underpins that there's a "split-brain" issue with both HMS and 
Glue, since they both have their own metadata and also rely on the Iceberg 
metadata.

In both HMS and Glue, the catalog implementations reflect some (but not all) 
information in their managed metadata.  For example, HMS stores the schema and 
properties, but not the partitioning or sort order.  Glue stores the schema, 
but not other properties, partitioning, or sort order.  In both cases, the 
schema may not be converted/reflected perfectly (time/UUID/ect).

What's worse is that in both HMS and Glue, if you use tools that work natively 
with the catalog APIs (e.g. the AWS console or SDK), you can update the 
representation so that it isn't consistent with the underlying Iceberg table 
metadata.

I don't know how much you can trust that the metadata is being kept in sync and 
would definitely not rely on it for security or governance purposes.

-Dan

On Thu, Oct 1, 2026 at 10:11 AM Shawn Chang 
<[email protected]<mailto:[email protected]>> wrote:

Hi Javier,

I think the informational metadata footprint in Glue is intentional, and I 
would be cautious about making Glue mirror HMS here.

GlueCatalog and Lake Formation have a different design principle from HMS. For 
Glue, the catalog entry is primarily a control-plane/catalog representation of 
the Iceberg table, while the Iceberg metadata remains the source of truth. Lake 
Formation then defines the authorization boundary for access to that catalog 
metadata and to the underlying data/iceberg metadata.

HMS serves a different purpose. It is also an interoperability surface for Hive 
clients, which historically inspect HMS table parameters directly. That is why 
Iceberg projects much more state into HMS, including table properties and 
derived metadata such as schema, partition spec, sort order, and snapshot 
information.

Best,

Shawn

On Thu, Oct 1, 2026 at 8:28 AM Javier Sanchez Beltran via dev 
<[email protected]<mailto:[email protected]>> wrote:
Hi all, a question about GlueCatalog and what it writes to the Glue table on 
commit.

GlueTableOperations.prepareProperties<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/aws/src/main/java/org/apache/iceberg/aws/glue/GlueTableOperations.java#L291-L301>
 sets only table_type, metadata_location and previous_metadata_locationin the 
Glue table Parameters. It carries over whatever parameters the Glue table 
already had, but never copies TableMetadata.properties().

The Hive catalog works differently. 
HMSTablePropertyHelper.updateHmsTableForIcebergTable<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HMSTablePropertyHelper.java#L78-L130>
 pushes all Iceberg table properties into HMS params, plus the current snapshot 
summary, schema, partition spec and sort order, all capped by 
iceberg.hive.table-property-max-size 
(https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HiveOperationsBase.java#L53).
 So the same table shows a lot more metadata through HMS than through Glue.

Questions:

  1.  Is the minimal Glue footprint intentional? The javadoc says Glue info is 
"informational only" and metadata_location is the source of truth. Or has 
nobody needed more yet?
  2.  Would the community accept a change that mirrors the Hive behavior for 
Glue? That would mean syncing table properties (and maybe the snapshot 
summary/schema/spec) into Glue Parameters, with a size cap and an opt-in flag. 
Open design questions would be Glue's parameter size limits, removing keys that 
were deleted in Iceberg, and protecting reserved keys.


Our use case: tools that only read the Glue catalog (governance, discovery, 
Lake Formation-based tooling) can't see Iceberg table properties without 
opening the metadata JSON.

Slack message - 
https://apache-iceberg.slack.com/archives/C025PH0G1D4/p1790773851197099


Reply via email to