Thank you, Daniel and Shawn, for explaining your thoughts on this. I will see what we can do in our side to evolve our platform tooling to be able to pull the Iceberg metadata instead.
Javier From: Daniel Weeks <[email protected]> Date: Monday, 5 October 2026 at 22:06 To: [email protected] <[email protected]> Cc: EGDL <[email protected]>; Javier Sanchez Beltran <[email protected]> Subject: [External] Re: GlueCatalog and what it writes to the Glue table on commit Javier, I think issue underpins that there's a "split-brain" issue with both HMS and Glue, since they both have their own metadata and also rely on the Iceberg metadata. In both HMS and Glue, the catalog implementations reflect some (but not all) information in their managed metadata. For example, HMS stores the schema and properties, but not the partitioning or sort order. Glue stores the schema, but not other properties, partitioning, or sort order. In both cases, the schema may not be converted/reflected perfectly (time/UUID/ect). What's worse is that in both HMS and Glue, if you use tools that work natively with the catalog APIs (e.g. the AWS console or SDK), you can update the representation so that it isn't consistent with the underlying Iceberg table metadata. I don't know how much you can trust that the metadata is being kept in sync and would definitely not rely on it for security or governance purposes. -Dan On Thu, Oct 1, 2026 at 10:11 AM Shawn Chang <[email protected]<mailto:[email protected]>> wrote: Hi Javier, I think the informational metadata footprint in Glue is intentional, and I would be cautious about making Glue mirror HMS here. GlueCatalog and Lake Formation have a different design principle from HMS. For Glue, the catalog entry is primarily a control-plane/catalog representation of the Iceberg table, while the Iceberg metadata remains the source of truth. Lake Formation then defines the authorization boundary for access to that catalog metadata and to the underlying data/iceberg metadata. HMS serves a different purpose. It is also an interoperability surface for Hive clients, which historically inspect HMS table parameters directly. That is why Iceberg projects much more state into HMS, including table properties and derived metadata such as schema, partition spec, sort order, and snapshot information. Best, Shawn On Thu, Oct 1, 2026 at 8:28 AM Javier Sanchez Beltran via dev <[email protected]<mailto:[email protected]>> wrote: Hi all, a question about GlueCatalog and what it writes to the Glue table on commit. GlueTableOperations.prepareProperties<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/aws/src/main/java/org/apache/iceberg/aws/glue/GlueTableOperations.java#L291-L301> sets only table_type, metadata_location and previous_metadata_locationin the Glue table Parameters. It carries over whatever parameters the Glue table already had, but never copies TableMetadata.properties(). The Hive catalog works differently. HMSTablePropertyHelper.updateHmsTableForIcebergTable<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HMSTablePropertyHelper.java#L78-L130> pushes all Iceberg table properties into HMS params, plus the current snapshot summary, schema, partition spec and sort order, all capped by iceberg.hive.table-property-max-size (https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HiveOperationsBase.java#L53). So the same table shows a lot more metadata through HMS than through Glue. Questions: 1. Is the minimal Glue footprint intentional? The javadoc says Glue info is "informational only" and metadata_location is the source of truth. Or has nobody needed more yet? 2. Would the community accept a change that mirrors the Hive behavior for Glue? That would mean syncing table properties (and maybe the snapshot summary/schema/spec) into Glue Parameters, with a size cap and an opt-in flag. Open design questions would be Glue's parameter size limits, removing keys that were deleted in Iceberg, and protecting reserved keys. Our use case: tools that only read the Glue catalog (governance, discovery, Lake Formation-based tooling) can't see Iceberg table properties without opening the metadata JSON. Slack message - https://apache-iceberg.slack.com/archives/C025PH0G1D4/p1790773851197099
