Hi Karthik,

Thanks for the detailed write-up. Since GitHub issue creation was not
working on your side, I filed it on your behalf so it stays tracked in one
place:

https://github.com/apache/incubator-xtable/issues/871

On the initialization: yes, creating the table with an empty schema is
intentional, so the Iceberg API does not remap the field IDs derived from
the source. The leftover unpartitioned spec 0 is a side effect of that
rather than a design goal, and passing the final spec into createTable is
not enough on its own, since the spec has to be bound to a schema whose
field IDs we control. Details and a possible direction are in the issue.

That said, I think the blocker sits on the Databricks side rather than in
the metadata XTable produces. Partition evolution is a core Iceberg feature
and the table is spec-correct, which you confirmed by reading it from
Athena and EMR. Databricks does not read Iceberg natively in this path, it
converts to Delta [1], and that conversion is what rejects multiple specs.

A user running this in production pointed to a Databricks-side config,
*spark.databricks.delta.convert.iceberg.partitionEvolution.enabled*, which
appears to relax the check. It is not in the Databricks docs for this path,
so this comes from their experience rather than published documentation.
Could you try it and report back? If it works, it would also unblock the
tables you have already synced.

Thanks again for the report.

Best,
Vinish

[1]
https://docs.databricks.com/gcp/en/ingestion/data-migration/convert-to-delta#converting-a-directory-of-parquet-or-iceberg-files-in-an-external-location-to-delta-lake


On Wed, Jul 29, 2026 08:14 AM, "Vutukuru, Karthik via dev" <
[email protected]> wrote:

> Internal
>
> Hi Xtable Team,
>
> Hope this message finds you well. I tried to create an issue on github but
> I always run into issue. Hence, emailing below to get your attention.
>
> Iceberg metadata generated by XTable records partition evolution during
> initial sync, causing Databricks Delta clone incompatibility
>
> We are using Apache XTable to generate Iceberg metadata over existing Hudi
> tables registered in AWS Glue.
> The generated Iceberg tables are readable by Iceberg-compatible engines,
> such as Athena (Trino) and AWS EMR(spark)  but Databricks fails when
> attempting to use the table through its Iceberg-to-Delta clone path with
> the following error:
>
> [DELTA_CLONE_INCOMPATIBLE_SOURCE.ICEBERG_UNDERGONE_PARTITION_EVOLUTION]
>
> The clone source has valid format, but has unsupported feature with Delta.
> Source iceberg table has undergone partition evolution.
>
> SQLSTATE: 0AKDC
>
> After inspecting the generated Iceberg metadata, it appears that XTable
> creates the table initially as unpartitioned and then applies the actual
> source partition specification in a subsequent metadata commit.
> This creates two Iceberg partition specifications:
>
> spec-id 0: unpartitioned
> spec-id 1: identity(df_meta_p_cppib)
>
> Although no data files were written using spec 0, Databricks detects the
> presence of multiple partition specifications and classifies the table as
> having undergone partition evolution.
> Environment
> XTable artifact:
>
> xtable-utilities_2.12-0.2.0-SNAPSHOT-bundled.jar
>
> Source table format:
>
> Apache Hudi
> COPY_ON_WRITE
> Parquet
>
> Target format:
>
> Apache Iceberg format version 2
> AWS Glue catalog
> S3FileIO
>
> Target consumers include Databricks.
> Source table partitioning
> The Hudi source table is partitioned by:
>
> df_meta_p_cppib
>
> The partition definition was not intentionally changed after table
> creation.
> The expected Iceberg partition specification is:
>
> identity(df_meta_p_cppib)
>
> XTable execution
> XTable was run against the Hudi base path to generate an Iceberg target.
> The generated Iceberg metadata directory contained:
>
> 00000-444ce851-9abd-4d4b-98ea-947f11a2f914.metadata.json
> 00001-87e5211c-cfb3-4b10-ae04-4fca25f0a5b2.metadata.json
> 00002-312c2a3b-a430-4f4b-803a-9262c71cf013.metadata.json
> snap-5065718056232190869-1-2b42ae3a-edc2-4357-8eff-1fd32dfb6ddb.avro
> 2b42ae3a-edc2-4357-8eff-1fd32dfb6ddb-m0.avro
>
> These files were created within approximately two seconds:
>
> 2026-07-20 13:41:12  00000-...metadata.json
> 2026-07-20 13:41:13  00001-...metadata.json
> 2026-07-20 13:41:14  00002-...metadata.json
>
> The current Glue table points to:
>
> metadata_location:
>
> s3://cppib-data-fabric-trusted-qa/om/om_europe_country/metadata/00002-312c2a3b-a430-4f4b-803a-9262c71cf013.metadata.json
>
> previous_metadata_location:
>
> s3://cppib-data-fabric-trusted-qa/om/om_europe_country/metadata/00001-87e5211c-cfb3-4b10-ae04-4fca25f0a5b2.metadata.json
>
> Current partition metadata
> The current Iceberg metadata contains:
>
> {
>   "format_version": 2,
>   "default_spec_id": 1,
>   "last_partition_id": 1000,
>   "partition_spec_count": 2,
>   "partition_specs": [
>     {
>       "spec_id": 0,
>       "fields": []
>     },
>     {
>       "spec_id": 1,
>       "fields": [
>         {
>           "name": "df_meta_p_cppib",
>           "transform": "identity",
>           "source-id": 14,
>           "field-id": 1000
>         }
>       ]
>     }
>   ]
> }
>
> The raw Iceberg representation is effectively:
>
> {
>   "partition-specs": [
>     {
>       "spec-id": 0,
>       "fields": []
>     },
>     {
>       "spec-id": 1,
>       "fields": [
>         {
>           "name": "df_meta_p_cppib",
>           "transform": "identity",
>           "source-id": 14,
>           "field-id": 1000
>         }
>       ]
>     }
>   ],
>   "default-spec-id": 1
> }
>
> The manifest metadata confirms that all actual data files use only
> partition spec 1:
>
> partition_spec_id = 1
>
> Example manifest output:
>
> path:
>
> s3://cppib-data-fabric-trusted-qa/om/om_europe_country/metadata/2b42ae3a-edc2-4357-8eff-1fd32dfb6ddb-m0.avro
>
> partition_spec_id: 1
> added_data_files_count: 1
> added_rows_count: 260
>
> No data files appear to have been written under partition spec 0.
> Metadata history
> The latest metadata file contains the following metadata log:
>
> {
>   "metadata_log": [
>     {
>       "timestamp-ms": 1784554871636,
>       "metadata-file":
> "s3://cppib-data-fabric-trusted-qa/om/om_europe_country/metadata/00000-444ce851-9abd-4d4b-98ea-947f11a2f914.metadata.json"
>     },
>     {
>       "timestamp-ms": 1784554872506,
>       "metadata-file":
> "s3://cppib-data-fabric-trusted-qa/om/om_europe_country/metadata/00001-87e5211c-cfb3-4b10-ae04-4fca25f0a5b2.metadata.json"
>     }
>   ],
>   "snapshots": [
>     {
>       "snapshot_id": 5065718056232190869,
>       "timestamp_ms": 1784554873567,
>       "operation": "append"
>     }
>   ]
> }
>
> There is only one data snapshot, and its operation is append.
> A comparison of the partition metadata in 00001 and 00002 shows no
> difference:
>
> diff -u \
>   <(jq -S '{
>       default_spec_id: ."default-spec-id",
>       last_partition_id: ."last-partition-id",
>       partition_specs: ."partition-specs"
>     }' 00001.metadata.json) \
>   <(jq -S '{
>       default_spec_id: ."default-spec-id",
>       last_partition_id: ."last-partition-id",
>       partition_specs: ."partition-specs"
>     }' 00002.metadata.json)
>
> The command produces no output.
> This suggests that the partition evolution occurred between metadata
> versions 00000 and 00001, before the data snapshot was added.
> Suspected XTable behavior
> The observed metadata is consistent with the following initialization
> sequence:
>
> 1. Create a new Iceberg table with an empty schema and an unpartitioned
> spec.
> 2. Synchronize the Hudi schema.
> 3. Apply identity(df_meta_p_cppib) as a new partition specification.
> 4. Add the Hudi data files under the new spec.
>
> The result is:
>
> metadata 00000:
>   spec 0 = unpartitioned
>
> metadata 00001:
>   spec 0 = unpartitioned
>   spec 1 = identity(df_meta_p_cppib)
>
> metadata 00002:
>   append data files under spec 1
>
> From a business and data perspective, the partitioning did not evolve. The
> intended partition definition was always:
>
> identity(df_meta_p_cppib)
>
> However, from the Iceberg metadata perspective, the table evolved from
> unpartitioned to partitioned.
> Impact
> Databricks rejects the table because it sees multiple partition
> specifications:
>
> [DELTA_CLONE_INCOMPATIBLE_SOURCE.ICEBERG_UNDERGONE_PARTITION_EVOLUTION]
>
> This occurs even though:
>
>   *   There is only one data snapshot.
>   *   All data files use the same partition specification.
>   *   No data files were written while the table was unpartitioned.
>   *   The source Hudi table’s partitioning was not changed.
>   *   The extra unpartitioned spec appears to be an initialization
> artifact.
>
> This makes XTable-generated Iceberg metadata incompatible with Databricks
> workflows that require the source Iceberg table not to have undergone
> partition evolution.
> Expected behavior
> For a newly generated Iceberg target, XTable should create the table with
> the final source schema and partition specification in the initial Iceberg
> metadata commit.
> The expected initial metadata should contain:
>
> {
>   "default-spec-id": 0,
>   "partition-specs": [
>     {
>       "spec-id": 0,
>       "fields": [
>         {
>           "name": "df_meta_p_cppib",
>           "transform": "identity",
>           "source-id": 14,
>           "field-id": 1000
>         }
>       ]
>     }
>   ]
> }
>
> The expected result is:
>
> spec count = 1
> default spec ID = 0
> spec 0 = identity(df_meta_p_cppib)
>
> All generated manifests should then use:
>
> partition_spec_id = 0
>
> Actual behavior
> The generated table contains:
>
> spec count = 2
> default spec ID = 1
> spec 0 = unpartitioned
> spec 1 = identity(df_meta_p_cppib)
>
> All data uses spec 1, but the unused spec 0 remains in table metadata.
> Requested guidance
> Could the XTable maintainers clarify:
> 1. Is creating an initially unpartitioned Iceberg table an intentional
> part of XTable’s table-creation flow?
> 2. Is this done to preserve source schema field IDs from Hudi or the
> existing Parquet files?
> 3. Can the final Hudi-derived schema and partition specification be
> supplied during the initial Iceberg table creation so the intended
> partition becomes spec 0?
> 4. Is there an existing configuration or creation mode that avoids the
> empty-schema and unpartitioned-spec initialization?
> 5. Would the project accept a change that introduces a
> Databricks-compatible creation mode, for example:
>
> Create the Iceberg table with the final schema and final partition spec
> in the first metadata commit.
>
> 6. If preserving source field IDs requires the current initialization
> approach, is there another supported way to preserve IDs without first
> creating an unpartitioned spec?
> 7. Is this behavior already addressed in a newer XTable version or branch?
> Possible fix
> A possible implementation would be to create a new Iceberg table using the
> complete schema and partition spec together:
>
> catalog.createTable(
>     tableIdentifier,
>     finalSchema,
>     finalPartitionSpec,
>     basePath,
>     tableProperties
> );
>
> Instead of first creating:
>
> catalog.createTable(
>     tableIdentifier,
>     new Schema(),
>     PartitionSpec.unpartitioned(),
>     basePath,
>     tableProperties
> );
>
> and subsequently applying the actual schema and partition specification.
> The fix would need to account for XTable’s schema field-ID preservation
> requirements.
> A useful validation test would assert that, after the initial
> synchronization of a partitioned Hudi source:
>
> assertThat(table.specs()).hasSize(1);
> assertThat(table.spec().specId()).isEqualTo(0);
> assertThat(table.spec().fields()).hasSize(1);
> assertThat(table.spec().fields().get(0).name())
>     .isEqualTo("df_meta_p_cppib");
>
> A second identical sync should also leave the table with only one
> partition spec.
> Reproduction outline
> 1. Create a partitioned Hudi table with partition column:
>
> df_meta_p_cppib
>
> 2. Run XTable to generate an Iceberg target in AWS Glue.
> 3. Download the generated Iceberg metadata JSON.
> 4. Inspect the partition specifications:
>
> jq '{
>   default_spec_id: ."default-spec-id",
>   partition_spec_count: (."partition-specs" | length),
>   partition_specs: ."partition-specs"
> }' current.metadata.json
>
> 5. Observe:
>
> spec 0 = unpartitioned
> spec 1 = identity(df_meta_p_cppib)
>
> 6. Query the manifest metadata and observe that only spec 1 is used:
>
> SELECT
>     partition_spec_id,
>     COUNT(*) AS manifest_count
> FROM "<database>"."<table>$manifests"
> GROUP BY partition_spec_id;
>
> 7. Attempt the relevant operation from Databricks.
> 8. Observe:
>
> DELTA_CLONE_INCOMPATIBLE_SOURCE.ICEBERG_UNDERGONE_PARTITION_EVOLUTION
>
> Additional related observation
> We also attempted to pre-create an Iceberg table so that the intended
> partition could be present as spec 0 before XTable synchronization.
> The pre-created metadata included fields such as:
>
> {
>   "initial-default": null,
>   "write-default": null
> }
>
> The Iceberg parser bundled with the XTable snapshot failed while loading
> that metadata:
>
> java.lang.NullPointerException:
> Cannot create expression literal from null
>
> at org.apache.iceberg.expressions.Literals.from
> at org.apache.iceberg.expressions.Expressions.lit
> at org.apache.iceberg.SchemaParser.defaultFromJson
>
> This appears to be a separate Iceberg-version compatibility issue, but it
> makes manually pre-creating the target difficult as a workaround.
> The preferred solution would therefore be for XTable itself to create the
> initial Iceberg metadata using the final schema and partition specification.
> Summary
> XTable appears to record an unused unpartitioned spec during initial
> Iceberg target creation:
>
> spec 0 = unpartitioned
> spec 1 = intended partition
>
> Databricks interprets this as partition evolution and rejects the table,
> even though all data files use only the intended partition specification.
> It would be helpful to have either:
>
>   *   an XTable creation mode that initializes the Iceberg target with the
> final partition spec as spec 0;
>   *   a fix to the default table-creation sequence; or
>   *   guidance on a supported workaround that preserves field IDs without
> generating partition evolution.
>
> Thanks,
> Karthik
>

Reply via email to