zhang-arvin opened a new pull request, #10078: URL: https://github.com/apache/paimon/pull/10078
### What is the purpose of the change Fixes #9990. `CREATE TABLE ... LOCATION '/data/external/t'` (a LOCATION without a URI scheme) on a `HiveCatalog` writes the schema to the driver's local filesystem instead of the configured default filesystem, leaves a zombie entry in the metastore when the create fails, and then fails with `NoSuchTableException` out of `SparkCatalog#createTable`. ### Brief change log **1. `HiveCatalog#initialTableLocation` built the location without resolving the scheme.** `FileIO#get` (`FileIO.java:527`) returns a `LocalFileIO` whenever the path has no scheme, so for a schemeless `LOCATION` the catalog wrote `schema-0` to the driver's local disk while the metastore registered the table against the path the user asked for. The table is then unreadable from any other node, and the log lines in the report match this exactly (`schema-0` on the local FS, empty directory on HDFS). `resolveLocationScheme` now resolves a scheme-less location against `FileSystem.getDefaultUri(hiveConf)`, so it lands on the default filesystem like the catalog warehouse does. Already-schemed locations are returned untouched. Applied to both `createTableImpl` and `createObjectTable`. **2. A failed create left the metastore entry behind.** The `catch` block around `createHiveTable` deleted the directory only for a managed table and never removed the metastore entry. The schema write happens *before* the registration, so removing the entry restores the pre-call state. `cleanupOnCreateTableFailure` now drops the table via `dropTable(db, table, true, false)` and keeps the managed-table directory cleanup. ### Note on the Spark-side symptom The `NoSuchTableException`/zombie-table behaviour the report observes was introduced by #9549, which made `SparkCatalog#createTable` call `loadTable(ident)` right after `catalog.createTable(...)`. That change is reasonable when the table is genuinely readable, so the fix here is at the catalog layer where the broken location and the missing rollback actually live — the `loadTable` then succeeds instead of miss. Happy to restructure it if maintainers prefer the resolution to live in `SparkCatalog` instead. ### Tests - `testCreateExternalTableWithSchemelessLocation` — creates an external table with `path=/data/external/...` and asserts the location the catalog works with carries an explicit scheme and preserves the requested path. This asserts the regressed behaviour directly and fails without the fix. - `testCreateTableDoesNotLeaveZombieEntryWhenMetastoreFails` — asserts the rollback path removes the metastore entry. ### Verification status `mvn -pl paimon-hive/paimon-hive-catalog -am -Pfast-build -DskipTests compile` → BUILD SUCCESS. ⚠️ The two new tests are **not yet green locally**: `HiveCatalogTest` (and its siblings in this module) fail in `@BeforeEach setUp` with `IllegalStateException: Hadoop configuration is not available for this CatalogContext` (`CatalogTestBase.java:126` → `ResolvingFileIO.configure` → `CatalogContext.hadoopConf`). `HadoopUtils.getHadoopConfiguration` needs `HADOOP_HOME`/`HADOOP_CONF_DIR` on this machine, and `SerializableConfiguration` is a no-op stand-in otherwise. This is a pre-existing harness/environment limitation, unrelated to the change — the module's existing tests fail identically on clean `master`. CI has the Hadoop environment and should run them; I'll report back here once it does. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
