bydeath opened a new issue, #9990: URL: https://github.com/apache/paimon/issues/9990
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar. ### Paimon version - `1.4.1` (Bug occurs) - `2.0.0` (Bug occurs) - `1.3.1` (Works fine - baseline) ### Compute Engine - `Spark 3.2.3` (spark-sql / Scala 2.12) - Catalog: `HiveMetastore 2.3` via `org.apache.paimon.spark.SparkCatalog` (or `SparkGenericCatalog`) - Storage: Hadoop HDFS 3.3.6 (`fs.defaultFS = hdfs://namenode:8020`) ### Minimal reproduce step ### Environment Configuration ```properties # spark-defaults.conf or spark-sql parameters spark.sql.catalog.paimon_hive=org.apache.paimon.spark.SparkCatalog spark.sql.catalog.paimon_hive.metastore=hive spark.sql.catalog.paimon_hive.uri=thrift://hive-metastore:9083 spark.sql.catalog.paimon_hive.warehouse=hdfs://namenode:8020/warehouse spark.sql.defaultCatalog=paimon_hive spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions ``` ### Step 1: Create Table with schemeless LOCATION ```sql CREATE DATABASE IF NOT EXISTS paimon_db; USE paimon_db; CREATE TABLE my_table_external ( user_id BIGINT, item_id BIGINT, behavior STRING, dt STRING, hh STRING ) PARTITIONED BY (dt, hh) TBLPROPERTIES ( 'primary-key' = 'dt,hh,user_id' ) LOCATION '/data/external/my_table_external'; ``` ### Step 2: Observe the Stack Trace ```text java.lang.RuntimeException: org.apache.spark.sql.catalyst.analysis.NoSuchTableException: Table paimon_db.my_table_external not found at org.apache.paimon.spark.SparkCatalog.createTable(SparkCatalog.java:383) at org.apache.spark.sql.execution.datasources.v2.CreateTableExec.run(CreateTableExec.scala:42) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.result$lzycompute(V2CommandExec.scala:43) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.result(V2CommandExec.scala:43) at org.apache.spark.sql.execution.datasources.v2.V2CommandExec.executeCollect(V2CommandExec.scala:49) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.$anonfun$applyOrElse$1(QueryExecution.scala:97) at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withNewExecutionId$5(SQLExecution.scala:103) at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:163) at org.apache.spark.sql.execution.SQLExecution$.$anonfun$withNewExecutionId$1(SQLExecution.scala:90) at org.apache.spark.sql.SparkSession.withActive(SparkSession.scala:775) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:64) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.applyOrElse(QueryExecution.scala:97) at org.apache.spark.sql.execution.QueryExecution$$anonfun$eagerlyExecuteCommands$1.applyOrElse(QueryExecution.scala:93) at org.apache.spark.sql.catalyst.trees.TreeNode.$anonfun$transformDownWithPruning$1(TreeNode.scala:481) at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:82) at org.apache.spark.sql.catalyst.trees.TreeNode.transformDownWithPruning(TreeNode.scala:481) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.org$apache$spark$sql$catalyst$plans$logical$AnalysisHelper$$super$transformDownWithPruning(LogicalPlan.scala:30) at org.apache.spark.sql.catalyst.plans.logical.AnalysisHelper.transformDownWithPruning(AnalysisHelper.scala:267) at org.apache.spark.sql.catalyst.plans.logical.AnalysisHelper.transformDownWithPruning$(AnalysisHelper.scala:263) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.transformDownWithPruning(LogicalPlan.scala:30) at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.transformDownWithPruning(LogicalPlan.scala:30) at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:457) at org.apache.spark.sql.execution.QueryExecution.eagerlyExecuteCommands(QueryExecution.scala:93) at org.apache.spark.sql.execution.QueryExecution.commandExecuted$lzycompute(QueryExecution.scala:80) at org.apache.spark.sql.execution.QueryExecution.commandExecuted(QueryExecution.scala:78) at org.apache.spark.sql.Dataset.<init>(Dataset.scala:219) at org.apache.spark.sql.Dataset$.$anonfun$ofRows$2(Dataset.scala:99) at org.apache.spark.sql.SparkSession.withActive(SparkSession.scala:775) at org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:96) at org.apache.spark.sql.SparkSession.$anonfun$sql$1(SparkSession.scala:618) at org.apache.spark.sql.SparkSession.withActive(SparkSession.scala:775) at org.apache.spark.sql.SparkSession.sql(SparkSession.scala:613) at org.apache.spark.sql.SQLContext.sql(SQLContext.scala:651) at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:67) at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:384) Caused by: org.apache.spark.sql.catalyst.analysis.NoSuchTableException: Table paimon_db.my_table_external not found at org.apache.paimon.spark.SparkCatalog.loadSparkTable(SparkCatalog.java:837) at org.apache.paimon.spark.SparkCatalog.loadTable(SparkCatalog.java:309) at org.apache.paimon.spark.SparkCatalog.createTable(SparkCatalog.java:377) ... 58 more ``` ### Step 3: Inspect Local Filesystem vs HDFS ```bash # Check local filesystem on Spark client/driver: ls -la /data/external/my_table_external/schema/ # Output: schema-0 exists! (WRONG TARGET) # Check HDFS: hdfs dfs -ls -R /data/external/my_table_external/ # Output: Empty directory! (schema-0 missing on HDFS) ``` ### What doesn't meet your expectations? In a Hadoop cluster environment with Hive Metastore and HDFS (`fs.defaultFS`), executing `CREATE TABLE ... LOCATION '/data/external/...'` with a schemeless absolute path should: 1. Resolve the path against `fs.defaultFS` and initialize `schema-0` on HDFS. 2. Register the table in Hive Metastore and return successfully (as it did in Paimon 1.3.1). ### Anything else? ### Workaround Explicitly specify the URI scheme in `LOCATION`: ```sql -- Use hdfs:/// prefix: CREATE TABLE my_table_external ( ... ) LOCATION 'hdfs:///data/external/my_table_external'; ``` Adding `hdfs:///` allows `FileIO.get` to detect the `hdfs` scheme and correctly instantiate `HadoopFileIO`. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
