andygrove commented on code in PR #6326:
URL: https://github.com/apache/datafusion-comet/pull/6326#discussion_r4125286487


##########
spark/src/main/scala/org/apache/comet/serde/operator/CometNativeScan.scala:
##########
@@ -143,6 +143,21 @@ object CometNativeScan extends 
CometOperatorSerde[CometScanExec] with CometTypeS
       withFallbackReason(scanExec, unsupportedDefaultReason)
     }
 
+    // Under the case-insensitive resolver Spark's reader matches each 
requested field to the file
+    // fields with the same folded name and raises when more than one answers. 
The native scan
+    // only resolves nested names while casting a column whose file type 
differs from the
+    // requested one. When they are equal DataFusion reads the column 
positionally, and its opener
+    // skips the expression adapter altogether when the schemas match and no 
predicate is pushed.
+    // Spark's analyzer rejects such a requested schema, but a DataFrame 
analyzed under the
+    // case-sensitive resolver still reaches here, so let Spark's reader 
resolve it (#6136).

Review Comment:
   The comment says a DataFrame analyzed under the case-sensitive resolver is 
how this schema gets past the analyzer, but plain SQL reaches here too. Spark 
caches the resolved relation for a catalog table, and later reads reuse it 
without running `checkSchemaColumnNameDuplication` again. So `CREATE TABLE t 
(id bigint, s struct<x: bigint, X: bigint>) USING parquet`, one `SELECT * FROM 
t` with `spark.sql.caseSensitive=true`, and then the same query under the 
default all land on this gate. A temp view created under the case-sensitive 
resolver does as well. The gate handles both correctly, so this is only about 
the comment. Could it mention the cached-table path? It's plain SQL, so it's 
the more likely way in, and whoever eventually swaps this gate for a native 
check will want to test it.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to