infvg commented on code in PR #12726:
URL: https://github.com/apache/gluten/pull/12726#discussion_r4144408066
##########
gluten-substrait/src/main/scala/org/apache/gluten/extension/columnar/PushDownInputFileExpression.scala:
##########
@@ -104,15 +103,44 @@ object PushDownInputFileExpression {
ProjectExec(f.output, FilterExec(newCondition, newChild))
}
+ /**
+ * Returns true when any of the injected metadata attribute names (all
lowercase) matches a
+ * column already present in the scan output at the Velox/case-insensitive
level.
+ *
+ * Velox is case-insensitive. If the scan has a user data column named
e.g. "Input_File_Name"
+ * its lowercase form "input_file_name" collides with the metadata
function of the same name.
+ * Velox rejects a TableScan that maps the same lowercase column name to
both a Regular handle
+ * (data column) and a PartitionKey/metadata handle (file-path metadata).
When this conflict is
+ * detected the scan must fall back to Vanilla so that Spark's own
FilePartitionReader sets the
+ * InputFileBlockHolder thread-local and input_file_name() returns the
correct file path.
+ */
+ private def hasVeloxColumnNameConflict(
+ scanOutput: Seq[org.apache.spark.sql.catalyst.expressions.Attribute],
+ replacedExprs: mutable.Map[String, Alias]): Boolean = {
+ val scanOutputLowerNames =
scanOutput.map(_.name.toLowerCase(java.util.Locale.ROOT)).toSet
Review Comment:
This causes an unnecessary fallback, with `spark.sql.caseSensitive=true`,
`Input_File_Name` and `input_file_name` are distinct, but this check forces
fallback.
##########
gluten-substrait/src/main/scala/org/apache/gluten/execution/BasicScanExecTransformer.scala:
##########
@@ -123,7 +123,9 @@ trait BasicScanExecTransformer extends LeafTransformSupport
with BaseDataSource
InputFileBlockLength().prettyName)
val neededInputFileRelatedMetadataKeys =
- inputFileRelatedMetadataKeys.filter(k => output.exists(_.name == k))
+ inputFileRelatedMetadataKeys.filter {
+ k => output.exists(a => ConverterUtils.normalizeColName(a.name) == k)
+ }
Review Comment:
I tested this locally, and the column only case returns incorrect results
when caseSensitive is false:
```
SET spark.sql.caseSensitive=false;
CREATE TABLE table1 (`Input_File_Name` STRING) USING parquet;
INSERT INTO table1 VALUES ('value1'), ('value2');
SELECT `Input_File_Name` FROM table1;
```
Expected:
```
value1
value2
```
Actual:
```
file:///tmp/pr12726-column-data-12367331577223769254/part-00000-21512ca9-b174-40be-8888-d148f6e54129-c000.snappy.parquet
file:///tmp/pr12726-column-data-12367331577223769254/part-00000-21512ca9-b174-40be-8888-d148f6e54129-c000.snappy.parquet
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]