andygrove opened a new issue, #6707:
URL: https://github.com/apache/datafusion-comet/issues/6707
### Describe the bug
With the native Iceberg scan, which is on by default
(`spark.comet.scan.icebergNative.enabled=true`), `input_file_name()`,
`input_file_block_start()` and `input_file_block_length()` return `""`, `-1`
and `-1` for every row of an Iceberg table. Spark returns each row's data file
and block offsets. The query succeeds, so the wrong values are silent.
These expressions read `InputFileBlockHolder`, which Iceberg's Spark reader
sets as it opens each data file. The native scan never sets it. #3312 made
`CometScanRule` fall back from the native Parquet scan when the plan uses these
expressions, but `transformV2Scan` has no such check, and it is not handed the
plan to make one.
No Comet operator is needed between the scan and the projection:
```
*(1) Project [input_file_name() AS input_file_name()#27,
input_file_block_start() AS input_file_block_start()#28L,
input_file_block_length() AS input_file_block_length()#29L, id#26L]
+- *(1) CometColumnarToRow
+- CometIcebergNativeScan [id#26L], .../db/t/metadata/v4.metadata.json
```
### Steps to reproduce
```scala
spark.sql("CREATE TABLE cat.db.t (id BIGINT) USING iceberg")
for (i <- 0 until 3) {
spark.sql(s"INSERT INTO cat.db.t SELECT id FROM range(${i * 1000}, ${(i +
1) * 1000})")
}
spark
.sql("SELECT input_file_name(), input_file_block_start(),
input_file_block_length(), id FROM cat.db.t")
.show(3, false)
```
Comet returns `["", -1, -1, <id>]` for all 3000 rows. `SELECT
input_file_name(), id FROM cat.db.t WHERE id >= 0`, which puts a `CometFilter`
above the scan, returns `""` for every row too.
### Expected behavior
Each row reports the data file it was read from, as Spark does. Falling back
to Spark's Iceberg reader when the plan uses these expressions, as the native
Parquet scan does, would also be acceptable.
### Additional context
- Reproduced on `main` at `b80bf4e08` with the default Spark 4.1 profile and
a Hadoop catalog, comparing against the same query with
`spark.comet.enabled=false`.
- From reading `branch-1.0` and `branch-1.1`, the native Iceberg scan is on
by default in both and neither has a guard, so this most likely ships in 1.0
and 1.1.
- #6703 fixes the same problem for the Spark-to-Arrow conversions (#6573)
and notes this gap. Its `CometScanRule.readsInputFileBlock(plan)` could be
reused here once `transformV2Scan` has the plan.
- The "Current limitations" list in
`docs/source/user-guide/latest/iceberg.md` should name the fallback once there
is one.
- The native CSV V2 scan (`spark.comet.scan.csv.v2.enabled`, off by default)
goes through the same `transformV2Scan`. I have not tested it.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]