andygrove opened a new issue, #6707:
URL: https://github.com/apache/datafusion-comet/issues/6707

   ### Describe the bug
   
   With the native Iceberg scan, which is on by default 
(`spark.comet.scan.icebergNative.enabled=true`), `input_file_name()`, 
`input_file_block_start()` and `input_file_block_length()` return `""`, `-1` 
and `-1` for every row of an Iceberg table. Spark returns each row's data file 
and block offsets. The query succeeds, so the wrong values are silent.
   
   These expressions read `InputFileBlockHolder`, which Iceberg's Spark reader 
sets as it opens each data file. The native scan never sets it. #3312 made 
`CometScanRule` fall back from the native Parquet scan when the plan uses these 
expressions, but `transformV2Scan` has no such check, and it is not handed the 
plan to make one.
   
   No Comet operator is needed between the scan and the projection:
   
   ```
   *(1) Project [input_file_name() AS input_file_name()#27, 
input_file_block_start() AS input_file_block_start()#28L, 
input_file_block_length() AS input_file_block_length()#29L, id#26L]
   +- *(1) CometColumnarToRow
      +- CometIcebergNativeScan [id#26L], .../db/t/metadata/v4.metadata.json
   ```
   
   ### Steps to reproduce
   
   ```scala
   spark.sql("CREATE TABLE cat.db.t (id BIGINT) USING iceberg")
   for (i <- 0 until 3) {
     spark.sql(s"INSERT INTO cat.db.t SELECT id FROM range(${i * 1000}, ${(i + 
1) * 1000})")
   }
   spark
     .sql("SELECT input_file_name(), input_file_block_start(), 
input_file_block_length(), id FROM cat.db.t")
     .show(3, false)
   ```
   
   Comet returns `["", -1, -1, <id>]` for all 3000 rows. `SELECT 
input_file_name(), id FROM cat.db.t WHERE id >= 0`, which puts a `CometFilter` 
above the scan, returns `""` for every row too.
   
   ### Expected behavior
   
   Each row reports the data file it was read from, as Spark does. Falling back 
to Spark's Iceberg reader when the plan uses these expressions, as the native 
Parquet scan does, would also be acceptable.
   
   ### Additional context
   
   - Reproduced on `main` at `b80bf4e08` with the default Spark 4.1 profile and 
a Hadoop catalog, comparing against the same query with 
`spark.comet.enabled=false`.
   - From reading `branch-1.0` and `branch-1.1`, the native Iceberg scan is on 
by default in both and neither has a guard, so this most likely ships in 1.0 
and 1.1.
   - #6703 fixes the same problem for the Spark-to-Arrow conversions (#6573) 
and notes this gap. Its `CometScanRule.readsInputFileBlock(plan)` could be 
reused here once `transformV2Scan` has the plan.
   - The "Current limitations" list in 
`docs/source/user-guide/latest/iceberg.md` should name the fallback once there 
is one.
   - The native CSV V2 scan (`spark.comet.scan.csv.v2.enabled`, off by default) 
goes through the same `transformV2Scan`. I have not tested it.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to