dwsmith1983 opened a new issue, #6746:
URL: https://github.com/apache/datafusion-comet/issues/6746

   ## Describe the bug
   
   The native Parquet scan builds one object store per file partition, from 
that partition's first file 
([planner.rs#L1950-L1967](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/native/core/src/execution/planner.rs#L1950-L1967)),
 and `get_partitioned_files` keeps only each file's path, dropping the scheme 
and authority 
([planner.rs#L491-L512](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/native/core/src/execution/planner.rs#L491-L512)).
 Spark's `FilePartition.getFilePartitions` packs small files from every root 
path into the same partitions, so a scan over two buckets can put files from 
both into one partition. The files from the second bucket are then fetched from 
the first. When the same key exists in both buckets the scan returns the first 
bucket's rows twice, with no error.
   
   `CometScanRule` already falls back for this when the paths use an opt-in 
S3-compatible alias scheme, and its comment notes that plain `s3://` and 
`s3a://` have the same flaw 
([CometScanRule.scala#L293-L307](https://github.com/apache/datafusion-comet/blob/e038cf05fa665dc61fdecbdde1c07b8b6a1ce3cb/spark/src/main/scala/org/apache/comet/rules/CometScanRule.scala#L293-L307)).
 Plain multi-bucket scans still run natively. #6059 declines the ABFS case 
(more than one container or account) for the same reason.
   
   ## Steps to reproduce
   
   MinIO through `CometS3TestBase`, two buckets `repro-a` and `repro-b`. Each 
holds two small Parquet files written by Spark with Comet off: ids 1-3 and 4-6, 
plus a column `bucket` = `'A'` or `'B'`. To get all four files into one 
partition, set `spark.sql.files.minPartitionNum=1`, 
`spark.sql.files.openCostInBytes=1` and 
`spark.sql.files.maxPartitionBytes=128MB`. Then run:
   
   ```scala
   spark.read.parquet("s3a://repro-a/pq-same", 
"s3a://repro-b/pq-same").collect()
   ```
   
   The plan is `CometNativeScan parquet ... InMemoryFileIndex(2 paths)`. Rows 
below are sorted after collecting.
   
   - **Same keys in both buckets** (`pq-same/f1.parquet`, 
`pq-same/f2.parquet`): Spark returns `[1,A] [1,B] [2,A] [2,B] ... [6,A] [6,B]`. 
Comet returns `[1,A] [1,A] [2,A] [2,A] ... [6,A] [6,A]`, with no error.
   - **Different keys** (`repro-a/pq-a/a1.parquet`, `repro-b/pq-b/b1.parquet`, 
...): the scan fails with `[FAILED_READ_FILE.FILE_NOT_EXIST] ... reading file 
pq-b/b1.parquet`. The message carries the key without the bucket, so it does 
not show that the wrong bucket was queried.
   - **One file per partition** (`openCostInBytes=128MB`): the results match 
Spark.
   
   ## Expected behavior
   
   The same rows as Spark, or a fallback to Spark when the native scan cannot 
serve more than one bucket.
   
   ## Additional context
   
   Reproduced on `main` at e038cf05f with the default Spark 4.1 profile. The 
CSV native scan has a related but broader problem, filed separately.
   
   There are two ways out. The scan could decline when its files span more than 
one bucket or authority, which extends the alias check to every scheme. Or 
native planning could keep each file's store, either by grouping a partition's 
files by store or by carrying the full URL into the file's location.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to