unikdahal opened a new issue, #6737:
URL: https://github.com/apache/datafusion-comet/issues/6737

   I've been working on runtime pruning for `IcebergScanExec` (#6588, #6641, 
follow-up to #4807/#4808). It works. On the benchmark in #6641, the sorted join 
and TopK cases cut reader bytes by 94-99%. But getting there meant adding a 
fair amount of adaptive-pruning machinery to iceberg-rust (tracked in 
apache/iceberg-rust#3343). That includes a runtime predicate provider, 
rejecting file tasks from retained stats, live row-group refresh, and keeping 
position/equality deletes correct through all of it.
    
   While writing it I kept noticing that DataFusion's Parquet path already does 
most of this for plain Parquet scans. apache/datafusion#22450 (closing #22407) 
re-checks a live `DynamicFilterPhysicalExpr` at row-group boundaries, and the 
sort pushdown epic (apache/datafusion#23036) covers stats-based file/RG 
ordering for TopK. So for the same dynamic filters we end up with two parallel 
stacks:
    
   ```
   native Parquet scan:  DynamicFilterPhysicalExpr -> ParquetSource -> file / 
RG / page pruning
   Iceberg scan:         DynamicFilterPhysicalExpr -> translate -> iceberg-rust 
runtime predicate
                         -> ArrowReader -> file / RG / page pruning 
(reimplemented)
   ```
    
   To be clear, I'm not proposing to rip anything out, and #6641 stands as is. 
I also don't want Comet re-planning tables in Rust. Iceberg Java should keep 
producing the `FileScanTask`s.
    
   What I'm wondering about is `apache/datafusion-iceberg`, now that it lives 
under DataFusion. Long term, could it expose a source that takes 
already-planned `FileScanTask`s and runs them through 
`DataSourceExec`/`ParquetSource`, with the Iceberg-specific parts wrapped 
around it (field-ID schema mapping, defaults, position/equality deletes, DVs, 
encryption)?
    
   ```
   Spark / Iceberg Java -> FileScanTask[] -> Comet -> datafusion-iceberg
                        -> DataSourceExec / ParquetSource -> Parquet
   ```
    
   Comet would hand it the Java-planned tasks, and dynamic filters, TopK 
ordering and whatever DataFusion adds next would come for free instead of being 
redone for Iceberg.
    
   I know the hard part is semantics. Position deletes need physical row 
positions before any filtering, equality deletes need extra columns pulled into 
the internal projection, and then there are DVs, split tasks and schema 
evolution. `ArrowReader` handles all of that today and Comet relies on it. 
There's also a smaller practical point in #6126: for the native Parquet scan 
Comet owns the `AsyncFileReader`, but for Iceberg it's iceberg-rust's 
`ArrowFileReader` and we can't swap it, which makes things like memory 
accounting harder.
    
   ## Questions
    
   1. Is this a direction people would want, or is `ArrowReader` meant to stay 
the physical execution path for Iceberg in Comet?
   2. If it is, is datafusion-iceberg the right home, and would a "from 
pre-planned tasks" entry point be reasonable there?
   3. Meanwhile, is it fine to keep going with #6641 in its current shape, 
knowing parts of it might be replaced later? I mainly want to decide how much 
more adaptive logic (live refresh, ordering) to put into iceberg-rust.
   I haven't dug into datafusion-iceberg's internals yet, so some of this may 
already exist or be planned.
    


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to