unikdahal commented on issue #6737: URL: https://github.com/apache/datafusion-comet/issues/6737#issuecomment-6026242683
Thanks @comphead, your diagram is right, and I worded it badly. iceberg-rust would still own the Iceberg semantics. I only meant: who owns the physical Parquet read. `IcebergTableScan` builds a `TableScan` and calls `to_arrow()`, with filters fixed into an Iceberg `Predicate` at plan time. Comet's `IcebergScanExec` also reads through `ArrowReader`. So neither gets DataFusion's ParquetSource runtime filter pruning (join/TopK), which is why #6641 currently needs equivalent support on the iceberg-rust ArrowReader path. By "pre-planned tasks" I meant the `FileScanTask`s Comet already gets from Iceberg Java. By "DataFusion-native source" I meant a `DataSource` behind `DataSourceExec`, so runtime filters have a standard place to land. So I agree with @mbutrovich: start with a `DataSource` that still reads through `ArrowReader`. Decoding through `ParquetSource` itself can be a later question. Deletes, especially equality deletes, are the hard part there. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
