andygrove commented on PR #6664: URL: https://github.com/apache/datafusion-comet/pull/6664#issuecomment-6060918310
@viirya `CometIcebergWriteBenchmark` for the four arms, at 1cb149f6e on an Apple M3 Ultra (local filesystem, 4M rows, `local[5]`, average of five runs): | Case | Spark | Comet scan, iceberg-java writer | Comet scan, native writer | Native vs iceberg-java writer | | --- | --: | --: | --: | --: | | Unpartitioned `INSERT` | 1426 ms | 1250 ms | 606 ms | 2.1x | | Partitioned `INSERT`, clustered writer | 1948 ms | 1496 ms | 1040 ms | 1.4x | | Partitioned `INSERT`, fanout writer | 1922 ms | 1787 ms | 951 ms | 1.9x | | Copy-on-write `DELETE` | 2293 ms | 2116 ms | 1559 ms | 1.4x | The native writer is faster than iceberg-java's in all four. On the footer read: for every file it writes natively, the task reads the file's footer back before it returns its commit message, so that iceberg-java's own code computes the file's metrics. Through `S3FileIO` that is two ranged GETs per file, the 8-byte tail and then the footer, made one file after another. iceberg-java takes the same metrics from the footer it still holds in memory. A task that writes a few large files hardly notices, but a fanout task that writes many small files pays two request round trips for each of them. The user guide, under Native Parquet write eligibility, and the 1.2.0 upgrade entry now say so, and #6772 tracks removing the reads by handing each footer from the native writer to the JVM. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
