andygrove opened a new issue, #5690:
URL: https://github.com/apache/datafusion-comet/issues/5690

   ### Describe the bug
   
   The native (iceberg-rust) writer rolls to a new data file at a different 
point than iceberg-java for the same `write.target-file-size-bytes`, producing 
consistently half as many files.
   
   `TestSparkDataWrite` in Iceberg's `spark` suite, with 
`spark.comet.iceberg.write.enabled=true`:
   
   - `testUnpartitionedCreateWithTargetFileSizeViaTableProperties`: expected 4 
data files, got 2
   - `testPartitionedCreateWithTargetFileSizeViaOption`: expected 8 data files, 
got 4
   
   Both fail identically for `branch = null`, `branch = main` and `branch = 
testBranch`. The written data is correct; only the file count differs. Each 
produced file holds 2000 records where iceberg-java produced two files of 1000.
   
   The same accounting difference propagates into compaction, where bin-packing 
decisions depend on the sizes of the files the writer produced. In the same run:
   
   - `TestRewriteDataFilesAction.testBinPackCombineMediumFiles` (formatVersion 
= 2): "Action should add 3 data files", expected 3 but was 4
   - 
`TestRewriteDataFilesAction.testRemoveDangledPositionDeletesPartitionEvolution` 
(formatVersion = 2): expected `[1, 2, 1]`, got `[2, 2, 0]`
   
   ### Steps to reproduce
   
   Spark 4.1.3, Iceberg 1.11.0, with `spark.comet.iceberg.write.enabled=true` 
and the Iceberg Spark SQL test setup from `dev/diffs/iceberg/1.11.0.diff`:
   
   ```
   ./gradlew -DsparkVersions=4.1 -DscalaVersion=2.13 -DflinkVersions= 
-DkafkaVersions= \
     :iceberg-spark:iceberg-spark-4.1_2.13:test --tests '*TestSparkDataWrite*' 
-Pquick=true -x javadoc
   ```
   
   ### Expected behavior
   
   For a given `write.target-file-size-bytes`, the native writer rolls to a new 
file at the same point as iceberg-java, so file counts and sizes match what the 
JVM writer would have produced.
   
   ### Additional context
   
   Found by turning the two Iceberg write flags on by default in #5677.
   
   Comet's own `CometIcebergWriteActionSuite` has a "target-file-size rolls one 
task across multiple files" case that passes, so the writer does roll. It just 
does not roll at the same threshold, which suggests the size estimate compared 
against the target differs (for example compressed versus in-memory size, or 
the point in the row-group lifecycle at which the check happens) rather than 
the rolling logic being absent.
   
   Part of #5649.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to