andygrove opened a new pull request, #6808:
URL: https://github.com/apache/datafusion-comet/pull/6808

   ## Which issue does this PR close?
   
   No issue. This follows up on #6441.
   
   ## Rationale for this change
   
   #6441 makes a native Iceberg write to S3 fall back to Spark when the Hadoop 
configuration has an `fs.s3a.*` setting the native writer doesn't support. It 
already ignores Hadoop's built-in defaults and a few S3A read settings that 
Spark seeds into every session. Some other settings are also common in S3 
deployments, and any one of them makes every native S3 write fall back:
   
   - Since Spark 4.1, `SparkContext` sets `fs.s3a.committer.magic.enabled=true` 
and `fs.s3a.committer.name=magic` for every application when 
`spark-hadoop-cloud` is on the classpath (SPARK-47618).
   - Many deployments also set `fs.s3a.committer.threads`, 
`fs.s3a.experimental.input.fadvise` (Hadoop's S3A docs recommend `random` for 
columnar formats) or `fs.s3a.bulk.delete.page.size` cluster-wide.
   
   The fallback reason looks like this:
   
   ```
   unsupported Hadoop S3A settings: fs.s3a.bulk.delete.page.size, 
fs.s3a.committer.magic.enabled, fs.s3a.committer.name, 
fs.s3a.committer.threads, fs.s3a.experimental.input.fadvise
   ```
   
   None of these settings affects an Iceberg data-file write:
   
   - The native writer doesn't go through S3A, so S3A's delete batching doesn't 
apply.
   - `fadvise` is a read hint.
   - Iceberg commits through table metadata, not through a Hadoop output 
committer.
   
   ## What changes are included in this PR?
   
   The five keys are added to `IgnoredHadoopS3Keys` in 
`CometIcebergNativeWrite`, alongside the Spark-seeded read settings, with a 
comment explaining why.
   
   ## How are these changes tested?
   
   There's a new test in `CometIcebergWriteDetectionSuite`. It checks that 
these settings produce no unsupported keys, and that a setting the native 
writer can't honor (`fs.s3a.encryption.algorithm`) is still reported next to 
them. The whole suite passes locally.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to