andygrove opened a new pull request, #6306:
URL: https://github.com/apache/datafusion-comet/pull/6306

   Backport of #6106 to `branch-1.1`.
   
   Cherry-picked from `786bbc8fd630abfd0b29bdaffea6168f9f228863` without 
conflicts. The five files it changes are identical on `branch-1.1` and on 
`main` just before #6106, so the diff is byte-identical to upstream.
   
   ## Which issue does this PR close?
   
   Closes #6105 on `branch-1.1`. #6106 already closed it on `main`.
   
   ## Rationale for this change
   
   Every native Iceberg scan or write task builds its own `FileIO`. #6105 
counted 40,317 of them for a single query on a production workload. Each one 
rebuilds the storage factory and its parsed configuration, and for S3 with a 
custom access provider, the JNI bridge. #6106 merged after `branch-1.1` was cut 
at 36ab57c68, so 1.1.0 would ship without it.
   
   ## What changes are included in this PR?
   
   The original change, so see #6106 for the details. No adaptations were 
needed. In short:
   
   - `load_file_io` serves clones from a per-executor LRU cache of 64 entries 
and builds a `FileIO` only on a miss.
   - The cache key holds everything that shapes the client: the access mode, 
the catalog name, the reference path and the catalog properties.
   - `memory:///` is never cached. Neither is a read whose S3 access provider 
failed to initialize, so the next task retries the provider.
   - `release_runtime` drains the cache.
   
   ## How are these changes tested?
   
   Run locally on `branch-1.1` with the default profile (Spark 4.1, Scala 2.13, 
JDK 17), against a `libcomet` built from this branch:
   
   - The six new tests in `iceberg_common.rs` pass, along with the rest of the 
core crate's lib tests, 503 in all.
   - `cargo fmt --all -- --check` and `cargo clippy --all-targets --workspace 
-- -D warnings` pass on rustc 1.97.0.
   - `CometIcebergNativeSuite` and `CometIcebergWriteActionSuite` pass: 182 
tests, and 1 is canceled by its Spark 4.1 `assume` (SPARK-55626).
   
   #6106 was also verified with the MinIO-backed `IcebergReadFromS3Suite` and 
the manual `CometS3CredentialBridgeSuite`, which no CI job runs. I did not 
re-run them here. With this PR, the branch's native tree and its Java S3 SPI 
are byte-identical to `main` at #6106, apart from #6181's opt-in local TopK 
fusion in `planner.rs` and #5322's `contains` kernel. So those runs covered the 
code this PR ships.
   
   Against `branch-1.1`, the changed paths route this pull request to every 
suite except Spark 3.4's SQL job and the benchmark check. That includes every 
Spark profile, macOS and all four Iceberg versions.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to