wangzhigang1999 opened a new pull request, #9715:
URL: https://github.com/apache/paimon/pull/9715

   ### Purpose
   
   Fixes https://github.com/apache/paimon/issues/9714.
   
   Concurrent PyPaimon writers to OSS can both pass the destination existence 
check and overwrite the same snapshot during temporary-file-and-rename 
publication. Both commits can return successfully while only one remains 
visible.
   
   Introduce a thin `OssFileIO(PyArrowFileIO)` that overrides atomic metadata 
creation with OSS `PutObject` and `x-oss-forbid-overwrite=true`, following 
Java's approach in #8228. `FileAlreadyExists` returns `False`; other OSS 
failures propagate with their cause. Ordinary Arrow/Jindo operations remain 
inherited. Route the factory, REST token refresh and ResolvingFileIO through 
the OSS implementation, and migrate internal OSS callers.
   
   Forward Java-aligned SSE options for these metadata PUTs, including KMS 
key/data encryption and the legacy algorithm fallback. Add `oss2>=2.18,<3` to 
the optional `oss` and `jindo` extras; OSS atomic writes require explicit 
endpoint/credential options. The README documents setup and the encryption 
scope.
   
   Conditional creation applies only to buckets that have never enabled 
versioning. Enabled/suspended/unknown states and `GetBucketVersioning` access 
denial fall back to the inherited write path with a warning, preserving legacy 
behavior without claiming concurrent-write protection or new SSE behavior for 
that fallback. Other query failures propagate. All writers must use conditional 
creation, and bucket versioning must remain disabled.
   
   ### Tests
   
   - After rebasing onto master `8f5ce6b84`: **242 passed** (208 related 
regression cases plus 34 OSS protocol cases), Flake8 and `git diff --check` 
passed. Local runtime: Python 3.12.6, PyArrow 19.0.1, oss2 2.19.1.
   - The same 208 regression cases passed against a clean baseline and the 
implementation. The 34 protocol cases also passed with the minimum oss2 2.18.0. 
They exercise real SDK HTTP requests, concurrent creation/content preservation, 
lost responses, error propagation, versioning fallback, refreshed STS headers, 
URI handling and SSE validation.
   - Real OSS, same settings per version: **64 independent processes × 300 
rounds**, 19,200 attempts with 64 KiB payloads. Baseline: 300 rounds with 
multiple successful creators and 7,545 overwritten successful writes. New: 
exactly one creator per round, zero overwritten successful writes.
   - Real append commits: **16 processes × 10 commits × 2,000 rows**, 
`commit.max-retries=64` in both runs. Baseline: 160 completed calls, 232,000 
rows and 116 snapshots. New: all 320,000 rows and 160 snapshots, no 
missing/duplicate rows. Lost-response recovery controls passed in both versions.
   - Real OSS ordinary FileIO operations passed for both factory and Resolving 
routes. Seven SSE scenarios passed with the change (AES256, legacy AES256, KMS, 
KMS+SM4, SM4, option precedence and an explicit OSS-managed KMS key); 
encryption headers were absent in the baseline. Catalog/Resolving metadata 
writes also produced encrypted schema/snapshot objects.
   
   The real OSS runs used a never-versioned bucket. Versioning/access-denied 
fallback and STS forwarding were validated locally; real Jindo, STS renewal, 
customer-managed KMS permissions and other object stores were not exercised. An 
extra Arrow read through ResolvingFileIO reproduced the same existing 
missing-`filesystem` AttributeError in both versions; reading the same 
committed data through the ordinary route succeeded. That reader issue is 
unchanged and outside this PR.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to