LuciferYang opened a new pull request, #10266:
URL: https://github.com/apache/paimon/pull/10266

   ### Purpose
   
   `PostponeBucketFileStoreWrite.getWriteId` recovered the write id from a data 
file name by splitting on the first `-s-`. The name is built as 
`{data-file.prefix}-u-{commitUser}-s-{writeId}-w-{uuid}-{count}`, so when 
`data-file.prefix` or the commit user contains `-s-`, the first `-s-` is not 
the writer's separator: the parse either throws `Data file name ... does not 
match the pattern` or returns a wrong write id, which then misroutes splits 
under `getWriteId(...) % parallelism`.
   
   This anchors the parse to the `-w-` writer marker that always terminates the 
interpolated prefix. The write id is the digits between the last `-s-` before 
`-w-` and `-w-`, which is unaffected by any `-s-` or `-w-` inside the prefix or 
commit user (the uuid and count that follow `-w-` never contain `w`).
   
   This closes #10265.
   
   ### Tests
   
   - `PostponeBucketWriterTest.testGetWriteIdWithSeparatorInCommitUser` pins 
that a name whose commit user contains `-s-` parses to the correct write id. 
The old first-`-s-` split returned a wrong value or threw on the same input.
   
   ### API and Format
   
   No.
   
   ### Documentation
   
   No.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to