LuciferYang opened a new pull request, #10266:
URL: https://github.com/apache/paimon/pull/10266
### Purpose
`PostponeBucketFileStoreWrite.getWriteId` recovered the write id from a data
file name by splitting on the first `-s-`. The name is built as
`{data-file.prefix}-u-{commitUser}-s-{writeId}-w-{uuid}-{count}`, so when
`data-file.prefix` or the commit user contains `-s-`, the first `-s-` is not
the writer's separator: the parse either throws `Data file name ... does not
match the pattern` or returns a wrong write id, which then misroutes splits
under `getWriteId(...) % parallelism`.
This anchors the parse to the `-w-` writer marker that always terminates the
interpolated prefix. The write id is the digits between the last `-s-` before
`-w-` and `-w-`, which is unaffected by any `-s-` or `-w-` inside the prefix or
commit user (the uuid and count that follow `-w-` never contain `w`).
This closes #10265.
### Tests
- `PostponeBucketWriterTest.testGetWriteIdWithSeparatorInCommitUser` pins
that a name whose commit user contains `-s-` parses to the correct write id.
The old first-`-s-` split returned a wrong value or threw on the same input.
### API and Format
No.
### Documentation
No.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]