LuciferYang opened a new issue, #10265: URL: https://github.com/apache/paimon/issues/10265
### Search before asking - [X] I searched in the [issues](https://github.com/apache/paimon/issues) and found no similar issues. ### Paimon version master (1.5-SNAPSHOT) ### Compute Engine Flink (postpone-bucket write path). ### Minimal reproduce step 1. Create a postpone-bucket table and set `data-file.prefix` to a value that contains the `-s-` separator (for example `a-s-b`). 2. Write data so postpone-bucket data files are produced. Each file name is `{data-file.prefix}-u-{commitUser}-s-{writeId}-w-{uuid}-{count}`. 3. Read or compact the table, which routes splits by `getWriteId(fileName) % parallelism`. ### What doesn't meet your expectations? `PostponeBucketFileStoreWrite.getWriteId` recovered the write id by splitting on the first `-s-` in the file name. When `data-file.prefix` (or a custom commit user) already contains `-s-`, the first `-s-` is not the writer's separator, so the parse either throws `Data file name ... does not match the pattern` or returns a wrong write id. A wrong id then routes the split to the wrong reader under `getWriteId(...) % parallelism`. ### Anything else? The correct anchor is the `-w-` writer marker that always terminates the interpolated prefix, so the write id is the segment between the last `-s-` before `-w-` and `-w-` itself. ### Are you willing to submit a PR? - [X] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
