JingsongLi commented on PR #10295: URL: https://github.com/apache/paimon/pull/10295#issuecomment-5935709744
Thanks for the patch. I reviewed the production bootstrap paths and ran the core regression suite (`GlobalIndexAssignerTest`: 10 passing tests, with normal Maven checks). `FlinkSinkTest` also passed all 12 tests, including the real operator-commit unaligned-checkpoint rejection and coordinator primary-key rejection. The advertised trigger is not reachable through the currently supported Flink sink: `GlobalDynamicBucketSink.build()` calls `sinkFrom()` → `FlinkSink.doCommit()`, whose streaming configuration check rejects unaligned checkpoints. This is also true in the merge-base. The coordinator-commit exception cannot enable this global-primary-key topology: it requires an append table without primary keys and `BUCKET_UNAWARE`. Spark consumes all bootstrap records before calling `endBoostrapWithoutEmit`, and bounded/no-checkpoint execution has no checkpoint barrier that can overtake these records. The added tests manually end bootstrap before submitting a `KEY_PART`; they establish helper behavior, but do not establish a supported job that encounters the reported failure. Also, allowing that ordering before fixing the operator/checkpoint protocol leaves the acknowledged same-key cross-partition retraction gap. Closing because this change does not currently resolve an end-to-end problem in a supported production path. A reproducer through a supported Flink topology, together with a bootstrap/checkpoint ordering solution that preserves global primary-key uniqueness, would provide a concrete basis to reopen this work. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
