wangmingzhou1986 commented on PR #4370: URL: https://github.com/apache/flink-cdc/pull/4370#issuecomment-6077443289
Hi @dangerousfeng, thanks for this fix — could you please reopen this PR? We hit this problem in production with Flink CDC 3.6.0 (MySQL → Paimon pipeline, application mode on Kubernetes, `schema.change.behavior: TRY_EVOLVE`). When a job restarts while ADD COLUMN DDL is still inside the binlog range being replayed, the same column is appended to the evolved schema again on every restart (the column count grows 16 → 17 → 18 …), until the job fails with `java.lang.IllegalStateException: Duplicate key ...` (same stack as FLINK-37537) and cannot recover without manual intervention. Making `SchemaUtils.applyAddColumnEvent` idempotent, as this PR does, looks like the right fix for the duplicated columns, since FLINK-37537 / FLINK-38830 / this issue all end up appending an already-existing column through this path, while FLINK-37710 explains why the replay happens in the first place. We are going to verify the patch against our reproduction and will report the results here. If you no longer have time for it, we would be happy to continue it in a new PR (crediting you as the original author). cc @lvyanquan (requested reviewer) Related: FLINK-39412, FLINK-37537, FLINK-38830, FLINK-37710 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
