dilverse commented on issue #17193: URL: https://github.com/apache/iceberg/issues/17193#issuecomment-5857623730
We reproduced this on stock Iceberg 1.11.0 (Kafka Connect 4.3.1, classic consumer protocol, default sink settings) and can confirm the mechanism, including which group fails. **The failing group** is the Worker's `cg-control-<uuid>`, not the Connect or `-coord` group. **Why the background heartbeat does not help:** the classic consumer starts its heartbeat thread only after it has handled its JoinGroup response inside `poll()` (`AbstractCoordinator`: disabled while the member has not joined). The Worker is created on the first non-empty `put()`, `Channel.start()` polls for only 1s (a new group's first join waits `group.initial.rebalance.delay.ms`, 3s by default), and afterwards the control consumer is polled only with `poll(Duration.ZERO)` from `put()`/`flush()`. On an idle topic Connect calls `put()` once per `offset.flush.interval.ms` (60s), so the JoinGroup response is handled about 60s later, after the 45s session has expired: `UNKNOWN_MEMBER_ID`, rejoin, forever. A member that has already joined keeps its session through the heartbeat thread, so only a join that starts just before the topic goes idle is affected (a small snapshot, a low-volume topic). **Repro:** two connectors side by side, schemaless JSON, `iceberg.control.commit.interval-ms` left at its default. - Idle (3 records, then nothing): JoinGroup sent at 04:23:40, response handled at 04:24:36, then `SyncGroup failed: The coordinator is not aware of this member`. 9 commit rounds over 45 minutes all ended "committed to 0 table(s)". - Busy (1 record every 5s): joined in 5s and committed every round. **Not a 1.11 regression:** 1.10.1 has the same lazy Worker creation and the same 1s poll in `Channel.start()`. #14395 changed only leader election, which matches #11818 seeing this on 1.7.1. **Fix we are running:** `Worker.start()` polls until the control consumer has its assignment, bounded by `iceberg.control.commit.timeout-ms`; after that the heartbeat thread holds the membership. With only this change and default settings, the stuck idle connector joined within 3s and committed its 3 records within the next rounds. The patch and a `TestWorker` case (fails on 1.11.0, passes with the patch) apply cleanly to `main`; we can open a PR. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
