dilverse commented on issue #17193:
URL: https://github.com/apache/iceberg/issues/17193#issuecomment-5857623730

   We reproduced this on stock Iceberg 1.11.0 (Kafka Connect 4.3.1, classic 
consumer protocol, default sink settings) and can confirm the mechanism, 
including which group fails.
   
   **The failing group** is the Worker's `cg-control-<uuid>`, not the Connect 
or `-coord` group.
   
   **Why the background heartbeat does not help:** the classic consumer starts 
its heartbeat thread only after it has handled its JoinGroup response inside 
`poll()` (`AbstractCoordinator`: disabled while the member has not joined). The 
Worker is created on the first non-empty `put()`, `Channel.start()` polls for 
only 1s (a new group's first join waits `group.initial.rebalance.delay.ms`, 3s 
by default), and afterwards the control consumer is polled only with 
`poll(Duration.ZERO)` from `put()`/`flush()`. On an idle topic Connect calls 
`put()` once per `offset.flush.interval.ms` (60s), so the JoinGroup response is 
handled about 60s later, after the 45s session has expired: 
`UNKNOWN_MEMBER_ID`, rejoin, forever. A member that has already joined keeps 
its session through the heartbeat thread, so only a join that starts just 
before the topic goes idle is affected (a small snapshot, a low-volume topic).
   
   **Repro:** two connectors side by side, schemaless JSON, 
`iceberg.control.commit.interval-ms` left at its default.
   
   - Idle (3 records, then nothing): JoinGroup sent at 04:23:40, response 
handled at 04:24:36, then `SyncGroup failed: The coordinator is not aware of 
this member`. 9 commit rounds over 45 minutes all ended "committed to 0 
table(s)".
   - Busy (1 record every 5s): joined in 5s and committed every round.
   
   **Not a 1.11 regression:** 1.10.1 has the same lazy Worker creation and the 
same 1s poll in `Channel.start()`. #14395 changed only leader election, which 
matches #11818 seeing this on 1.7.1.
   
   **Fix we are running:** `Worker.start()` polls until the control consumer 
has its assignment, bounded by `iceberg.control.commit.timeout-ms`; after that 
the heartbeat thread holds the membership. With only this change and default 
settings, the stuck idle connector joined within 3s and committed its 3 records 
within the next rounds. The patch and a `TestWorker` case (fails on 1.11.0, 
passes with the patch) apply cleanly to `main`; we can open a PR.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to