[ 
https://issues.apache.org/jira/browse/KAFKA-21132?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18117291#comment-18117291
 ] 

Ojasva Jain commented on KAFKA-21132:
-------------------------------------

If the community agrees this is worth fixing, I'll be happy to discuss 
approaches and work on it.

> Producer close() hangs indefinitely if the in-flight transaction abort never 
> succeeds
> -------------------------------------------------------------------------------------
>
>                 Key: KAFKA-21132
>                 URL: https://issues.apache.org/jira/browse/KAFKA-21132
>             Project: Kafka
>          Issue Type: Bug
>          Components: clients, producer 
>            Reporter: Ojasva Jain
>            Assignee: Ojasva Jain
>            Priority: Minor
>         Attachments: AbortCloseBlockingDemoTest.java
>
>
> h2. Problem
> {{KafkaProducer#close()}} (no arguments) is documented to close the producer 
> and abort any in-flight transaction. If the transaction coordinator is 
> unreachable, or keeps returning retriable errors while the abort is 
> attempted, this call never returns and there is no bound on how long it will 
> wait.
> h2. Root cause
> {{KafkaProducer.close()}} delegates to 
> {{close(Duration.ofMillis(Long.MAX_VALUE))}}. That method only force-closes 
> the sender ({{sender.forceClose()}}) if the I/O thread is still alive after 
> waiting for the given timeout. With an effectively-infinite timeout, that 
> check never fires, so nothing ever interrupts the shutdown.
> On the I/O thread, {{Sender.run()}}'s shutdown path calls 
> {{transactionManager.beginAbort()}} once, which enqueues a single 
> {{EndTxnHandler}}, then loops calling {{runOnce()}} to send it. If the 
> coordinator
>   is unreachable or returns a retriable error, the handler's {{onComplete()}} 
> callback just re-enqueues itself ({{reenqueue()}}) with no retry-count or 
> elapsed-time cap. The loop only exits when the abort actually succeeds or 
> {{forceClose}} is set and per above, that never happens for no-arg 
> {{close()}}. So the loop, and {{close()}}, retry forever.
> Note: {{close(Duration timeout)}} with a real, finite timeout is unaffected 
> h2. Repro
> Attached {{{}AbortCloseBlockingDemoTest{}}}. It uses a {{MockClient}}, which 
> answers every {{{}EndTxn{}}}(abort) attempt with {{CONCURRENT_TRANSACTIONS}} 
> (a retriable error) for many attempts in a row. The test shows 
> {{producer.close()}} still blocked after several seconds and 15+ retries, 
> only returning once the mocked coordinator finally accepts the abort.
> h2. Expectation
>  Every transaction already has a bounded lifetime on the broker. After 
> {{transaction.timeout.ms}}, the coordinator expires it regardless of what the 
> client does. Given that, the client-side abort attempted during {{close()}} 
> should never need to wait longer than {{transaction.timeout.ms}}, if the 
> abort hasn't completed by then, the broker will have already discarded the 
> transaction on its own, so there is nothing left to wait for. {{close()}} 
> should not block past that point.
> h2. Impact
> Any caller of {{close()}} without args (a very common pattern) can hang 
> forever if the transaction coordinator is unreachable or persistently 
> erroring at shutdown time.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to