MartijnVisser opened a new pull request, #29338: URL: https://github.com/apache/flink/pull/29338
Unchanged backport of https://github.com/apache/flink/pull/27624 (b5fbf2b4aca, 45c62cacc0a), which went to master and release-2.3 but not to this branch. Replaces #28727. When an upstream task fails, the consumer's remote channel receives the producer's error. If the consumer still holds a partially read buffer, releasing its deserializer during cleanup rethrows that error, and since FLINK-38180 a failure in that cleanup shuts down the TaskManager. The fix logs a warning instead when the channel already had a transport error. On release-2.2 this fails `EventTimeWindowCheckpointingITCase` and `LocalRecoveryITCase` in about half the nightlies, the last one on 2026-09-26 (`test_ci tests`): https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79462&view=logs&j=5c8e7682-d68f-54d1-16a2-a09310218a49 Verified locally on JDK 17: both ITCases pass, and `testCloseIgnoresReleaseFailureFromChannelWithError` fails with the production change reverted. With the channel error forced at cleanup, `EventTimeWindowCheckpointingITCase` lost a TaskManager in 2 of 12 cases on release-2.2 and in 0 of 24 with this change (the new warning was logged 4 times). --- ##### Was generative AI tooling used to co-author this PR? - [X] Yes (please specify the tool below) Generated-by: Claude Code (Claude Opus 5.5) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
