[
https://issues.apache.org/jira/browse/FLINK-40788?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124959#comment-18124959
]
Martijn Visser commented on FLINK-40788:
----------------------------------------
Failed again on the last two master nightlies, both in the Java 11 e2e_4 leg:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79879&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79930&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e
This time only Metaspace: after the test's artificial failure one TaskManager
runs out of Metaspace (64m) and dies, and the remaining four cannot provide the
100 slots.
{code}
2026-10-07 01:04:01,860 ERROR org.apache.flink.util.FatalExitExceptionHandler
[] - FATAL: Thread 'flink-pekko.remote.default-remote-dispatcher-11' produced
an uncaught exception. Stopping the process...
java.lang.OutOfMemoryError: Metaspace
...
Caused by:
org.apache.flink.runtime.jobmanager.scheduler.NoResourceAvailableException:
Could not acquire the minimum required resources.
[FAIL] 'Heavy deployment end-to-end test' failed after 5 minutes and 13
seconds! Test exited with exit code 1
{code}
The GHA Java 11 nightlies on the same two commits passed this test.
> 'Heavy deployment end-to-end test' fails because all TaskManagers run out of
> memory on Java 11
> ----------------------------------------------------------------------------------------------
>
> Key: FLINK-40788
> URL: https://issues.apache.org/jira/browse/FLINK-40788
> Project: Flink
> Issue Type: Bug
> Components: Tests
> Affects Versions: 2.4.0
> Reporter: Martijn Visser
> Priority: Major
>
> The heavy deployment end-to-end test failed in the Java 11 nightly on master:
> https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79325&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e
> {code}
> Caused by: org.apache.flink.runtime.JobException: Recovery is suppressed by
> FixedDelayRestartBackoffTimeStrategy(maxNumberRestartAttempts=3,
> backoffTimeMS=0)
> ...
> Caused by:
> org.apache.flink.runtime.jobmanager.scheduler.NoResourceAvailableException:
> Could not acquire the minimum required resources.
> [FAIL] 'Heavy deployment end-to-end test' failed after 5 minutes and 12
> seconds! Test exited with exit code 1
> {code}
> All five TaskManagers logged an OutOfMemoryError during deployment, three
> "Java heap space" and two "Metaspace", before the job gave up. FLINK-16456
> had the same symptom on Java 11 in 2020.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)