[ 
https://issues.apache.org/jira/browse/FLINK-40788?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18124959#comment-18124959
 ] 

Martijn Visser commented on FLINK-40788:
----------------------------------------

Failed again on the last two master nightlies, both in the Java 11 e2e_4 leg:
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79879&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e
https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79930&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e

This time only Metaspace: after the test's artificial failure one TaskManager 
runs out of Metaspace (64m) and dies, and the remaining four cannot provide the 
100 slots.

{code}
2026-10-07 01:04:01,860 ERROR org.apache.flink.util.FatalExitExceptionHandler 
[] - FATAL: Thread 'flink-pekko.remote.default-remote-dispatcher-11' produced 
an uncaught exception. Stopping the process...
java.lang.OutOfMemoryError: Metaspace
...
Caused by: 
org.apache.flink.runtime.jobmanager.scheduler.NoResourceAvailableException: 
Could not acquire the minimum required resources.
[FAIL] 'Heavy deployment end-to-end test' failed after 5 minutes and 13 
seconds! Test exited with exit code 1
{code}

The GHA Java 11 nightlies on the same two commits passed this test.

> 'Heavy deployment end-to-end test' fails because all TaskManagers run out of 
> memory on Java 11
> ----------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40788
>                 URL: https://issues.apache.org/jira/browse/FLINK-40788
>             Project: Flink
>          Issue Type: Bug
>          Components: Tests
>    Affects Versions: 2.4.0
>            Reporter: Martijn Visser
>            Priority: Major
>
> The heavy deployment end-to-end test failed in the Java 11 nightly on master: 
> https://dev.azure.com/apache-flink/98463496-1af2-4620-8eab-a2ecc1a2e6fe/_build/results?buildId=79325&view=logs&j=320b3f94-fe78-5d7f-e965-9db30de9270e
> {code}
> Caused by: org.apache.flink.runtime.JobException: Recovery is suppressed by 
> FixedDelayRestartBackoffTimeStrategy(maxNumberRestartAttempts=3, 
> backoffTimeMS=0)
> ...
> Caused by: 
> org.apache.flink.runtime.jobmanager.scheduler.NoResourceAvailableException: 
> Could not acquire the minimum required resources.
> [FAIL] 'Heavy deployment end-to-end test' failed after 5 minutes and 12 
> seconds! Test exited with exit code 1
> {code}
> All five TaskManagers logged an OutOfMemoryError during deployment, three 
> "Java heap space" and two "Metaspace", before the job gave up. FLINK-16456 
> had the same symptom on Java 11 in 2020.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to