dalelane opened a new pull request, #29379: URL: https://github.com/apache/flink/pull/29379
## What is the purpose of the change This PR adds `bin/minicluster.sh`, which runs one application in a MiniCluster: the JobManager and a TaskManager in a single JVM. Unlike `flink run -t local`, the process behaves like an Application Mode cluster. It keeps a fixed JobID across restarts, recovers from HA metadata, supports `execution.shutdown-on-application-finish` and `execution.submit-failed-job-on-application-error`, and exits with codes that a process supervisor can act on. This is part of [FLIP-611](https://cwiki.apache.org/confluence/spaces/FLINK/pages/451974246/FLIP-611+Running+Flink+jobs+in+MiniCluster+using+the+Kubernetes+Operator). The Kubernetes Operator's proposed `FlinkMiniCluster` resource will start pods with this script, in the same way that it uses `standalone-job.sh` today. ## Brief change log - Move the shared configuration and program-loading code out of `StandaloneApplicationClusterEntryPoint` into `ApplicationClusterEntryPointUtils` (refactor - no functional change) - Add `ApplicationModeMiniCluster` and `MiniClusterJobEntryPoint`, which takes the same arguments as `StandaloneApplicationClusterEntryPoint` - Add `bin/minicluster.sh` (always runs in the foreground, with application arguments after `--`), and a `minicluster` service in `flink-console.sh` - Exit the process on TaskManager fatal errors - Docs: a new *Application Mode in a MiniCluster* section in the standalone deployment overview ## Verifying this change - `MiniClusterJobEntryPointTest` and `ApplicationModeMiniClusterTest`: argument parsing, configuration precedence, and close and fatal-error handling. - `MiniClusterJobEntryPointITCase` starts the entry point as separate JVMs, so that it can send real SIGTERM and SIGKILL signals. It uses ZooKeeper HA and the configuration that the Operator uses for application deployments. It covers: - checkpoint recovery after SIGTERM, SIGKILL and repeated SIGKILLs - no second run of `main()` after a restart once the application has finished or failed - `--fromSavepoint` after stop-with-savepoint - non-zero exits on JobManager and TaskManager fatal errors, with HA data kept - exit codes with `shutdown-on-application-finish` - `local://` jars - a reference test that runs the same restart-after-finish scenario against a distributed standalone application cluster - Manually tested with a distribution built from this branch: - The same `WordCount.jar` produced identical output from `minicluster.sh`, `standalone-job.sh` + `taskmanager.sh`, and `flink run -t local`. - After `kill -9`, `StateMachineExample.jar` restarted with `--fromSavepoint` under the same JobID, and checkpoint numbering carried on from where it stopped. ## Does this pull request potentially affect one of the following parts: - Dependencies (does it add or upgrade a dependency): no (test scope only: `flink-runtime` test-jar and `curator-test`) - The public API, i.e., is any changed class annotated with `@Public(Evolving)`: no - The serializers: no - The runtime per-record code paths (performance sensitive): no - Anything that affects deployment or recovery: JobManager (and its components), Checkpointing, Kubernetes/Yarn, ZooKeeper: yes (a new deployment entry point; existing default behaviour is unchanged) - The S3 file system connector: no ## Documentation - Does this pull request introduce a new feature? yes - If yes, how is the feature documented? docs and JavaDocs --- ##### Was generative AI tooling used to co-author this PR? - [X] Yes (please specify the tool below) Generated-by: Claude Code (Claude Opus 5.5) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
