Dennis-Mircea Ciupitu created FLINK-40581:
---------------------------------------------

             Summary: Fix examples that cannot run against their declared Flink 
version
                 Key: FLINK-40581
                 URL: https://issues.apache.org/jira/browse/FLINK-40581
             Project: Flink
          Issue Type: Improvement
          Components: Kubernetes Operator
            Reporter: Dennis-Mircea Ciupitu


Currently, on 1.16.0 operator release, 4 of the shipped examples do not work 
against the Flink version they declare.
h2. flink-beam-example

This example has a Dockerfile that is still on {{flink:1.16,}} while its 
Kubernetes YAML declares {{{}v2_2{}}}, so the operator writes a {{config.yaml}} 
that a 1.16 entrypoint cannot read and the JobManager crash loops on {{{}"The 
Flink config file '/opt/flink/conf/flink-conf.yaml' does not exist"{}}}. This 
example also has a different version on the Beam runner itself 
({{{}beam-runners-flink-1.19{}}}), creating a mixture of 3 different Flink 
versions that breaks the stability of the example.

Besides fixing the versions, there are also another Beam-related problems:
 * The example doesn't integrate with the operator's {{local:}} scheme 
({{{}"unknown protocol: local"{}}}).
 * The example has a malformed {{--output}} argument.
 * The example should run in a detached mode.

h2. flink-sql-runner-example

This example is also shipped with {{{}flink:1.16{}}}, which wraps Java 11 while 
the jar is built with Java 17, so the container dies on 
{{UnsupportedClassVersionError}} before Flink starts. The declared {{v1_20}} 
does not match the image either.
h2. hpa/basic-hpa.yaml

This example has the {{scaleTargetRef}} that references 
{{{}FlinkDeployment/basic{}}}, which no example creates. Renaming it is not 
enough, because the scale subresource maps to 
{{{}spec.taskManager.replicas{}}}, so the target has to declare that field, and 
it only has relevance in standalone mode.
h2. snapshot savepoint examples

The {{savepoint-job-defaults.yaml}} and {{savepoint-job-full-spec.yaml}} 
examples always fail on Flink 2.0 to 2.3 versions, while the checkpoint example 
passes, because {{SavepointSpec.formatType}} defaults to {{{}CANONICAL{}}}, and 
the {{StateMachineExample}} job enables async state, which cannot produce 
canonical savepoints.

This is Flink-related behaviour, not an operator bug, because the same request 
against Flink's REST endpoint fails identically. FLINK-39061 removed the async 
state usage on the master branch, but versions 2.0 to 2.3 still have it. The 
Flink 1.20 version is fine, and only savepoints of a running job are affected, 
not the stop-with-savepoint used by upgrades.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to