[ 
https://issues.apache.org/jira/browse/FLINK-40627?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18113905#comment-18113905
 ] 

Dale Lane commented on FLINK-40627:
-----------------------------------

After seeing this happen once accidentally, I've been able to manually 
reproduce it against a local test HTTP server configured to serve half a jar on 
the first request and the full jar on subsequent requests. I tried a few 
different scenarios and in each case the Job Manager start up fails as 
described above. 

> Partially fetched user artifacts are reused on subsequent starts
> ----------------------------------------------------------------
>
>                 Key: FLINK-40627
>                 URL: https://issues.apache.org/jira/browse/FLINK-40627
>             Project: Flink
>          Issue Type: Bug
>          Components: Client / Job Submission
>            Reporter: Dale Lane
>            Priority: Minor
>
> {{ArtifactFetchManager.fetchArtifact}} returns any file already present at 
> the target path without re-fetching ("Already fetched user artifacts are 
> kept"). {{HttpArtifactFetcher}} and {{FsArtifactFetcher}} write directly to 
> that final path via {{{}FileUtils.copyToFile{}}}, so if a transfer ends part 
> way through, a truncated file is left at the final path. Nothing cleans up 
> files from the artifacts directory.
> A transfer can end part way through in several ways, all of which leave a 
> truncated file:
>  * fetch fails with an exception (e.g. connection reset)
>  * Job Manager process is killed mid-transfer (e.g. OOMKilled, liveness-probe 
> kill in Kubernetes)
>  * an HTTP server without a {{Content-Length}} ends the response early (the 
> fetch then reports success, and the job fails on its first start as well as 
> every later one)
> Any later start that sees the same directory reuses the truncated file and 
> never retries the download. The job fails on every start with no fetch 
> traffic:
> {{org.apache.flink.client.program.ProgramInvocationException: Error while 
> opening jar file '.../job.jar'}}
> {{Caused by: java.io.IOException: Error while opening jar file '.../job.jar'}}
> {{Caused by: java.util.zip.ZipException: zip END header not found}}
> This affects:
>  * Native Kubernetes Application Mode: the artifacts dir 
> ({{{}<user.artifacts.base-dir>/<namespace>/<cluster-id>{}}}) is an 
> {{{}emptyDir{}}}, which survives Job Manager container restarts within a pod, 
> so the pod is stuck until it is deleted. With a persistent volume or 
> {{hostPath}} mounted at {{{}base-dir{}}}, it is stuck permanently.
>  * Standalone Application Mode with a persistent 
> {{{}user.artifacts.base-dir{}}}.
> Clearing it today needs manual deletion of the file (or pod deletion for 
> {{{}emptyDir{}}}).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to