[
https://issues.apache.org/jira/browse/BEAM-12792?focusedWorklogId=754835&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-754835
]
ASF GitHub Bot logged work on BEAM-12792:
-----------------------------------------
Author: ASF GitHub Bot
Created on: 08/Apr/22 22:14
Start Date: 08/Apr/22 22:14
Worklog Time Spent: 10m
Work Description: tvalentyn commented on PR #16658:
URL: https://github.com/apache/beam/pull/16658#issuecomment-1093407575
Here's a repro without Dataflow in the picture:
```
gcloud compute instances create valentyn-cos2 --zone=us-central1-f
--project=<my gcp project> --image-family cos-stable --image-project=cos-cloud
--restart-on-failure
gcloud compute ssh valentyn-cos2 --zone=us-central1-f --project=<my gcp
project>
docker run -it --entrypoint=/bin/bash -v /tmp:/var/opt/google
python:3.7-bullseye
python -m venv /var/opt/google/env
root@018e7f3cac24:/# /var/opt/google/env/bin/pip
bash: /var/opt/google/env/bin/pip: Permission denied
```
Issue Time Tracking
-------------------
Worklog Id: (was: 754835)
Time Spent: 16h 40m (was: 16.5h)
> Multiple jobs running on Flink session cluster reuse the persistent Python
> environment.
> ---------------------------------------------------------------------------------------
>
> Key: BEAM-12792
> URL: https://issues.apache.org/jira/browse/BEAM-12792
> Project: Beam
> Issue Type: Bug
> Components: sdk-py-harness
> Affects Versions: 2.27.0, 2.28.0, 2.29.0, 2.30.0, 2.31.0
> Environment: Kubernetes 1.20 on Ubuntu 18.04.
> Reporter: Jens Wiren
> Priority: P1
> Labels: FlinkRunner, beam
> Time Spent: 16h 40m
> Remaining Estimate: 0h
>
> I'm running TFX pipelines on a Flink cluster using Beam in k8s. However,
> extra python packages passed to the Flink runner (or rather beam worker
> side-car) are only installed once per deployment cycle. Example:
> # Flink is deployed and is up and running
> # A TFX pipeline starts, submits a job to Flink along with a python whl of
> custom code and beam ops.
> # The beam worker installs the package and the pipeline finishes succesfully.
> # A new TFX pipeline is build where a new beam fn is introduced, the pipline
> is started and the new whl is submitted as in step 2).
> # This time, the new package is not being installed in the beam worker
> causing the job to fail due to a reference which does not exist in the beam
> worker, since it didn't install the new package.
>
> I started using Flink from beam version 2.27 and it has been an issue all the
> time.
--
This message was sent by Atlassian Jira
(v8.20.1#820001)