[ https://issues.apache.org/jira/browse/BEAM-6777?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16905654#comment-16905654 ]
Yueyang Qiu commented on BEAM-6777: ----------------------------------- Hi Oded, Dataflow has implemented a solution to this issue but I don't think the change is fully rolled out to production yet. > SDK Harness Resilience > ---------------------- > > Key: BEAM-6777 > URL: https://issues.apache.org/jira/browse/BEAM-6777 > Project: Beam > Issue Type: Improvement > Components: runner-dataflow > Reporter: Sam Rohde > Assignee: Yueyang Qiu > Priority: Major > Time Spent: 7h 20m > Remaining Estimate: 0h > > If the Python SDK Harness crashes in any way (user code exception, OOM, etc) > the job will hang and waste resources. The fix is to add a daemon in the SDK > Harness and Runner Harness to communicate with Dataflow to restart the VM > when stuckness is detected. -- This message was sent by Atlassian JIRA (v7.6.14#76016)