The GitHub Actions job "Tests (AMD)" on airflow.git/fix-ssh-remote-job-pty-race 
has failed.
Run started by GitHub user potiuk (triggered by potiuk).

Head commit for run:
69ad2ea8f9d4ffe0a06f63dce84c919833fa258b / Jarek Potiuk <[email protected]>
Make SSH remote-job kill test diagnosable and stop it reddening main

test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly 
and
the whole budget went on polling for a pid file that never appeared.

That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and 
both
produce the same ~5s runtime. Two changes separate them:

- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
  creates that directory and the log file before it backgrounds anything, so a
  missing directory means the launcher never got there and an empty one means 
the
  job was launched and died before its first statement

The teardown also has a real ordering hazard, closed here: the marker only says 
the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty 
master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty 
is
now held open until the job proves it left the session by recording its own pid.

I could not reproduce the failure. macOS has no setsid(1) so the class skips 
there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 
25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is 
hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the 
first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.

Generated-by: Claude Opus 5 (1M context)

Report URL: https://github.com/apache/airflow/actions/runs/30304570196

With regards,
GitHub Actions via GitBox


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to