bryancall opened a new pull request, #13755:
URL: https://github.com/apache/trafficserver/pull/13755

   `lua_global_shutdown` fails intermittently on CI in the "Shut down while a 
Lua state is running Lua code" and "Traffic Server exited" runs. 
`shutdown_race_client.py` returns 1 after about 15 seconds, and ATS is still 
running at the next check.
   
   `shutdown_race_client.py` waits for `lua-state-1.active` (global state 1 
inside `/hold`), and only then starts waiting for `lua-remap-state.active`. 
`remap_shutdown.lua` writes its marker once (`held_once`), for half a second, 
and never again. When the global marker is slow to appear, the remap window has 
already passed. The client then waits out `STATES_ACTIVE_TIMEOUT_SECONDS`, 
prints `lua-remap-state.active never appeared`, and returns 1 without sending 
SIGTERM. That leaves ATS up for the "Traffic Server exited" check.
   
   This change watches both markers together from the start, and treats each 
one as satisfied once it has been seen. The load on both plugin instances still 
runs across the SIGTERM, so the remap requests keep entering Lua and queuing on 
its mutex through shutdown, as before.
   
   ## Testing
   
   On master in `ci.trafficserver.apache.org/ats/fedora:44` (`ci-fedora-autest` 
preset, run as a non-root user as on CI), with 16 copies of the test running at 
once:
   
   | | Runs | Failed |
   |---|---|---|
   | Before this change | 182 | 19, every one `lua-remap-state.active never 
appeared` |
   | With this change | 191 | 0 |
   
   To force the ordering, I started the global `/hold` load one second late. 
The old client then failed 3 of 3 runs, the same way; this one passed 3 of 3.
   
   The two CI failures whose consoles were still available (ATSUnderground 
sec-master builds 4576 and 4601) fail in exactly those two runs, with the 
client exiting 1 after 15.6 seconds.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to