This is an automated email from the ASF dual-hosted git repository.
github-merge-queue[bot] pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/texera.git
The following commit(s) were added to refs/heads/main by this push:
new d2fe4ba34e fix(pyamber, test): stabilize flaky AtomicInteger
get_and_set deadlock test (#7295)
d2fe4ba34e is described below
commit d2fe4ba34e0e76679eb14499c854d487ed953a09
Author: Meng Wang <[email protected]>
AuthorDate: Fri Aug 7 17:23:11 2026 -0700
fix(pyamber, test): stabilize flaky AtomicInteger get_and_set deadlock test
(#7295)
### What changes were proposed in this PR?
`test_get_and_set_does_not_deadlock_on_non_reentrant_lock` (added in
#5010)
waited on an `Event` for a fixed 0.5s and then asserted `not
worker.is_alive()`.
The worker sets that event *inside* `attempt()`, before the thread
exits, so the
assertion could fire while a perfectly correct `get_and_set` was still
tearing
down — and if the worker didn't finish inside the 0.5s budget at all,
the assert
failed outright. Both are wall-clock races unrelated to the deadlock the
test
guards, and both surface under CI load.
Replaced the fixed window with `worker.join(timeout=5)`. `join` returns
only
once the thread is really dead, which is the precondition `is_alive()`
needs, and
it returns in microseconds on a correct implementation — so the timeout
costs
nothing in practice while still letting a real deadlock keep the worker
alive
past it and trip the same assertion. The regression-detection intent is
unchanged. Test-only change; no production code touched.
### Any related issues, documentation, discussions?
Closes #7294.
### How was this PR tested?
`pytest amber/src/test/python/core/util/test_atomic.py` locally on
Python 3.12 —
11 passed, run repeatedly, all green in ~0.6s per run (the join adds no
measurable time).
Failure path verified: temporarily reintroducing the #4794 deadlock in
`AtomicInteger.get_and_set` (`old_value = self.value` while holding the
non-reentrant lock) makes the test fail red with the same `assert not
True` and
pytest exit code 1 after the 5s join timeout.
`ruff check` and `ruff format --check` clean on the touched file.
### Was this PR authored or co-authored using generative AI tooling?
Generated-by: Claude Code (claude-fable-5)
Co-authored-by: Yicong Huang
<[email protected]>
---
amber/src/test/python/core/util/test_atomic.py | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/amber/src/test/python/core/util/test_atomic.py
b/amber/src/test/python/core/util/test_atomic.py
index 824b25b671..c8021234e0 100644
--- a/amber/src/test/python/core/util/test_atomic.py
+++ b/amber/src/test/python/core/util/test_atomic.py
@@ -85,8 +85,11 @@ class TestAtomicIntegerSingleThreaded:
# scheduling delay alone could let the assertions below pass even on
# a fixed implementation.
assert started.wait(timeout=2.0), "worker thread never started"
- # Give get_and_set a moment to either deadlock or return.
- completed.wait(timeout=0.5)
+ # Join instead of waiting on `completed`: the worker sets that event
+ # before it exits, so `is_alive()` is only reliable after a join. A
+ # fixed implementation joins in microseconds; a deadlocked get_and_set
+ # stays alive past the timeout and still trips the assertion below.
+ worker.join(timeout=5)
assert not errors, (
f"get_and_set raised before reaching the deadlock spin:
{errors[0]!r}"
)