The GitHub Actions job "Lint" on tvm.git/main has failed. Run started by GitHub user tqchen (triggered by tqchen).
Head commit for run: 615abe2a568893f109bc0f90e26626aa03a0934c / Guanjie Wang <[email protected]> [TIRx][CUDA] Add the declared synchronization-word wait (#20353) ## Problem This exists for `tirx-kernels`, where every cross-thread protocol is a hand-rolled spin on a global word -- `ld; while (!done) { ld; }`, sometimes with a `__nanosleep` between polls and a fence after the loop. `sm100_fp8_fp4_mega_moe` alone has eight. To a race checker that loop is N ordinary reads racing another CTA's write, and it has to report every one: nothing in the source separates a read that retries until a condition holds from a read that took whatever it got. `wait_until` is that shape, spelled once, and it can replace all of them -- the 18 such loops in the corpus are an acquiring poll, a relaxed poll, or a volatile poll with a backoff, and all three are this operation. Adopting it costs nothing: the table below measures the emitted form against the hand-written loop it replaces, on every kernel that owns one. What it buys is that the word becomes declared -- the address the wait names is the protocol's, and accesses to it are judged against the protocol instead of as an ordinary pair. ## The operation T.cuda.wait_until(dst, ptr, predicate, scope=, space=, ptx_type=, backoff_ns=) Spin on one global word until `predicate` holds; the exit value stays in `dst`. Publishers stay raw PTX: `st.release`, `red`, `atom`. The name is shared with `T.nvshmem.wait_until` deliberately -- that one names a symmetric object and an enumerated comparison, this one a global address and a predicate. ## Generated code ld_relaxed(dst, ptr); // peeled if (!predicate) { #pragma unroll 1 do { ld_relaxed(dst, ptr); } while (!predicate); } { T edge; ld_acquire(edge, ptr); (void)edge; } // the edge `ld.relaxed.<scope>.<space>.<ptx_type>` polls, one `ld.acquire.<scope>` closes. The closing read is discarded on purpose: it may observe a value later than the predicate accepted, and these predicates are not all monotone -- `mega_moe`'s grid barrier tests a sign-bit flip, its ring waits test equality -- so reaching `dst` would hand the caller a value its own predicate rejects. The edge survives: such words are published by release RMWs, so acquiring any contribution synchronizes with all earlier ones. ## Why this shape Three spellings of the same acquire. The baseline is the loop these kernels wrote by hand -- `ld.acquire` on every poll, no closing read -- and both other columns are wall time relative to it, negative meaning faster. Interleaved, every benchmarkable kernel that owns a spin wait, first round dropped as warm-up. | kernel | sites | relaxed poll + acquire exit (emitted) | relaxed poll + `fence.acq_rel` exit | |---|---|---|---| | `sm100_fp8_fp4_mega_moe` | 8 | **-1.04%** | +49.9~51.8% | | `radix_topk_multi_cta` | 1 | **-0.50%** | +5.89~6.10% | | `agent_evolved_moe_fp8_blockscale_dsv3` | 3 | +0.03% | +1.53~2.15% | | `agent_evolved_kda_backward_packed` | 4 | -0.08% | -0.10~+0.03% | | `cudnn_sm100_flex_attention_backward` | 1 | -0.25% | -0.06~+0.96% | The fence exit is never faster than the baseline on any of the five. `fence.sc` costs what `fence.acq_rel` costs, and keeping the acquiring poll *and* adding a fence is just as bad, which puts the cost in the fence rather than in the loop body. `ld.volatile` polls within 0.1% of `ld.relaxed` everywhere, which is ISA 8.4.2 measured; `.relaxed.<scope>` is emitted because it names the scope instead of resting on `volatile` meaning `.sys`. Peeling the first load gives ptxas one loop body to schedule, and the SASS then matches the hand-written loop this replaces. That axis was not swept separately. ## Other parameters - `scope`, `ptx_type` spell the access. `ptx_type` respells at the same width (a counter read as raw bits) and is required when the address is untyped. 128 bits is refused: that exit value is not a scalar a predicate can test. - `space` admits `global` only. A wait within a CTA or cluster belongs on an `mbarrier`; `shared` is refused with a message that says so. - `backoff_ns` puts `__nanosleep` between polls. Nothing helps: monotone degradation from 0, +1.35% at 20ns, +30.08% at 2us. Measured on short waits only, so it stays opt-in. Report URL: https://github.com/apache/tvm/actions/runs/35041454015 With regards, GitHub Actions via GitBox --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
