================
@@ -16,31 +16,81 @@ internally by the compiler. A thread that initiates one or
more async operations
An *asyncmark* created by a thread can be used to track async operations
initiated by that thread.
+### Stages
+
+A *stage* names a kind of async operation. Each async operation *belongs to*
the
+one stage determined by the instruction that initiates it.
+
+The stages are:
+
+| Bit | Stage | Async operations |
+|---|---|---|
+| 0 | `TENSOR` | tensor loads and stores |
+| 1 | `GLOBAL_LOAD_ASYNC_TO_LDS` | global loads async to LDS |
+| 2 | `GLOBAL_LOAD_ASYNC_TO_LDS_MCAST` | multicast (cluster) global loads
async to LDS |
+| 3 | `GLOBAL_STORE_ASYNC_FROM_LDS` | async global stores from LDS |
+| 5 | `BUFFER_GLOBAL_LOAD` | buffer loads to LDS and pre-gfx1250 global loads
to LDS |
+
+Bits 4 and 6 through 10 are reserved for future async operations, and no
+operation belongs to them yet.
+
+Which async operations a given subtarget actually has is described in
+{ref}`AMDGPU DMA Operations <amdgpu-dma-operations>`. A stage exists on every
+subtarget that supports asyncmarks, whether or not that subtarget has any
+operation belonging to it.
+
+### Stage Masks
+
+Both intrinsics take a *stage mask*: an 11-bit value in which a set bit names a
+stage.
+
+The mask `0` which names no stage is given the special meaning "every stage".
+This ensures if new stages are added that programs will continue to wait on all
+stages, if that was their intention.
+
+Bits not specified in the table above are reserved for future use. It is not an
+error to set them, but it could mean you have more conservative waits than
+necessary when the future stages are added.
+
+Setting a bit that is neither listed nor reserved is an error.
+
+Users are strongly advised to keep bitmasks disjoint in
asyncmark/wait_asyncmark
+operations, or else the resulting program may become rather confusing for them.
----------------
ssahasra wrote:
Confusing ... and mostly just an informational note. If stages are independent,
then any mix of masks should "just work", as long as the user is thinking of
each stage separately.
https://github.com/llvm/llvm-project/pull/220442
_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits