* H. J. Lu: > On Tue, Jul 21, 2026 at 4:01 PM H.J. Lu <[email protected]> wrote: >> >> On Tue, Jul 21, 2026 at 3:45 PM Florian Weimer <[email protected]> wrote: >> > >> > * H. J. Lu: >> > >> > > With the silicon vendor guarantees from Intel, AMD, Hygon and Zhaoxin in: >> > > >> > > https://gcc.gnu.org/bugzilla/show_bug.cgi?id=104688 >> > > >> > > many software developers would happily use inline 128-bit atomic loads >> > > and stores in their programs because they only target compatible CPUs. >> > > Add -m128bit-atomic to generate 128-bit atomic loads and stores to avoid >> > > the overhead of calling into libatomic. Enable -m128bit-atomic by >> > > default >> > > if supported by the targeting processor, which is one of x86-64-v3 >> > > capable >> > > processors as well as AVX capable processors from Intel, AMD, Hygon and >> > > Zhaoxin, with SEE2 and CMPXCHG16B enabled. >> > > >> > > gcc/ >> > > >> > > PR target/94649 >> > >> > This patch doesn't seem to produce lock cmpxchg16b for the reproducer in >> > PR94649 with just -mcx16. I don't think this optimization needs full >> > 128-bit atomics, just lock cmpxchg16b support is enough. >> >> -mcx16 alone won't generate "lock cmpxchg16b". My patch doesn't change >> atomic_compare_and_swap patterns. It only added atomic_loadti and >> atomic_storeti. I goess "lock cmpxchg16b" needs atomic_loadti and >> atomic_storeti somehow. >> >> > > +@opindex m128bit-atomic >> > > +@item -m128bit-atomic >> > > +Generate cmpxchg16b, 128-bit atomic vector load and store instructions. >> > > +This is safe to use only on x86-64-v3 capable processors as well as AVX >> > > +capable processors from Intel, AMD, Hygon and Zhaoxin, which guarantee >> > > +that 128-bit aligned vector loads and stores are atomic. This option >> > > +is enabled by default if supported by the targeting processor with >> > > +SEE2 and CMPXCHG16B enabled. >> > >> > This could reference -mcx16. >> > >> >> It mentions CMPXCHG16B. How should it be phrased? >> > > How about > > @opindex mcx16 > @opindex mno-cx16 > @item -mcx16 > This option enables GCC to generate @code{CMPXCHG16B} instructions in 64-bit > code to implement compare-and-exchange operations on 16-byte aligned 128-bit > objects. This is useful for atomic updates of data structures exceeding one > machine word in size. The compiler uses this instruction to implement > @ref{__sync Builtins}. However, for @ref{__atomic Builtins} operating on > 128-bit integers, @option{-m128bit-atomic} option, which implies > @option{-mcx16}, is required to avoid a library call.
It seems to describe the behavior of the patch. Whether this is the best possible behavior, I don't know. Thanks, Florian
