On 28/07/2026 12:54, Richard Biener wrote:
On Tue, Jul 28, 2026 at 12:40 PM Andrew Stubbs <[email protected]> wrote:

On 28/07/2026 11:22, Richard Biener wrote:
On Tue, Jul 28, 2026 at 11:00 AM Andrew Stubbs <[email protected]> wrote:

On 17/07/2026 18:19, Richard Biener wrote:


Am 17.07.2026 um 13:38 schrieb Andrew Stubbs <[email protected]>:

On 17/07/2026 11:48, Richard Biener wrote:
On Fri, Jul 17, 2026 at 12:13 PM Andrew Stubbs <[email protected]> wrote:

On 17/07/2026 09:13, Richard Biener wrote:
On Wed, Jul 8, 2026 at 4:48 PM Andrew Stubbs <[email protected]> wrote:

Hi all,

This patch series implements the "MEMs with address vectors" feature
that I want to use in AMD GCN, and have previously discussed on gcc@.
The motivation is to simplify a new back-end optimization pass I'm
working on and plan to post in the coming weeks.

The first patch, in this series, is the middle-end enablement and
documentation changes.

The second patch converts the amdgcn back-end to use the new scheme,
without adding any real new features. It depends on the
INSN_BASE_REG_CLASS patch I posted last week (still in review) to run
correctly.

The third and final patch demonstrates how I can use the new MEM to
extend my machine description without writing whole new insns for the
vector modes, just for them to evaluate to the same assembler code in
the end.

I have not yet finished testing the patches, so these might not be quite
the final form, but I'm not expecting big changes to the middle-end
portion.  (I can self-approve the amdgcn parts, once the other is
committed, but comments are welcome.)

OK for mainline? (Pending test results.)

I think this exposes what GCN does in a natural way.  I'm not sure
the advantage (exposing former gather/scatter as move) is not
also its disadvantage (moves are special!), and I somewhat question
the benefit over using some UNSPEC for this.  For example I would
assume that you side-step RTL alias/dependence analysis, meaning these
are barriers for optimization and scheduling?

This is intended to be an enabling technology for my real project, in
which I want to transform scalar operations into vector operations in
the backend, and the special casing for memory accesses was getting out
of hand. By expressing (unmasked) vector operations identically to the
scalar equivalent I can simplify things greatly, and the other compiler
passes can do (some of) their special MEM handling as they normally would.
I see.  I guess you ruled out a
    ([parallel
      (set (mem:<scalar> ...) (...))
      (set (mem:<scalar> ...) (...))])

I would need this insn for V2, V4, V8, V16, V32, and V64. And probably SI/SF, 
DI/DF, and TI variants of each of those too (maybe QI and HI also).

It would be very verbose, and what would the constraints even look like?

Not to mention it's right back to the special-case conversions I was trying to 
avoid in the first place.

etc. with vector extracts for the address.  At least for "transform scalar ops
to vectors" the actual representation of the MEM could be a black box
as long as you
have addresses in vectors (which is possible already).

I'm not really sure what this means. Can you give an example, please?

I mean you can write in your pass ‚turn MEM into a vec op‘ but how this is 
reflected into RTL is a detail that shouldn’t concern the pass itself (aka it 
might be an arbitrary complex but recognizable construct)


The discussion seems to have run out of steam without producing a
conclusion.

Heh, sorry.  It's still that I do not like (mem:<vector> ) much, I
think it does not capture
the GCN case well which the (parallel [(set (mem:<scalar> ...])
variant would.  That
doesn't solve your pattern explosion/typing problem, of course.

I don't know how the constraints would even work?

(parallel [(set (match_operand:<scalar> 0 "memory_operand" "v") (...))
             (set (match_operand:<scalar> 1 "memory_operand" "0") (...))
             ....

Or does it use (match_dup 0)? There are 64 notional scalar operands that
are somehow all in the same register. I don't know how to express that.

Maybe that can be
solved by parametrizing that (ick)

   (parallel_vec:<vector>
    [(set (mem:<scalar> (subreg:DI ...) ...)])

so basically have a parallel that implicitly enumerates operations on all vector
lanes of a mode.  The (parallel ..) bit implies it's unordered.

That could work, although that's a whole project in itself, and you'd
probably want to express all vector operations that way or else you'll
end up having to convert between them. And you'd have to stop the
combine pass trying to insert actual scalars into parallel_vec patterns
because they look like they'd fit.

  > I'm not exactly sure> about the subreg part for the address - some
magic needs to happen for the
subreg offset of course.

Subregs are already basically unusable for GCN because the implied
operation just doesn't fit when you can't reinterpret the entire vector
with a different number of lanes.

I want a "vsubreg" which refers to a portion of each lane, but I also
don't want to completely rewrite the compiler to handle that everywhere
it would come up.

Would this be
too restrictive?  I think GCN can actually do few (two?) lanes as actual vector
operation in an ALU, but I guess we're not modeling that.

As far as I know, the GPU does everything x64 in parallel. I'm sure
there's some hidden implementation details in there, but those are not
in the ISA manual.

That said, I do not want to stand in the way of (mem:<vector> ..) when there's
consensus this is the best general way to address this.

Just to add that (mem:<vector> ..) with unordered semantics ties us to that
unless we want to have a flag on the mem to have left-to-right semantics.
Given it in theory should be made to fit gather/scatter modelling (ignoring
the masking representational issue).  So I'd like at least that thought out,
if not documented and implemented (having the flag and printing it).

I believe the /i and /s flags are available, if we should choose to use them.

Would a target hook not be sufficient though? These things are likely to be true always or never, for any given mode, on most architectures.

Better even, as such flags are somewhat likely to "get lost" in all the random places that manipulate MEM.

Andrew

Reply via email to