On 21/09/26 3:07 PM, Avinash Jayakar wrote:
> Hi,
>
> I have incorporated the changes mentioned in the review of v1 and v2
> https://gcc.gnu.org/pipermail/gcc-patches/2026-September/730358.html
> https://gcc.gnu.org/pipermail/gcc-patches/2026-September/731788.html
>
> Changes from v2:
> - Commit message reworded.
> - Removed extra lines in rs6000-builtin.cc
> - Group builtin function (base and extended mnemonic) in documentation.
> - Update the comment in test cases for justifying load instructions.
> Changes from v1:
> - Add documentation relevant to this patch.
> - Expand the macros used in test sha-builtins-2.c
> - Move the parameter check in rs6000_expand_builtin after constant check
> is done.
> - Correct formatting in crypto.md and remove the & contraint.
>
> Bootstrapped and regtested on ppc64le. Ok for trunk?
>
> Thanks,
> Avinash Jayakar
>
> This patch adds new builtins for SHA2 and SHAPAD instructions which may
> or may not be supported in a future processor.
>
> Following are the new builtins added to support sha2 and shapad:
> void __builtin_dmsha2hash (__dmr1024*, __dmr1024*, uint1);
> void __builtin_dmxxshapad (__dmr1024*, vec_t, uint2, uint1, uint2);
>
> Following are the builtins to support the extended mnemonics:
> void __builtin_dmsha256hash (__dmr1024*, __dmr1024*);
> void __builtin_dmsha512hash (__dmr1024*, __dmr1024*);
> void __builtin_dmxxsha3512pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxsha3384pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxsha3256pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxsha3224pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxshake256pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxshake128pad (__dmr1024*, vec_t, uint1);
> void __builtin_dmxxsha384512pad (__dmr1024*, vec_t);
> void __builtin_dmxxsha224256pad (__dmr1024*, vec_t);
>
> The parameter check in dmxxshapad is done such that if argument 3
> corresponding to ID is 1 and argument 5 corresponding to BL is 0 or 1
> the compiler throws an error. And the test 3 validates this. Although
> the assembler captures this error, it would be better for the compiler
> as well to throw this error as it will be caught earlier with better
> error information.
>
> 2026-09-21 Avinash Jayakar <[email protected]>
>
> gcc/ChangeLog:
> * config/rs6000/crypto.md (UNSPEC_DMSHA2HASH): New unspec entry.
> (UNSPEC_DMSHA256HASH): Likewise.
> (UNSPEC_DMSHA512HASH): Likewise.
> (UNSPEC_DMXXSHAPAD): Likewise.
> (UNSPEC_DMXXSHA3512PAD): Likewise.
> (UNSPEC_DMXXSHA3384PAD): Likewise.
> (UNSPEC_DMXXSHA3256PAD): Likewise.
> (UNSPEC_DMXXSHA3224PAD): Likewise.
> (UNSPEC_DMXXSHAKE256PAD): Likewise.
> (UNSPEC_DMXXSHAKE128PAD): Likewise.
> (UNSPEC_DMXXSHA384512PAD): Likewise.
> (UNSPEC_DMXXSHA224256PAD): Likewise.
> (SHA2_code): Iterator for sha2 extended mnemonics.
> (SHA2_insn): Attribute for sha2 extended mnemonics.
> (SHAPAD_code): Iterator for sha3/shake padding extended mnemonics.
> (SHAPAD_insn): Attribute for sha3/shake padding extended mnemonics.
> (SHA2PAD_code): Iterator for sha2 padding extended mnemonics.
> (SHA2PAD_insn): Attribute for sha2 padding extended mnemonics.
> (dmsha2hash): New define_insn for base sha2.
> (<SHA2_insn>): New define_insn for sha2 extended mnemonics.
> (dmshapad): New define_insn for base shapad.
> (<SHAPAD_insn>): New define_insn for sha3/shake padding extended
> mnemonics.
> (<SHA2PAD_insn>): New define_insn for sha2 padding extended mnemonics.
> * config/rs6000/rs6000-builtin.cc (rs6000_expand_builtin): Error
> handling for shapad builtin arguments.
> * config/rs6000/rs6000-builtins.def: Add new builtin definitions for
> SHA2/3.
> * doc/extend.texi: Document the added builtins.
>
> gcc/testsuite/ChangeLog:
> * gcc.target/powerpc/sha-builtins-1.c: New test.
> * gcc.target/powerpc/sha-builtins-2.c: New test.
> * gcc.target/powerpc/sha-builtins-3.c: New test.
> * gcc.target/powerpc/sha-builtins-4.c: New test.
> ---
> gcc/config/rs6000/crypto.md | 82 ++++++++
> gcc/config/rs6000/rs6000-builtin.cc | 17 ++
> gcc/config/rs6000/rs6000-builtins.def | 72 +++++++
> gcc/doc/extend.texi | 32 +++
> .../gcc.target/powerpc/sha-builtins-1.c | 42 ++++
> .../gcc.target/powerpc/sha-builtins-2.c | 186 ++++++++++++++++++
> .../gcc.target/powerpc/sha-builtins-3.c | 33 ++++
> .../gcc.target/powerpc/sha-builtins-4.c | 54 +++++
> 8 files changed, 518 insertions(+)
> create mode 100644 gcc/testsuite/gcc.target/powerpc/sha-builtins-1.c
> create mode 100644 gcc/testsuite/gcc.target/powerpc/sha-builtins-2.c
> create mode 100644 gcc/testsuite/gcc.target/powerpc/sha-builtins-3.c
> create mode 100644 gcc/testsuite/gcc.target/powerpc/sha-builtins-4.c
>
> diff --git a/gcc/config/rs6000/crypto.md b/gcc/config/rs6000/crypto.md
> index a25d9f75439..06f9d000197 100644
> --- a/gcc/config/rs6000/crypto.md
> +++ b/gcc/config/rs6000/crypto.md
> @@ -48,6 +48,18 @@ (define_c_enum "unspec"
> UNSPEC_XXAES128GENLKP
> UNSPEC_XXAES192GENLKP
> UNSPEC_XXAES256GENLKP
> + UNSPEC_DMSHA2HASH
> + UNSPEC_DMSHA256HASH
> + UNSPEC_DMSHA512HASH
> + UNSPEC_DMXXSHAPAD
> + UNSPEC_DMXXSHA3512PAD
> + UNSPEC_DMXXSHA3384PAD
> + UNSPEC_DMXXSHA3256PAD
> + UNSPEC_DMXXSHA3224PAD
> + UNSPEC_DMXXSHAKE256PAD
> + UNSPEC_DMXXSHAKE128PAD
> + UNSPEC_DMXXSHA384512PAD
> + UNSPEC_DMXXSHA224256PAD
> UNSPEC_XXGFMUL128
> UNSPEC_XXGFMUL128GCM
> UNSPEC_XXGFMUL128XTS])
> @@ -83,6 +95,31 @@ (define_int_iterator AESACC_base_code [UNSPEC_XXAESENCP
> (define_int_attr AESACC_base_insn [(UNSPEC_XXAESENCP "xxaesencp")
> (UNSPEC_XXAESDECP "xxaesdecp")])
>
> +;; Iterator and attribute for SHA2 extended mnemonics
> +(define_int_iterator SHA2_code [UNSPEC_DMSHA256HASH
> + UNSPEC_DMSHA512HASH])
Add a new line here. Separate out SHA2_code and SHA2_insn.
> +(define_int_attr SHA2_insn [(UNSPEC_DMSHA256HASH "dmsha256hash")
> + (UNSPEC_DMSHA512HASH "dmsha512hash")])
> +
> +;; Iterator and attribute for SHAPAD extended mnemonics
> +(define_int_iterator SHAPAD_code [UNSPEC_DMXXSHA3512PAD
> + UNSPEC_DMXXSHA3384PAD
> + UNSPEC_DMXXSHA3256PAD
> + UNSPEC_DMXXSHA3224PAD
> + UNSPEC_DMXXSHAKE256PAD
> + UNSPEC_DMXXSHAKE128PAD])
Add a new line here.
> +(define_int_iterator SHA2PAD_code [UNSPEC_DMXXSHA384512PAD
> + UNSPEC_DMXXSHA224256PAD])
> +
> +(define_int_attr SHAPAD_insn [(UNSPEC_DMXXSHA3512PAD "dmxxsha3512pad")
> + (UNSPEC_DMXXSHA3384PAD "dmxxsha3384pad")
> + (UNSPEC_DMXXSHA3256PAD "dmxxsha3256pad")
> + (UNSPEC_DMXXSHA3224PAD "dmxxsha3224pad")
> + (UNSPEC_DMXXSHAKE256PAD "dmxxshake256pad")
> + (UNSPEC_DMXXSHAKE128PAD "dmxxshake128pad")])
Ditto.
With these changes, the patch is ok to be upstreamed (once the rest of the
SHA2/SHA3
patches are approved).
-Surya
> +(define_int_attr SHA2PAD_insn [(UNSPEC_DMXXSHA384512PAD "dmxxsha384512pad")
> + (UNSPEC_DMXXSHA224256PAD "dmxxsha224256pad")])
> +
> ;; Iterator and attribute for AES encrypt/decrypt extended mnemonics
> (define_int_iterator AESACC_code [UNSPEC_XXAES128ENCP
> UNSPEC_XXAES192ENCP
> @@ -211,3 +248,48 @@ (define_insn "<AESGF_insn>"
> AESGF_code))]
> "TARGET_FUTURE"
> "<AESGF_insn> %x0,%x1,%x2")
> +
> +(define_insn "dmsha2hash"
> + [(set (match_operand:TDO 0 "dmr_register_operand" "=wD")
> + (unspec:TDO [(match_operand:TDO 1 "dmr_register_operand" "0")
> + (match_operand:TDO 2 "dmr_register_operand" "wD")
> + (match_operand:SI 3 "const_0_to_1_operand" "n")]
> + UNSPEC_DMSHA2HASH))]
> + "TARGET_FUTURE"
> + "dmsha2hash %0,%2,%3")
> +
> +(define_insn "<SHA2_insn>"
> + [(set (match_operand:TDO 0 "dmr_register_operand" "=wD")
> + (unspec:TDO [(match_operand:TDO 1 "dmr_register_operand" "0")
> + (match_operand:TDO 2 "dmr_register_operand" "wD")]
> + SHA2_code))]
> + "TARGET_FUTURE"
> + "<SHA2_insn> %0,%2")
> +
> +(define_insn "dmxxshapad"
> + [(set (match_operand:TDO 0 "dmr_register_operand" "=wD")
> + (unspec:TDO [(match_operand:TDO 1 "dmr_register_operand" "0")
> + (match_operand:V16QI 2 "vsx_register_operand" "wa")
> + (match_operand:SI 3 "const_0_to_3_operand" "n")
> + (match_operand:SI 4 "const_0_to_1_operand" "n")
> + (match_operand:SI 5 "const_0_to_3_operand" "n")]
> + UNSPEC_DMXXSHAPAD))]
> + "TARGET_FUTURE"
> + "dmxxshapad %0,%2,%3,%4,%5")
> +
> +(define_insn "<SHAPAD_insn>"
> + [(set (match_operand:TDO 0 "dmr_register_operand" "=wD")
> + (unspec:TDO [(match_operand:TDO 1 "dmr_register_operand" "0")
> + (match_operand:V16QI 2 "vsx_register_operand" "wa")
> + (match_operand:SI 3 "const_0_to_1_operand" "n")]
> + SHAPAD_code))]
> + "TARGET_FUTURE"
> + "<SHAPAD_insn> %0,%2,%3")
> +
> +(define_insn "<SHA2PAD_insn>"
> + [(set (match_operand:TDO 0 "dmr_register_operand" "=wD")
> + (unspec:TDO [(match_operand:TDO 1 "dmr_register_operand" "0")
> + (match_operand:V16QI 2 "vsx_register_operand" "wa")]
> + SHA2PAD_code))]
> + "TARGET_FUTURE"
> + "<SHA2PAD_insn> %0,%2")
> diff --git a/gcc/config/rs6000/rs6000-builtin.cc
> b/gcc/config/rs6000/rs6000-builtin.cc
> index ad6f1fc09f4..2d3f7f4e275 100644
> --- a/gcc/config/rs6000/rs6000-builtin.cc
> +++ b/gcc/config/rs6000/rs6000-builtin.cc
> @@ -3788,6 +3788,23 @@ rs6000_expand_builtin (tree exp, rtx target, rtx /*
> subtarget */,
> }
> }
>
> + if (fcode == RS6000_BIF_DMXXSHAPAD_INTERNAL)
> + {
> + tree arg3 = CALL_EXPR_ARG (exp, 2);
> + tree arg5 = CALL_EXPR_ARG (exp, 4);
> +
> + HOST_WIDE_INT id_param = TREE_INT_CST_LOW (arg3);
> + HOST_WIDE_INT bl_param = TREE_INT_CST_LOW (arg5);
> +
> + if (id_param == 1 && (bl_param == 0 || bl_param == 1))
> + {
> + error ("Invalid combination for %qs: argument 5 cannot be %ld if "
> + "the argument 3 is 1",
> + rs6000_builtin_info[fcode].bifname, bl_param);
> + return const0_rtx;
> + }
> + }
> +
> if (bif_is_ldstmask (*bifaddr))
> return rs6000_expand_ldst_mask (target, arg[0]);
>
> diff --git a/gcc/config/rs6000/rs6000-builtins.def
> b/gcc/config/rs6000/rs6000-builtins.def
> index 4b7d21cdb81..1be92fda977 100644
> --- a/gcc/config/rs6000/rs6000-builtins.def
> +++ b/gcc/config/rs6000/rs6000-builtins.def
> @@ -4313,3 +4313,75 @@
> dmr1024 __builtin_mma_pmdmxvf16gerx2nn_internal (dmr1024, v256, vuc, const
> int<8>, \
> const int<4>, const int<2>);
> PMDMXVF16GERX2NN_INTERNAL dmf_pmdmxvf16gerx2nn {dm,pair}
> +
> + void __builtin_dmsha2hash (dmr1024 *, dmr1024 *, const int<1>);
> + DMSHA2HASH nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmsha2hash_internal (dmr1024, dmr1024, const int<1>);
> + DMSHA2HASH_INTERNAL dmsha2hash {dm}
> +
> + void __builtin_dmsha256hash (dmr1024 *, dmr1024 *);
> + DMSHA256HASH nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmsha256hash_internal (dmr1024, dmr1024);
> + DMSHA256HASH_INTERNAL dmsha256hash {dm}
> +
> + void __builtin_dmsha512hash (dmr1024 *, dmr1024 *);
> + DMSHA512HASH nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmsha512hash_internal (dmr1024, dmr1024);
> + DMSHA512HASH_INTERNAL dmsha512hash {dm}
> +
> + void __builtin_dmxxshapad (dmr1024 *, vuc, const int<2>, const int<1>,
> const int<2>);
> + DMXXSHAPAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxshapad_internal (dmr1024, vuc, const int<2>, const
> int<1>, const int<2>);
> + DMXXSHAPAD_INTERNAL dmxxshapad {dm}
> +
> + void __builtin_dmxxsha3512pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHA3512PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha3512pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHA3512PAD_INTERNAL dmxxsha3512pad {dm}
> +
> + void __builtin_dmxxsha3384pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHA3384PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha3384pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHA3384PAD_INTERNAL dmxxsha3384pad {dm}
> +
> + void __builtin_dmxxsha3256pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHA3256PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha3256pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHA3256PAD_INTERNAL dmxxsha3256pad {dm}
> +
> + void __builtin_dmxxsha3224pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHA3224PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha3224pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHA3224PAD_INTERNAL dmxxsha3224pad {dm}
> +
> + void __builtin_dmxxshake256pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHAKE256PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxshake256pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHAKE256PAD_INTERNAL dmxxshake256pad {dm}
> +
> + void __builtin_dmxxshake128pad (dmr1024 *, vuc, const int<1>);
> + DMXXSHAKE128PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxshake128pad_internal (dmr1024, vuc, const int<1>);
> + DMXXSHAKE128PAD_INTERNAL dmxxshake128pad {dm}
> +
> + void __builtin_dmxxsha384512pad (dmr1024 *, vuc);
> + DMXXSHA384512PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha384512pad_internal (dmr1024, vuc);
> + DMXXSHA384512PAD_INTERNAL dmxxsha384512pad {dm}
> +
> + void __builtin_dmxxsha224256pad (dmr1024 *, vuc);
> + DMXXSHA224256PAD nothing {dm,dmint,dmr}
> +
> + dmr1024 __builtin_dmxxsha224256pad_internal (dmr1024, vuc);
> + DMXXSHA224256PAD_INTERNAL dmxxsha224256pad {dm}
> diff --git a/gcc/doc/extend.texi b/gcc/doc/extend.texi
> index 88213d25a89..884656dbf15 100644
> --- a/gcc/doc/extend.texi
> +++ b/gcc/doc/extend.texi
> @@ -27189,6 +27189,38 @@ vec_t __builtin_galois_field_mult_gcm (vec_t, vec_t);
> vec_t __builtin_galois_field_mult_xts (vec_t, vec_t);
> @end smallexample
>
> +@subsubheading PowerPC Secure Hash Algorithm (SHA) Acceleration Built-in
> Functions
> +
> +Future ISA of the PowerPC may add new instructions for accelerating SHA
> +hash function. GCC provides support for these new instructions through the
> +following built-in functions.
> +
> +@smallexample
> +void __builtin_dmsha2hash (__dmr1024*, __dmr1024*, uint1);
> +
> +void __builtin_dmsha256hash (__dmr1024*, __dmr1024*);
> +
> +void __builtin_dmsha512hash (__dmr1024*, __dmr1024*);
> +
> +void __builtin_dmxxshapad (__dmr1024*, vec_t, uint2, uint1, uint2);
> +
> +void __builtin_dmxxsha3512pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxsha3384pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxsha3256pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxsha3224pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxshake256pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxshake128pad (__dmr1024*, vec_t, uint1);
> +
> +void __builtin_dmxxsha384512pad (__dmr1024*, vec_t);
> +
> +void __builtin_dmxxsha224256pad (__dmr1024*, vec_t);
> +@end smallexample
> +
> @subsubheading PowerPC Vector Integer Multiply High Built-in Functions
>
> Future PowerPC processors may add new instructions for vector integer
> diff --git a/gcc/testsuite/gcc.target/powerpc/sha-builtins-1.c
> b/gcc/testsuite/gcc.target/powerpc/sha-builtins-1.c
> new file mode 100644
> index 00000000000..86a7d610fcd
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/powerpc/sha-builtins-1.c
> @@ -0,0 +1,42 @@
> +/* { dg-do compile } */
> +/* { dg-require-effective-target powerpc_future_compile_ok } */
> +/* { dg-require-effective-target lp64 } */
> +/* { dg-options "-mdejagnu-cpu=future -O2" } */
> +
> +void
> +sha2hash_256 (__dmr1024 *dst, __dmr1024 *src)
> +{
> + __builtin_dmsha2hash (dst, src, 0);
> +}
> +
> +void
> +sha2hash_512 (__dmr1024 *dst, __dmr1024 *src)
> +{
> + __builtin_dmsha2hash (dst, src, 1);
> +}
> +
> +void
> +sha256hash (__dmr1024 *dst, __dmr1024 *src)
> +{
> + __builtin_dmsha256hash (dst, src);
> +}
> +
> +void
> +sha512hash (__dmr1024 *dst, __dmr1024 *src)
> +{
> + __builtin_dmsha512hash (dst, src);
> +}
> +
> +/* Each function loads both dst and src. src is loaded since it is an input
> to
> + the instruction. dst is loaded to preserve the state of hashing. */
> +/* { dg-final { scan-assembler-times {\mlxvp\M} 32 } } */
> +
> +/* Each function stores only one dmr into dst. */
> +/* { dg-final { scan-assembler-times {\mstxvp\M} 16 } } */
> +
> +/* { dg-final { scan-assembler-times {\mdmsha2hash\M} 2 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmsha2hash\M\s+[^,\n]+,\s*[^,\n]+,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmsha2hash\M\s+[^,\n]+,\s*[^,\n]+,\s*1\M} 1 } } */
> +
> +/* { dg-final { scan-assembler-times {\mdmsha256hash\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmsha512hash\M} 1 } } */
> diff --git a/gcc/testsuite/gcc.target/powerpc/sha-builtins-2.c
> b/gcc/testsuite/gcc.target/powerpc/sha-builtins-2.c
> new file mode 100644
> index 00000000000..22c25f626c8
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/powerpc/sha-builtins-2.c
> @@ -0,0 +1,186 @@
> +/* { dg-do compile } */
> +/* { dg-require-effective-target powerpc_future_compile_ok } */
> +/* { dg-require-effective-target lp64 } */
> +/* { dg-options "-mdejagnu-cpu=future -O2" } */
> +
> +typedef unsigned char vec_t __attribute__((vector_size(16)));
> +
> +void
> +shapad_0_0_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 0, 0);
> +}
> +void
> +shapad_0_0_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 0, 1);
> +}
> +void
> +shapad_0_0_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 0, 2);
> +}
> +void
> +shapad_0_0_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 0, 3);
> +}
> +void
> +shapad_0_1_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 1, 0);
> +}
> +void
> +shapad_0_1_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 1, 1);
> +}
> +void
> +shapad_0_1_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 1, 2);
> +}
> +void
> +shapad_0_1_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 0, 1, 3);
> +}
> +
> +void
> +shapad_1_0_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 1, 0, 2);
> +}
> +void
> +shapad_1_0_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 1, 0, 3);
> +}
> +void
> +shapad_1_1_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 1, 1, 2);
> +}
> +void
> +shapad_1_1_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 1, 1, 3);
> +}
> +
> +void
> +shapad_2_0_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 0, 0);
> +}
> +void
> +shapad_2_0_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 0, 1);
> +}
> +void
> +shapad_2_0_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 0, 2);
> +}
> +void
> +shapad_2_0_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 0, 3);
> +}
> +void
> +shapad_2_1_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 1, 0);
> +}
> +void
> +shapad_2_1_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 1, 1);
> +}
> +void
> +shapad_2_1_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 1, 2);
> +}
> +void
> +shapad_2_1_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 2, 1, 3);
> +}
> +
> +void
> +shapad_3_0_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 0, 0);
> +}
> +void
> +shapad_3_0_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 0, 1);
> +}
> +void
> +shapad_3_0_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 0, 2);
> +}
> +void
> +shapad_3_0_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 0, 3);
> +}
> +void
> +shapad_3_1_0 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 1, 0);
> +}
> +void
> +shapad_3_1_1 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 1, 1);
> +}
> +void
> +shapad_3_1_2 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 1, 2);
> +}
> +void
> +shapad_3_1_3 (__dmr1024 *res, vec_t pd)
> +{
> + __builtin_dmxxshapad (res, pd, 3, 1, 3);
> +}
> +
> +/* Each function loads dst into dmr to preserve state of padding. */
> +/* { dg-final { scan-assembler-times {\mlxvp\M} 112 } } */
> +
> +/* Each function stores only one dmr into dst. */
> +/* { dg-final { scan-assembler-times {\mstxvp\M} 112 } } */
> +
> +/* { dg-final { scan-assembler-times {\mdmxxshapad\M} 28 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*0,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*0,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*0,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*0,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*1,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*1,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*1,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*0,\s*1,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*1,\s*0,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*1,\s*0,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*1,\s*1,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*1,\s*1,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*0,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*0,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*0,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*0,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*1,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*1,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*1,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*2,\s*1,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*0,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*0,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*0,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*0,\s*3\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*1,\s*0\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*1,\s*1\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*1,\s*2\M} 1 } } */
> +/* { dg-final { scan-assembler-times
> {\mdmxxshapad\M\s+[^,\n]+,\s*[^,\n],\s*3,\s*1,\s*3\M} 1 } } */
> diff --git a/gcc/testsuite/gcc.target/powerpc/sha-builtins-3.c
> b/gcc/testsuite/gcc.target/powerpc/sha-builtins-3.c
> new file mode 100644
> index 00000000000..4551af8ef14
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/powerpc/sha-builtins-3.c
> @@ -0,0 +1,33 @@
> +/* { dg-do compile } */
> +/* { dg-require-effective-target powerpc_future_compile_ok } */
> +/* { dg-require-effective-target lp64 } */
> +/* { dg-options "-mdejagnu-cpu=future -O2" } */
> +
> +typedef unsigned char vec_t __attribute__((vector_size(16)));
> +
> +void
> +shapad_1_0_0(__dmr1024 *res, vec_t pd) /* { dg-error "Invalid combination" }
> */
> +{
> + __builtin_dmxxshapad(res, pd, 1, 0, 0);
> +}
> +
> +void
> +shapad_1_0_1(__dmr1024 *res, vec_t pd) /* { dg-error "Invalid combination" }
> */
> +{
> + __builtin_dmxxshapad(res, pd, 1, 0, 1);
> +}
> +
> +void
> +shapad_1_1_0(__dmr1024 *res, vec_t pd) /* { dg-error "Invalid combination" }
> */
> +{
> + __builtin_dmxxshapad(res, pd, 1, 1, 0);
> +}
> +
> +void
> +shapad_1_1_1(__dmr1024 *res, vec_t pd) /* { dg-error "Invalid combination" }
> */
> +{
> + __builtin_dmxxshapad(res, pd, 1, 1, 1);
> +}
> +
> +// There are other error messages emitted in expand phase. Ignore them.
> +/* { dg-excess-errors "" } */
> diff --git a/gcc/testsuite/gcc.target/powerpc/sha-builtins-4.c
> b/gcc/testsuite/gcc.target/powerpc/sha-builtins-4.c
> new file mode 100644
> index 00000000000..8a7fc8b7518
> --- /dev/null
> +++ b/gcc/testsuite/gcc.target/powerpc/sha-builtins-4.c
> @@ -0,0 +1,54 @@
> +/* { dg-do compile } */
> +/* { dg-require-effective-target powerpc_future_compile_ok } */
> +/* { dg-require-effective-target lp64 } */
> +/* { dg-options "-mdejagnu-cpu=future -O2" } */
> +
> +typedef unsigned char vec_t __attribute__((vector_size(16)));
> +
> +void sha3512pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha3512pad (dst, pd, 0);
> +}
> +void sha3384pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha3384pad (dst, pd, 0);
> +}
> +void sha3256pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha3256pad (dst, pd, 0);
> +}
> +void sha3224pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha3224pad (dst, pd, 0);
> +}
> +void shake256pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxshake256pad (dst, pd, 0);
> +}
> +void shake128pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxshake128pad (dst, pd, 0);
> +}
> +void sha384512pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha384512pad (dst, pd);
> +}
> +void sha224256pad (__dmr1024* dst, vec_t pd)
> +{
> + __builtin_dmxxsha224256pad (dst, pd);
> +}
> +
> +/* Each function loads dst into dmr to preserve state of padding. */
> +/* { dg-final { scan-assembler-times {\mlxvp\M} 32 } } */
> +
> +/* Each function stores only one dmr into dst. */
> +/* { dg-final { scan-assembler-times {\mstxvp\M} 32 } } */
> +
> +/* { dg-final { scan-assembler-times {\mdmxxsha3512pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxsha3384pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxsha3256pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxsha3224pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxshake256pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxshake128pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxsha384512pad\M} 1 } } */
> +/* { dg-final { scan-assembler-times {\mdmxxsha224256pad\M} 1 } } */