> -----Original Message----- > From: Kyrylo Tkachov <[email protected]> > Sent: Wednesday, August 19, 2026 6:54 PM > To: Cui, Lili <[email protected]> > Cc: [email protected]; [email protected] > Subject: Re: [PATCH 3/3] ifcvt: Extend noce_convert_multiple_sets to > diamonds [PR125557] > > Hi Lili > > > On 19 Aug 2026, at 10:42, Cui, Lili <[email protected]> wrote: > > > > Hi Kyrill, > > > > Thanks for the patch. I have also been working on this code recently, and > expanding the RTL if-conversion pass to cover more conditional-move shapes > is clearly the right direction. > > > > The patch restricts the shape of the diamond, though. For example, it does > not handle the following case: > > > > long > > f (long a, long b, long c, long d, long s, long x, long y, long z, > > long w) { if (s > 0) { x = a; y = b - 2; } > > else { z = c; w = d - 4; } > > return x + y + z + w; > > } > > > > IF-THEN-ELSE-JOIN block found, pass 1, test 2, then 3, else 4, join 5 > > bb 2: cmp s, 0 ; jle bb 3: insn 17: x(r106) = a + 1 > > insn 18: y(r107) = b - 2 > > bb 4: insn 23: z(r108) = c + 3 > > insn 24: w(r109) = d - 4 > > bb 5: x + y + z + w > > > > I wonder whether these restrictions could be removed by extending your > patch to handle the general case (without distinguishing a primary from a > secondary arm). If so, the multiple-set implementation could also cover > cond_move_process_if_block, avoiding two separate code paths for closely > related transformations -- and perhaps noce_try_cmove_arith as well. > > > > I had prototyped the generalization initially, I thought maybe landing this > version first was simpler as it targets just the particular shape in the PR > for > Snappy. But I’ll rebase and retest that prototype and send it out to see if > you > and others prefer that. > Thanks, > Kyrill > Hi Kyrill,
That sounds great. I am looking forward to the generalized version, but there is no urgency. I realize that supporting the general case will require a substantial amount of work. I have also been experimenting in this area locally, so I have a rough prototype and several less common x86 test cases. Once your version is posted, I will be happy to help test its correctness on x86 and collect some performance data. I would also be glad to discuss the implementation with you. I hope these efforts will make the RTL if-conversion pass more capable. Thanks for working on this, Lili > > > Thanks, > > Lili > > > >> -----Original Message----- > >> From: Kyrylo Tkachov <[email protected]> > >> Sent: Tuesday, August 11, 2026 6:21 PM > >> To: [email protected] > >> Cc: [email protected] > >> Subject: Re: [PATCH 3/3] ifcvt: Extend noce_convert_multiple_sets to > >> diamonds [PR125557] > >> > >> The Linaro CI has informed me that this patch would have regressed an > >> IT blocks test on arm because it doesn’t consider the target-specific > >> MAX_CONDITIONAL_EXECUTE restrictions. I fixed it and sent out an > >> updated version. > >> Thanks, > >> Kyrill > >> > >>> On 10 Aug 2026, at 16:26, Kyrylo Tkachov <[email protected]> > wrote: > >>> > >>> From: Kyrylo Tkachov <[email protected]> > >>> > >>> noce_convert_multiple_sets handles multi-set IF-THEN-JOIN regions > >>> but not IF-THEN-ELSE-JOIN diamonds. After GIMPLE factors a value > >>> shared by both arms, such a diamond can still define several > >>> live-out values and retain a data-dependent branch. > >>> > >>> Validate both arms with the existing multiple-set checks and collect > >>> the distinct pseudo destinations that they define and that are live > >>> out. Choose as primary an arm with at least two such destinations > >>> whose bitmap is a superset of the secondary arm's bitmap. This > >>> leaves single-result diamonds to the existing specialized transforms. > >>> > >>> For each live-out destination, also classify the raw source of its > >>> last RTL definition in the arm. A constant is a non-register source. > >>> Require at least one destination shared by both arms to have a > >>> non-register final source in both arms. This keeps mixed > >>> register-copy and computed diamonds outside the dependency-aware > >>> path while allowing it to own diamonds with a common computed > >>> result. The classification is deliberately not transitive through > >>> arm-local > register copies. > >>> > >>> Evaluate each secondary-arm SET_SRC in original instruction order > >>> into a fresh pseudo with copy_to_mode_reg. Rewire reads of earlier > >>> secondary definitions to their temporaries and preserve each SET's > >>> source > >> location. > >>> When the secondary arm defines a primary destination more than once, > >>> use the temporary for its last definition as the alternate value. > >>> Use the incoming value if the secondary arm does not define that > destination. > >>> > >>> Either CFG arm may be primary. When the branch-target arm is > >>> primary, swap the arms in a local noce_if_info and reverse the > >>> select orientation. Reject secondary speculation that modifies > >>> canonical condition inputs. When it clobbers a reusable > >>> condition-code comparison, stop sharing the comparison and form > >>> later conditional moves from the canonical condition, rematerializing > comparisons as needed. > >>> > >>> In outline, convert > >>> > >>> secondary: > >>> (set advance secondary_advance) > >>> primary: > >>> (set next primary_next) > >>> (set advance primary_advance) > >>> > >>> to > >>> > >>> (set secondary_tmp secondary_advance) (set next > >>> (if_then_else primary_p primary_next incoming_next)) (set > >>> advance > >>> (if_then_else primary_p primary_advance secondary_tmp)) > >>> > >>> where primary_p selects the primary arm. > >>> > >>> The existing simple conditional-move handling owns assignments whose > >>> sources are constants or registers. Factor that operand > >>> classification into noce_simple_cmove_operand_p and use it in both > >>> paths. Apply the dependency-aware diamond conversion only when at > >>> least one arm contains a non-simple source. All-simple diamonds > >>> therefore continue through the existing handling, while > >>> register-copy permutations that it cannot safely convert retain their > branch. > >>> > >>> For speed, weight arm costs by edge probability. For size, sum them. > >>> The target noce_conversion_profitable_p hook retains the final > >>> profitability decision. > >>> > >>> Add execution and code-generation tests for symmetric and asymmetric > >>> diamonds, arm dependencies, repeated definitions, unmatched > >>> live-outs, condition-code clobbers, both arm orientations, > >>> constant-only fallback, the all-simple ownership boundary, and a > >>> shared output whose final source is a register in one arm. > >>> > >>> With this patch we if-convert the unpredictable branch in the Snappy > >>> decompress loop and get ~20% improvement on my aarch64 machine. > >>> > >>> Bootstrapped and tested on aarch64-none-linux-gnu and x86_64-linux. > >>> Ok for trunk? > >>> Thanks, > >>> Kyrill > >>> > >>> gcc/ChangeLog: > >>> > >>> PR tree-optimization/125557 > >>> * ifcvt.cc: Include "explow.h". > >>> (noce_simple_cmove_operand_p): New function. > >>> (noce_try_cmove): Use it. > >>> (noce_convert_multiple_sets): Handle diamond CFG cleanup. > >>> (noce_convert_multiple_sets_1): Evaluate secondary-arm values in > >>> fresh pseudos, preserve their source locations, and use their final > >>> values as conditional-move inputs. Use copy_to_mode_reg. Reject > >>> secondary sequences that modify condition inputs and handle > >>> condition-code clobbers. > >>> (bb_ok_for_noce_convert_multiple_sets): Add REQUIRE_MULTIPLE, > >>> LIVE_OUT_DESTS, HAS_NON_SIMPLE_SRC, and > NONREG_LIVE_OUT_DESTS. > >> Record > >>> distinct live-out destinations, non-simple sources, and the raw > >>> source class of each final live-out definition. > >>> (noce_process_if_block): Recognize and cost multi-set diamonds. > >>> Choose the compatible primary arm from the validated live-out > destinations. > >>> Require a shared destination with non-register final sources in both > >>> arms. Defer all-simple diamonds to existing conditional-move handling. > >>> (check_cond_move_block): Use noce_simple_cmove_operand_p. > >>> > >>> gcc/testsuite/ChangeLog: > >>> > >>> PR tree-optimization/125557 > >>> * gcc.c-torture/execute/ifcvt-diamond-1.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_2.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_3.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_4.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_5.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_9.c: New test. > >>> * gcc.target/aarch64/ifcvt_multiple_sets_diamond_10.c: New test. > >>> * gcc.target/i386/ifcvt-multiple-sets-diamond-1.c: New test. > >>> > >>> Signed-off-by: Kyrylo Tkachov <[email protected]> > >>> --- > >>> gcc/ifcvt.cc | 273 ++++++++++++++++-- > >>> .../gcc.c-torture/execute/ifcvt-diamond-1.c | 143 +++++++++ > >>> .../aarch64/ifcvt_multiple_sets_diamond.c | 68 +++++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_10.c | 26 ++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_2.c | 53 ++++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_3.c | 57 ++++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_4.c | 65 +++++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_5.c | 31 ++ > >>> .../aarch64/ifcvt_multiple_sets_diamond_9.c | 30 ++ > >>> .../i386/ifcvt-multiple-sets-diamond-1.c | 44 +++ > >>> 10 files changed, 758 insertions(+), 32 deletions(-) create mode > >>> 100644 gcc/testsuite/gcc.c-torture/execute/ifcvt-diamond-1.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_10.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_2.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_3.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_4.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_5.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_9.c > >>> create mode 100644 > >>> gcc/testsuite/gcc.target/i386/ifcvt-multiple-sets-diamond-1.c > >>> > >>> diff --git a/gcc/ifcvt.cc b/gcc/ifcvt.cc index > >>> b6ce4da8fe3..06f7623b3eb 100644 > >>> --- a/gcc/ifcvt.cc > >>> +++ b/gcc/ifcvt.cc > >>> @@ -37,6 +37,7 @@ > >>> #include "cfgrtl.h" > >>> #include "cfganal.h" > >>> #include "cfgcleanup.h" > >>> +#include "explow.h" > >>> #include "expr.h" > >>> #include "output.h" > >>> #include "cfgloop.h" > >>> @@ -2263,6 +2264,15 @@ noce_emit_cmove (struct noce_if_info > >>> *if_info, > >> rtx x, enum rtx_code code, > >>> return NULL_RTX; > >>> } > >>> > >>> +/* Return true if X is a constant or register operand suitable for simple > >>> + conditional-move handling. */ > >>> + > >>> +static bool > >>> +noce_simple_cmove_operand_p (rtx x) { > >>> + return CONSTANT_P (x) || register_operand (x, VOIDmode); } > >>> + > >>> /* Try only simple constants and registers here. More complex cases > >>> are handled in noce_try_cmove_arith after noce_try_store_flag_arith > >>> has had a go at it. */ > >>> @@ -2277,8 +2287,8 @@ noce_try_cmove (struct noce_if_info *if_info) > >>> if (!noce_simple_bbs (if_info)) > >>> return false; > >>> > >>> - if ((CONSTANT_P (if_info->a) || register_operand (if_info->a, > VOIDmode)) > >>> - && (CONSTANT_P (if_info->b) || register_operand (if_info->b, > >> VOIDmode))) > >>> + if (noce_simple_cmove_operand_p (if_info->a) > >>> + && noce_simple_cmove_operand_p (if_info->b)) > >>> { > >>> start_sequence (); > >>> > >>> @@ -3898,13 +3908,17 @@ try_emit_cmove_seq (struct noce_if_info > >> *if_info, rtx temp, > >>> conditional set to use the temporary we introduced earlier. > >>> > >>> IF_INFO contains the useful information about the block structure and > >>> - jump instructions. */ > >>> + jump instructions. For an IF-THEN-ELSE-JOIN, first evaluate the > >> secondary > >>> + arm's sets into temporaries and retain their final values for the > >>> + conditional moves. Return true if the replacement is valid and > profitable > >>> + and its CFG changes have been committed, otherwise return false. > >>> + */ > >>> > >>> static bool > >>> noce_convert_multiple_sets (struct noce_if_info *if_info) { > >>> basic_block test_bb = if_info->test_bb; basic_block then_bb = > >>> if_info->then_bb; > >>> + basic_block else_bb = if_info->else_bb; > >>> basic_block join_bb = if_info->join_bb; rtx_insn *jump = > >>> if_info->jump; rtx_insn *cond_earliest; @@ -4014,8 +4028,16 @@ > >>> noce_convert_multiple_sets (struct > >> noce_if_info *if_info) > >>> emit_insn_before_setloc (seq, if_info->jump, INSN_LOCATION > >>> (insn_info.last ()->unmodified_insn)); > >>> > >>> - /* Clean up THEN_BB and the edges in and out of it. */ > >>> - remove_edge (find_edge (test_bb, join_bb)); > >>> + /* Clean up the THEN (and, for a diamond, ELSE) block and the > >>> + edges into > >> and > >>> + out of the if-region. An IF-THEN-ELSE-JOIN has no test->join edge. > >>> + Deleting ELSE_BB removes the test->else and else->join edges > >>> + instead. */ if (else_bb) > >>> + { > >>> + delete_basic_block (else_bb); > >>> + num_true_changes++; > >>> + } > >>> + else > >>> + remove_edge (find_edge (test_bb, join_bb)); > >>> remove_edge (find_edge (then_bb, join_bb)); > >>> redirect_edge_and_branch_force (single_succ_edge (test_bb), > >>> join_bb); delete_basic_block (then_bb); @@ -4033,8 +4055,15 @@ > >>> noce_convert_multiple_sets (struct > >> noce_if_info *if_info) > >>> return true; > >>> } > >>> > >>> -/* This goes through all relevant insns of IF_INFO->then_bb and > >>> tries to > >> create > >>> - conditional moves. Information for the insns is kept in INSN_INFO. > >>> */ > >>> +/* Try to emit the multiple-set conversion described by IF_INFO. > >> INSN_INFO > >>> + holds the primary-arm metadata. For a diamond, evaluate the > >>> + secondary > >> arm > >>> + first and retain its final values for the conditional moves. > >>> + > >>> + LAST_NEEDS_COMPARISON is -1 on the first attempt. Record in it > >>> + the last > >> set > >>> + that needs a temporary to preserve the comparison, then use that > >> boundary > >>> + on the second attempt. Set USE_COND_EARLIEST if the emitted > >>> + sequence > >> uses > >>> + IF_INFO->cond_earliest. Return true if the complete sequence was > >>> + emitted. */ > >>> > >>> static bool > >>> noce_convert_multiple_sets_1 (struct noce_if_info *if_info, @@ > >>> -4060,6 > >>> +4089,73 @@ noce_convert_multiple_sets_1 (struct noce_if_info > >>> +*if_info, > >>> int count = 0; > >>> bool second_try = *last_needs_comparison != -1; *use_cond_earliest > >>> = false; > >>> + auto_delete_vec<noce_multiple_sets_info> else_insn_info; > >>> + > >>> + /* For an IF-THEN-ELSE-JOIN, emit the else block's computations first > into > >>> + fresh temporaries. This leaves the incoming register values > >>> available to > >>> + the then block. The conditional moves below select the then values > >>> or > >>> + these else values. */ > >>> + if (if_info->else_bb) > >>> + { > >>> + init_noce_multiple_sets_info (if_info->else_bb, else_insn_info); > >>> + int else_count = 0; > >>> + rtx_insn *before_else = get_last_insn (); > >>> + location_t saved_location = curr_insn_location (); > >>> + rtx_insn *else_insn; > >>> + FOR_BB_INSNS (if_info->else_bb, else_insn) { if > >>> + (!active_insn_p (else_insn)) > >>> + continue; > >>> + > >>> + noce_multiple_sets_info *info = else_insn_info[else_count]; rtx > >>> + set = single_set (else_insn); gcc_checking_assert (set && REG_P > >>> + (SET_DEST (set)) > >>> + && !HARD_REGISTER_P (SET_DEST (set))); > >>> + > >>> + rtx target = SET_DEST (set); > >>> + rtx value = copy_rtx (SET_SRC (set)); int i, ii; > >>> + FOR_EACH_VEC_ELT (info->rewired_src, i, ii) > >>> + value = simplify_replace_rtx (value, > >>> + else_insn_info[ii]->target, else_insn_info[ii]->temporary); > >>> + > >>> + set_curr_insn_location (INSN_LOCATION (else_insn)); rtx > >>> + temporary = copy_to_mode_reg (GET_MODE (target), value); > >>> + > >>> + info->target = target; > >>> + info->temporary = temporary; > >>> + info->unmodified_insn = else_insn; else_count++; } > >>> + > >>> + set_curr_insn_location (saved_location); > >>> + > >>> + gcc_checking_assert (else_count == (int) > >>> + else_insn_info.length ()); > >>> + > >>> + /* These insns run ahead of the conditional moves. If they > >>> + change a register the comparison reads we cannot reuse it, so bail. > >>> + If they only clobber the condition code, drop the shared compare > >>> + so that every move re-materializes its own. */ > >>> + > >>> + rtx_insn *first_else > >>> + = before_else ? NEXT_INSN (before_else) : get_insns (); > >>> + for (rtx_insn *ei = first_else; ei; ei = NEXT_INSN (ei)) { > >>> + if (modified_in_p (cond, ei)) > >>> + { > >>> + end_sequence (); > >>> + return false; > >>> + } > >>> + if (cc_cmp > >>> + && (modified_in_p (cc_cmp, ei) || (rev_cc_cmp && > >>> + modified_in_p (rev_cc_cmp, ei)))) > >>> + { > >>> + cc_cmp = NULL_RTX; > >>> + rev_cc_cmp = NULL_RTX; > >>> + } > >>> + } > >>> + } > >>> > >>> FOR_BB_INSNS (then_bb, insn) > >>> { > >>> @@ -4084,6 +4180,17 @@ noce_convert_multiple_sets_1 (struct > >>> noce_if_info *if_info, > >>> > >>> rtx old_val = target; > >>> > >>> + /* Use the final value assigned to TARGET on the else arm. > >>> + Scanning in reverse is important when the arm assigns the same > >>> + register more than once. */ > >>> + if (if_info->else_bb) > >>> + for (int j = else_insn_info.length () - 1; j >= 0; --j) if > >>> + (rtx_equal_p (target, else_insn_info[j]->target)) > >>> + { > >>> + old_val = else_insn_info[j]->temporary; > >>> + break; > >>> + } > >>> + > >>> /* As we are transforming > >>> if (x > y) > >>> { > >>> @@ -4371,19 +4478,40 @@ init_noce_multiple_sets_info (basic_block > >>> bb, } > >>> > >>> /* Return true iff basic block TEST_BB is suitable for conversion to a > >>> - series of conditional moves. Also check that we have more than one > >>> - set (other routines can handle a single set better than we would), > >>> - and fewer than PARAM_MAX_RTL_IF_CONVERSION_INSNS sets. While > >> going > >>> - through the insns store the sum of their potential costs in COST. */ > >>> + series of conditional moves. Unless REQUIRE_MULTIPLE is false, > >>> + also > >> check > >>> + that we have more than one set (other routines can handle a single set > >>> + better than we would). A diamond arm may have a single set when it is > >>> + selected as the secondary arm. Require no more than > >>> + PARAM_MAX_RTL_IF_CONVERSION_INSNS sets. While going through > >> the insns store > >>> + the sum of their potential costs in COST. On success, if > >>> + LIVE_OUT_DESTS > >> is > >>> + nonnull, record the distinct pseudo destinations that are live out of > >>> + TEST_BB. If HAS_NON_SIMPLE_SRC is nonnull, set it when an > instruction > >>> + source is neither a constant nor a register operand. On success, if > >>> + NONREG_LIVE_OUT_DESTS is nonnull, record live-out destinations > >>> + whose > >> last > >>> + RTL definition has a non-register source. Constants are non-register > >>> + sources. */ > >>> > >>> static bool > >>> -bb_ok_for_noce_convert_multiple_sets (basic_block test_bb, unsigned > >>> *cost) > >>> +bb_ok_for_noce_convert_multiple_sets (basic_block test_bb, unsigned > >> *cost, > >>> + bool require_multiple = true, > >>> + bitmap live_out_dests = NULL, > >>> + bool *has_non_simple_src = NULL, > >>> + bitmap nonreg_live_out_dests = NULL) > >>> { > >>> rtx_insn *insn; > >>> unsigned count = 0; > >>> unsigned param = param_max_rtl_if_conversion_insns; > >>> bool speed_p = optimize_bb_for_speed_p (test_bb); unsigned > >>> potential_cost = 0; > >>> + if (live_out_dests) > >>> + bitmap_clear (live_out_dests); > >>> + if (has_non_simple_src) > >>> + *has_non_simple_src = false; > >>> + if (nonreg_live_out_dests) > >>> + bitmap_clear (nonreg_live_out_dests); bitmap bb_live_out = > >>> + NULL; if (live_out_dests || nonreg_live_out_dests) > >>> + bb_live_out = df_get_live_out (test_bb); > >>> > >>> FOR_BB_INSNS (test_bb, insn) > >>> { > >>> @@ -4399,6 +4527,9 @@ bb_ok_for_noce_convert_multiple_sets > >> (basic_block test_bb, unsigned *cost) > >>> rtx dest = SET_DEST (set); > >>> rtx src = SET_SRC (set); > >>> > >>> + if (has_non_simple_src && !noce_simple_cmove_operand_p (src)) > >>> + *has_non_simple_src = true; > >>> + > >>> /* Dependency rewiring is keyed by register number, so restrict > >>> destinations to pseudos. Hard-register definitions can overlap > >>> without having the same mode. Do not handle anything involving @@ > >>> -4418,6 +4549,19 @@ bb_ok_for_noce_convert_multiple_sets > >>> (basic_block > >> test_bb, unsigned *cost) > >>> if (!can_conditionally_move_p (GET_MODE (dest))) return false; > >>> > >>> + if (bb_live_out && bitmap_bit_p (bb_live_out, REGNO (dest))) > >>> + { if (live_out_dests) > >>> + bitmap_set_bit (live_out_dests, REGNO (dest)); if > >>> + (nonreg_live_out_dests) > >>> + { > >>> + if (register_operand (src, VOIDmode)) bitmap_clear_bit > >>> + (nonreg_live_out_dests, REGNO (dest)); > >>> + else > >>> + bitmap_set_bit (nonreg_live_out_dests, REGNO (dest)); > >>> + } > >>> + } > >>> + > >>> potential_cost += insn_cost (insn, speed_p); > >>> > >>> count++; > >>> @@ -4426,11 +4570,13 @@ bb_ok_for_noce_convert_multiple_sets > >> (basic_block test_bb, unsigned *cost) > >>> *cost += potential_cost; > >>> > >>> /* If we would only put out one conditional move, the other strategies > >>> - this pass tries are better optimized and will be more appropriate. > >>> + this pass tries are better optimized and will be more appropriate, > >>> so > >>> + require more than one set unless REQUIRE_MULTIPLE is false. A > >> diamond > >>> + arm may have one set when it is selected as the secondary arm. > >>> Some targets want to strictly limit the number of conditional moves > >>> that are emitted, they set this through PARAM, we need to respect > >>> that. */ > >>> - return count > 1 && count <= param; > >>> + return count >= (require_multiple ? 2u : 1u) && count <= param; > >>> } > >>> > >>> /* Compute average of two given costs weighted by relative > >>> probabilities @@ -4475,10 +4621,13 @@ noce_process_if_block (struct > >> noce_if_info *if_info) > >>> (2) x = b; if (...) x = a; > >>> (3) if (...) x = a; // as if with an initial x = x. > >>> (4) if (...) { x = a; y = b; z = c; } // Like 3, for multiple SETS. > >>> + (5) A multi-set IF-THEN-ELSE-JOIN. > >>> The later patterns require jumps to be more expensive. > >>> - For the if (...) x = a; else x = b; case we allow multiple insns > >>> - inside the then and else blocks as long as their only effect is > >>> - to calculate a value for x. > >>> + For the diamond case, defer diamonds with only register and constant > >>> + assignments in both arms to the existing simple conditional-move > >>> + handling. Use an arm that sets at least two distinct live-out > >>> pseudos as > >>> + the primary arm. Every live-out destination set by the secondary > >>> arm > >>> + must also be set by the primary arm. > >>> ??? For future expansion, further expand the "multiple X" rules. > >>> */ > >>> > >>> /* First look for multiple SETS. > >>> @@ -4487,25 +4636,85 @@ noce_process_if_block (struct noce_if_info > >> *if_info) > >>> If a target re-uses the existing CC comparison we keep track of that > >>> and add the costs before default noce_conversion_profitable_p. > >>> */ > >>> > >>> - unsigned potential_cost = if_info->original_cost; unsigned > >>> old_cost = if_info->original_cost; > >>> - if (!else_bb > >>> - && HAVE_conditional_move > >>> - && bb_ok_for_noce_convert_multiple_sets (then_bb, > &potential_cost)) > >>> - { > >>> - /* Temporarily set the original costs to what we estimated so > >>> - we can determine if the transformation is worth it. */ > >>> - if_info->original_cost = potential_cost; > >>> - if (noce_convert_multiple_sets (if_info)) > >>> + unsigned ms_then_cost = 0, ms_else_cost = 0; bool > >>> + ms_then_has_non_simple_src, ms_else_has_non_simple_src; > >> auto_bitmap > >>> + ms_then_live_out_dests, ms_else_live_out_dests; auto_bitmap > >>> + ms_then_nonreg_live_out_dests; auto_bitmap > >>> + ms_else_nonreg_live_out_dests; noce_if_info ms_if_info = > >>> + *if_info; bool multiple_sets_p = false; > >>> + > >>> + if (HAVE_conditional_move) > >>> + { > >>> + if (!else_bb) > >>> + multiple_sets_p > >>> + = bb_ok_for_noce_convert_multiple_sets (then_bb, &ms_then_cost); > >>> + else if (!if_info->then_else_reversed > >>> + && bb_ok_for_noce_convert_multiple_sets > >>> + (then_bb, &ms_then_cost, false, > >>> + ms_then_live_out_dests, &ms_then_has_non_simple_src, > >>> + ms_then_nonreg_live_out_dests) > >>> + && bb_ok_for_noce_convert_multiple_sets > >>> + (else_bb, &ms_else_cost, false, > >>> + ms_else_live_out_dests, &ms_else_has_non_simple_src, > >>> + ms_else_nonreg_live_out_dests)) { > >>> + /* Keep this path for diamonds that share a live-out whose last RTL > >>> + definition has a non-register source in both arms. */ bool > >>> + has_shared_nonreg_live_out > >>> + = bitmap_intersect_p (ms_then_nonreg_live_out_dests, > >>> + ms_else_nonreg_live_out_dests); > >>> + > >>> + if ((ms_then_has_non_simple_src || ms_else_has_non_simple_src) > >>> + && has_shared_nonreg_live_out) > >>> + { > >>> + if (bitmap_count_bits (ms_then_live_out_dests) >= 2 && > >>> + !bitmap_intersect_compl_p (ms_else_live_out_dests, > >>> + ms_then_live_out_dests)) > >>> + multiple_sets_p = true; > >>> + else if (bitmap_count_bits (ms_else_live_out_dests) >= 2 > >>> + && !bitmap_intersect_compl_p (ms_then_live_out_dests, > >>> + ms_else_live_out_dests)) > >>> + { > >>> + /* The branch-target arm is the compatible multi-set > >>> + superset. Make it the primary arm and reverse the select > >>> + orientation. */ > >>> + std::swap (ms_if_info.then_bb, ms_if_info.else_bb); > >>> + ms_if_info.then_else_reversed > >>> + = !ms_if_info.then_else_reversed; std::swap (ms_then_cost, > >>> + ms_else_cost); multiple_sets_p = true; } > >>> + } > >>> + } > >>> + } > >>> + > >>> + if (multiple_sets_p) > >>> + { > >>> + /* The original code runs the comparison and one arm. > >>> + Estimate that cost (for a diamond weight the two arms by their > >>> + probabilities) and let noce_convert_multiple_sets convert only if > >>> + the conditional moves come out cheaper. */ > >>> + unsigned potential_cost = old_cost + ms_then_cost; > >>> + if (ms_if_info.else_bb) > >>> { > >>> - if (dump_file && if_info->transform_name) > >>> + if (optimize_bb_for_speed_p (test_bb)) > >>> + potential_cost > >>> + = old_cost + average_cost (ms_then_cost, ms_else_cost, > >>> + find_edge (test_bb, > >>> + ms_if_info.then_bb)); > >>> + else > >>> + potential_cost = old_cost + ms_then_cost + ms_else_cost; } > >>> + ms_if_info.original_cost = potential_cost; > >>> + if (noce_convert_multiple_sets (&ms_if_info)) { if > >>> + (dump_file && ms_if_info.transform_name) > >>> fprintf (dump_file, "if-conversion succeeded through %s\n", > >>> - if_info->transform_name); > >>> + ms_if_info.transform_name); > >>> return true; > >>> } > >>> - > >>> - /* Restore the original costs. */ > >>> - if_info->original_cost = old_cost; > >>> } > >>> > >>> bool speed_p = optimize_bb_for_speed_p (test_bb); @@ -4813,7 > >>> +5022,7 @@ check_cond_move_block (basic_block bb, > >>> && targetm.small_register_classes_for_mode_p (GET_MODE (dest)))) > >>> return false; > >>> > >>> - if (!CONSTANT_P (src) && !register_operand (src, VOIDmode)) > >>> + if (!noce_simple_cmove_operand_p (src)) > >>> return false; > >>> > >>> if (side_effects_p (src) || side_effects_p (dest)) diff --git > >>> a/gcc/testsuite/gcc.c-torture/execute/ifcvt-diamond-1.c > >>> b/gcc/testsuite/gcc.c-torture/execute/ifcvt-diamond-1.c > >>> new file mode 100644 > >>> index 00000000000..b6a883e3de4 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.c-torture/execute/ifcvt-diamond-1.c > >>> @@ -0,0 +1,143 @@ > >>> +/* Runtime correctness of if-converted IF-THEN-ELSE-JOIN diamonds > with > >>> + multiple output registers (noce_convert_multiple_sets). */ > >>> + > >>> +long g1, g2, g3; > >>> + > >>> +__attribute__ ((noipa)) void > >>> +diamond2 (long c, long x, long y) > >>> +{ > >>> + long a, b; > >>> + if (c & 3) > >>> + { > >>> + a = x + 1; > >>> + b = y - 2; > >>> + } > >>> + else > >>> + { > >>> + a = x * 4; > >>> + b = y + 9; > >>> + } > >>> + g1 = a; > >>> + g2 = b; > >>> +} > >>> + > >>> +/* Then arm reads an earlier then output (arm-internal dependency). > >>> +*/ __attribute__ ((noipa)) void > >>> +diamond3 (long c, long p, long q) > >>> +{ > >>> + long a, b, d; > >>> + if (c > 0) > >>> + { > >>> + a = p ^ q; > >>> + b = a + 7; > >>> + d = q * 2; > >>> + } > >>> + else > >>> + { > >>> + a = p & q; > >>> + b = q | 1; > >>> + d = p - 3; > >>> + } > >>> + g1 = a; > >>> + g2 = b; > >>> + g3 = d; > >>> +} > >>> + > >>> +/* A single output in the else arm. The other register keeps its > >>> incoming > >>> + value on the else path. */ > >>> +__attribute__ ((noipa)) void > >>> +diamond_then2_else1 (long c, long x, long y) { > >>> + long a = x, b = y; > >>> + if (c < 0) > >>> + { > >>> + a = x + 100; > >>> + b = y + 200; > >>> + } > >>> + else > >>> + a = x - 50; > >>> + g1 = a; > >>> + g2 = b; > >>> +} > >>> + > >>> +/* Keep the single-set arm as the likely fallthrough block. The > >>> multi-set > >>> + arm must become the primary arm of the conversion. */ > >>> +__attribute__ ((noipa)) void > >>> +diamond_reversed_then2_else1 (long c, long x) { > >>> + long type = c & 3; > >>> + long next = type; > >>> + long advance; > >>> + if (__builtin_expect (type != 0, 1)) > >>> + advance = type + 1; > >>> + else > >>> + { > >>> + next = x + 1; > >>> + advance = x + 2; > >>> + } > >>> + g1 = next; > >>> + g2 = advance; > >>> +} > >>> + > >>> +/* Each arm produces a live-out value that the other arm does not. > >>> +*/ __attribute__ ((noipa)) void diamond_unmatched_liveouts (long c, > >>> +long p, long q) { > >>> + long a = 100, t = 300, e = 200; > >>> + if (c & 4) > >>> + { > >>> + a = p + q; > >>> + t = p * 2; > >>> + } > >>> + else > >>> + { > >>> + a = p - q; > >>> + e = q * 2; > >>> + } > >>> + g1 = a; > >>> + g2 = e; > >>> + g3 = t; > >>> +} > >>> + > >>> +int > >>> +main (void) > >>> +{ > >>> + for (long c = -4; c <= 12; c++) > >>> + for (long p = -6; p <= 6; p++) > >>> + for (long q = -6; q <= 6; q++) { > >>> + diamond2 (c, p, q); > >>> + if (g1 != ((c & 3) ? p + 1 : p * 4) > >>> + || g2 != ((c & 3) ? q - 2 : q + 9)) > >>> + __builtin_abort (); > >>> + > >>> + diamond3 (c, p, q); > >>> + { > >>> + long ea = (c > 0) ? (p ^ q) : (p & q); > >>> + long eb = (c > 0) ? ea + 7 : (q | 1); > >>> + long ed = (c > 0) ? (q * 2) : (p - 3); > >>> + if (g1 != ea || g2 != eb || g3 != ed) > >>> + __builtin_abort (); > >>> + } > >>> + > >>> + diamond_then2_else1 (c, p, q); > >>> + if (g1 != ((c < 0) ? p + 100 : p - 50) > >>> + || g2 != ((c < 0) ? q + 200 : q)) > >>> + __builtin_abort (); > >>> + > >>> + diamond_reversed_then2_else1 (c, p); { > >>> + long type = c & 3; > >>> + if (g1 != (type ? type : p + 1) > >>> + || g2 != (type ? type + 1 : p + 2)) > >>> + __builtin_abort (); > >>> + } > >>> + > >>> + diamond_unmatched_liveouts (c, p, q); > >>> + if (g1 != ((c & 4) ? p + q : p - q) > >>> + || g2 != ((c & 4) ? 200 : q * 2) > >>> + || g3 != ((c & 4) ? p * 2 : 300)) > >>> + __builtin_abort (); > >>> + } > >>> + return 0; > >>> +} > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond.c > >>> new file mode 100644 > >>> index 00000000000..616425b71a2 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond.c > >>> @@ -0,0 +1,68 @@ > >>> +/* Test if-conversion of IF-THEN-ELSE-JOIN diamonds with multiple > output > >>> + registers through noce_convert_multiple_sets. */ > >>> +/* { dg-do compile } */ > >>> +/* { dg-options "-O2 -fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-unpredictable-cost=100" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-predictable-cost=100" } */ > >>> + > >>> +void sink2 (long, long); > >>> + > >>> +/* Two outputs, both arms write the same registers. */ void > >>> +diamond_arith (long c, long x, long y) { > >>> + long a, b; > >>> + if (c > 7) > >>> + { > >>> + a = x + 1; > >>> + b = y - 2; > >>> + } > >>> + else > >>> + { > >>> + a = x * 4; > >>> + b = y + 9; > >>> + } > >>> + sink2 (a, b); > >>> +} > >>> + > >>> +/* Two outputs computed from constants on each arm. */ void > >>> +diamond_const (long c, long x, long y) { > >>> + long a, b; > >>> + if (c == 3) > >>> + { > >>> + a = 5; > >>> + b = 7; > >>> + } > >>> + else > >>> + { > >>> + a = 9; > >>> + b = 11; > >>> + } > >>> + sink2 (a, b); > >>> +} > >>> + > >>> +/* Two outputs in the then arm, a single output in the else arm. The > second > >>> + register keeps its incoming value on the else path. */ void > >>> +diamond_then2_else1 (long c, long x, long y) { > >>> + long a = x, b = y; > >>> + if (c < 0) > >>> + { > >>> + a = x + 100; > >>> + b = y + 200; > >>> + } > >>> + else > >>> + a = x - 50; > >>> + sink2 (a, b); > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-times "if-conversion succeeded > >>> +through noce_convert_multiple_sets" 2 "ce1" } } */ > >>> + > >>> +/* The converted diamonds are branchless: no conditional branch > >>> +remains. */ > >>> +/* { dg-final { scan-assembler-not > >>> +{\tb(eq|ne|cs|cc|mi|pl|vs|vc|hi|ls|ge|lt|gt|le)\t} } } */ > >>> +/* { dg-final { scan-assembler-not "\tcbn?z\t" } } */ > >>> +/* { dg-final { scan-assembler-not "\ttbn?z\t" } } */ > >>> +/* { dg-final { scan-assembler "\tcsel\t" } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_10.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_10.c > >>> new file mode 100644 > >>> index 00000000000..afe1b035eee > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_10. > >>> +++ c > >>> @@ -0,0 +1,26 @@ > >>> +/* Do not use dependency-aware conversion when the only shared > >>> +live-out > >> has a > >>> + register source in one arm. */ > >>> +/* { dg-do compile } */ > >>> +/* { dg-options "-O2 -fno-ssa-phiopt -fno-tree-ter > >>> +-fno-tree-coalesce-vars -fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-unpredictable-cost=100 > >>> +--param=max-rtl-if-conversion-predictable-cost=100" } */ > >>> + > >>> +void sink2 (long, long); > >>> + > >>> +/* STEP and MIDDLE are computed before the branch. COUNT is copied > >>> +in > >> one arm > >>> + and computed in the other. FIRST changes in only one arm. */ > >>> +void mixed_shared_with_one_arm_output (long c, long count, long > >>> +first) { > >>> + long step = count >> 1; > >>> + long middle = first + step; > >>> + if (c > 7) > >>> + count = step; > >>> + else > >>> + { > >>> + first = middle + 1; > >>> + count = count - step - 1; > >>> + } > >>> + sink2 (count, first); > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-not "if-conversion succeeded through > >>> +noce_convert_multiple_sets" "ce1" } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_2.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_2.c > >>> new file mode 100644 > >>> index 00000000000..6d2a4f7862b > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_2 > >>> +++ .c > >>> @@ -0,0 +1,53 @@ > >>> +/* Each arm of the diamond loads from a selected address and advances > a > >>> + pointer by a selected amount. Once the two arm loads are > >>> +commoned > >> the > >>> + diamond writes two registers on both arms (the load result and the > >>> + advance), which noce_convert_multiple_sets turns into > >>> +conditional moves. */ > >>> +/* { dg-do compile } */ > >>> +/* { dg-options "-O2 -fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options "-mtune=olympus" } */ > >>> + > >>> +#include <stddef.h> > >>> +#include <stdint.h> > >>> + > >>> +extern const int16_t lentab[256]; > >>> + > >>> +static inline uint32_t > >>> +extract (uint32_t val, size_t type) { > >>> + const uint64_t masks = 0x0000FFFF00FF0000ull; > >>> + return val & (uint32_t) ((masks >> (type * 16)) & 0xFFFF); } > >>> + > >>> +ptrdiff_t > >>> +f (const uint8_t *ip, size_t tag, const uint8_t *end, ptrdiff_t op) > >>> +{ > >>> + do > >>> + { > >>> + const uint8_t *old_ip = ip; > >>> + ptrdiff_t lmo = lentab[tag]; > >>> + size_t type = tag & 3; > >>> + if (type == 0) > >>> + { > >>> + size_t n = (tag >> 2) + 1; > >>> + tag = ip[n]; > >>> + ip += n + 1; > >>> + } > >>> + else > >>> + { > >>> + tag = ip[type]; > >>> + ip += type + 1; > >>> + } > >>> + uint32_t next = (uint32_t) old_ip[0] | ((uint32_t) old_ip[1] << 8); > >>> + ptrdiff_t extracted = extract (next, type); > >>> + op += lmo - extracted; > >>> + } > >>> + while (ip < end); > >>> + return op; > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump "if-conversion succeeded through > >>> +noce_convert_multiple_sets" "ce1" } } */ > >>> +/* { dg-final { scan-assembler-times "\tcsinc\t" 2 } } */ > >>> +/* { dg-final { scan-assembler-times "\tldrb\t" 1 } } */ > >>> +/* { dg-final { scan-assembler-not {\tb(eq|ne)\t} } } */ > >>> +/* { dg-final { scan-assembler-not {\tcbn?z\t} } } */ > >>> +/* { dg-final { scan-assembler-not {\ttbn?z\t} } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_3.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_3.c > >>> new file mode 100644 > >>> index 00000000000..c459703e9bd > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_3 > >>> +++ .c > >>> @@ -0,0 +1,57 @@ > >>> +/* Test dependencies between sets in both arms of an IF-THEN-ELSE- > JOIN > >>> + diamond. */ > >>> +/* { dg-do run } */ > >>> +/* { dg-options "-O2 > >>> +--param=max-rtl-if-conversion-unpredictable-cost=100 -fdump-rtl-ce1" > >>> +} */ > >>> +/* Keep both assignments to x in the same RTL pseudo. */ > >>> +/* { dg-additional-options "-fno-tree-ter -fno-tree-coalesce-vars" > >>> +} */ > >>> + > >>> +volatile long gx, gy; > >>> + > >>> +__attribute__ ((noipa)) void > >>> +diamond_dependencies (long c, long a, long b) { > >>> + long x, y; > >>> + if (c & 1) > >>> + { > >>> + x = a + 1; > >>> + y = x ^ b; > >>> + x = y + 3; > >>> + } > >>> + else > >>> + { > >>> + x = b - 1; > >>> + y = x ^ a; > >>> + x = y - 3; > >>> + } > >>> + gx = x; > >>> + gy = y; > >>> +} > >>> + > >>> +__attribute__ ((optimize ("O0"))) int main (void) { > >>> + for (long c = -3; c <= 3; ++c) > >>> + for (long a = -5; a <= 5; ++a) > >>> + for (long b = -5; b <= 5; ++b) { > >>> + long x, y; > >>> + diamond_dependencies (c, a, b); > >>> + if (c & 1) > >>> + { > >>> + long first_x = a + 1; > >>> + y = first_x ^ b; > >>> + x = y + 3; > >>> + } > >>> + else > >>> + { > >>> + long first_x = b - 1; > >>> + y = first_x ^ a; > >>> + x = y - 3; > >>> + } > >>> + if (gx != x || gy != y) > >>> + __builtin_abort (); > >>> + } > >>> + return 0; > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-times "if-conversion succeeded > >>> +through noce_convert_multiple_sets" 1 "ce1" } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_4.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_4.c > >>> new file mode 100644 > >>> index 00000000000..428d7f51cc5 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_4 > >>> +++ .c > >>> @@ -0,0 +1,65 @@ > >>> +/* { dg-do run } */ > >>> +/* { dg-options "-O2 > >>> +--param=max-rtl-if-conversion-unpredictable-cost=100 -fdump-rtl-ce1" > >>> +} */ > >>> + > >>> +volatile long ga, gt, ge; > >>> + > >>> +__attribute__ ((noipa)) void > >>> +convertible (long c, long p, long q) { > >>> + long a, t; > >>> + if (c > 7) > >>> + { > >>> + a = p + 1; > >>> + t = q + 2; > >>> + } > >>> + else > >>> + { > >>> + a = p - 3; > >>> + t = q - 4; > >>> + } > >>> + ga = a; > >>> + gt = t; > >>> +} > >>> + > >>> +__attribute__ ((noipa)) void > >>> +reject_arm_only_values (long c, long p, long q, long t, long e) { > >>> + long a; > >>> + if (c > 7) > >>> + { > >>> + a = p + 1; > >>> + t = q + 2; > >>> + } > >>> + else > >>> + { > >>> + a = p - 3; > >>> + e = q - 4; > >>> + } > >>> + ga = a; > >>> + gt = t; > >>> + ge = e; > >>> +} > >>> + > >>> +__attribute__ ((optimize ("O0"))) int main (void) { > >>> + convertible (8, 10, 20); > >>> + if (ga != 11 || gt != 22) > >>> + __builtin_abort (); > >>> + convertible (7, 10, 20); > >>> + if (ga != 7 || gt != 16) > >>> + __builtin_abort (); > >>> + > >>> + reject_arm_only_values (8, 10, 20, 31, 47); > >>> + if (ga != 11 || gt != 22 || ge != 47) > >>> + __builtin_abort (); > >>> + reject_arm_only_values (7, 10, 20, 31, 47); > >>> + if (ga != 7 || gt != 31 || ge != 16) > >>> + __builtin_abort (); > >>> + return 0; > >>> +} > >>> + > >>> +/* The first diamond is handled by the existing conditional-move path. > The > >>> + second must be rejected because each arm has a live-out value > >>> +not > >> assigned > >>> + by the other arm. */ > >>> +/* { dg-final { scan-rtl-dump-not "if-conversion succeeded through > >>> +noce_convert_multiple_sets" "ce1" } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_5.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_5.c > >>> new file mode 100644 > >>> index 00000000000..7490df7ddc0 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_5 > >>> +++ .c > >>> @@ -0,0 +1,31 @@ > >>> +/* Test a diamond whose likely fallthrough arm has one set and > >>> +whose > >> other > >>> + arm has multiple sets, including a live-out not changed by the > fallthrough > >>> + arm. */ > >>> +/* { dg-do compile } */ > >>> +/* { dg-options "-O2 -fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-predictable-cost=100" } */ > >>> + > >>> +unsigned long > >>> +f (const unsigned char *p, unsigned long tag) { > >>> + unsigned long type = tag & 3; > >>> + unsigned long next = type; > >>> + unsigned long advance; > >>> + if (__builtin_expect (type != 0, 1)) > >>> + advance = type + 1; > >>> + else > >>> + { > >>> + unsigned long base = tag >> 2; > >>> + next = base + 1; > >>> + advance = base + 2; > >>> + } > >>> + return p[next] + (advance << 8); > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-times "if-conversion succeeded > >>> +through noce_convert_multiple_sets" 1 "ce1" } } */ > >>> +/* { dg-final { scan-assembler-times "\tcsinc\t" 1 } } */ > >>> +/* { dg-final { scan-assembler-times "\tcsel\t" 1 } } */ > >>> +/* { dg-final { scan-assembler-times "\tldrb\t" 1 } } */ > >>> +/* { dg-final { scan-assembler-not {\tb(eq|ne)\t} } } */ > >>> +/* { dg-final { scan-assembler-not {\tcbn?z\t} } } */ > >>> +/* { dg-final { scan-assembler-not {\ttbn?z\t} } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_9.c > >>> b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_9.c > >>> new file mode 100644 > >>> index 00000000000..b3e95036b34 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/aarch64/ifcvt_multiple_sets_diamond_9 > >>> +++ .c > >>> @@ -0,0 +1,30 @@ > >>> +/* Do not extend dependency-aware conversion to register-copy > >>> +permutations. */ > >>> +/* { dg-do compile } */ > >>> +/* { dg-options "-O2 -fno-tree-ter -fno-tree-coalesce-vars > >>> +-fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-unpredictable-cost=100" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-predictable-cost=100" } */ > >>> + > >>> +volatile long gx, gy; > >>> + > >>> +void > >>> +f (long c, long a, long b) > >>> +{ > >>> + long x, y; > >>> + if (c > 7) > >>> + { > >>> + x = a; > >>> + y = x; > >>> + x = b; > >>> + } > >>> + else > >>> + { > >>> + x = b; > >>> + y = x; > >>> + x = a; > >>> + } > >>> + gx = x; > >>> + gy = y; > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-not "if-conversion succeeded through > >>> +noce_convert_multiple_sets" "ce1" } } */ > >>> +/* { dg-final { scan-assembler > >>> +{\tb(eq|ne|cs|cc|mi|pl|vs|vc|hi|ls|ge|lt|gt|le)\t} } } */ > >>> diff --git > >>> a/gcc/testsuite/gcc.target/i386/ifcvt-multiple-sets-diamond-1.c > >>> b/gcc/testsuite/gcc.target/i386/ifcvt-multiple-sets-diamond-1.c > >>> new file mode 100644 > >>> index 00000000000..fb4491074f3 > >>> --- /dev/null > >>> +++ b/gcc/testsuite/gcc.target/i386/ifcvt-multiple-sets-diamond-1.c > >>> @@ -0,0 +1,44 @@ > >>> +/* { dg-do run } */ > >>> +/* { dg-require-effective-target lp64 } */ > >>> +/* { dg-options "-O2 -mtune=generic -fdump-rtl-ce1" } */ > >>> +/* { dg-additional-options > >>> +"--param=max-rtl-if-conversion-predictable-cost=100" } */ > >>> + > >>> +/* The single-set arm is the likely fallthrough block. Speculative > >>> arithmetic > >>> + clobbers FLAGS, so each conditional move must re-materialize the > >>> + comparison. */ > >>> + > >>> +volatile long ga, gb; > >>> + > >>> +__attribute__ ((noipa)) void > >>> +f (long c, long x, long y, long b) > >>> +{ > >>> + long a; > >>> + if (__builtin_expect (c <= 7, 1)) > >>> + a = y + 3; > >>> + else > >>> + { > >>> + a = x + 1; > >>> + b = y + 2; > >>> + } > >>> + ga = a; > >>> + gb = b; > >>> +} > >>> + > >>> +/* Keep the runtime driver out of noce so that the dump count is specific > to > >>> + F. */ > >>> +__attribute__ ((optimize ("O0"))) int main (void) { > >>> + for (long c = 5; c != 11; ++c) > >>> + for (long x = -8; x != 9; ++x) > >>> + for (long y = -8; y != 9; ++y) { > >>> + f (c, x, y, 4); > >>> + if (ga != (c > 7 ? x + 1 : y + 3) > >>> + || gb != (c > 7 ? y + 2 : 4)) > >>> + __builtin_abort (); > >>> + } > >>> + return 0; > >>> +} > >>> + > >>> +/* { dg-final { scan-rtl-dump-times "if-conversion succeeded > >>> +through noce_convert_multiple_sets" 1 "ce1" } } */ > >>> -- > >>> 2.50.1 (Apple Git-155) > >>> > >
