On Wed, Jul 1, 2026 at 2:31 PM Nguyen Tran <[email protected]> wrote: > > Thanks Richard. > > On point 1 (get_precision on the int_range): I looked into this, and it > does not give the narrowed precision. irange has no get_precision (); > the only one is irange_bitmask::get_precision (), which returns the > mask's wide_int width, that is, the type precision (32 for int), not the > number of bits the values actually occupy. The current computation > instead derives the precision the values need (for example 20 when the > dividend lies in [0, 1000000]), which is what choose_multiplier takes. > > On point 2 (passing a value_range object to expand_divmod): I would > prefer to keep the int precision here. The multiply-shift narrowing > uses only the precision: choose_multiplier takes an int precision, and > once that number is known the lower bound, the sign and the exact > endpoints play no part in the computation. Passing the whole range > object would give the callee more than it uses, to arrive at the same > single number it needs today. > > The range object would help only a different job, the {1, power-of-two} > work you mentioned, which is not part of this patch. So I would prefer > to keep the int and let that work decide for itself how it wants the > range, rather than add a heavier parameter here that this optimization > does not need. > > This patch is small and self-contained, and it is already bootstrapped > and regression-tested, so I would prefer to keep it as it is. Would > that work?
Fair enough. Note I'm not familiar enough with the code paths to review the rest, I hope Jeff will chime in. Richard. > > On Wed, Jul 1, 2026 at 7:02 PM Richard Biener <[email protected]> > wrote: >> >> On Mon, Jun 22, 2026 at 4:30 AM Nguyen Tran <[email protected]> wrote: >> > >> > For divisions by a compile-time-constant divisor, expand_divmod calls >> > choose_multiplier with the dividend's full type precision (size or >> > size - 1). When the dividend's value range is known to be narrower >> > than the type, choose_multiplier could pick a smaller multiplier and >> > post-shift, but expand_divmod has no path to receive that range >> > information. >> > >> > This patch threads the VRP-derived value range from the GIMPLE side >> > into expand_divmod and forwards it to choose_multiplier: >> > >> > * expand_expr_divmod (gcc/expr.cc) computes a dividend precision from >> > treeop0's range via a new helper determine_value_range, which >> > queries the active range_query during pass_expand. >> > >> > * For unsigned (or non-negative-known signed) operands the precision >> > is wi::min_precision(max_val, UNSIGNED). For signed operands with >> > possibly negative values it is wi::min_precision of the larger >> > of -(min + 1) and max under unsigned interpretation, the bound >> > choose_multiplier itself relies on. >> > >> > * expand_divmod gets a new dividend_prec parameter and an inline >> > helper adjusted_dividend_prec that passes it to choose_multiplier >> > raised to LGUP, so the lgup <= precision precondition is preserved. >> > The helper is invoked at three call sites: unsigned TRUNC, signed >> > TRUNC, and signed FLOOR. Passing -1 selects the dividend's full >> > type precision and reproduces the baseline codegen. dividend_prec >> > is a required parameter, placed before the optional methods >> > argument; all callers other than expand_expr_divmod pass -1. >> > >> > For the example in PR target/91883, >> > >> > unsigned g (unsigned a) { >> > if (a > 1000000) __builtin_unreachable (); >> > return a / 10; >> > } >> > >> > baseline -O2 emits >> > >> > movl %edi, %eax >> > movl $3435973837, %edx >> > imulq %rdx, %rax >> > shrq $35, %rax >> > >> > with this patch: >> > >> > movl %edi, %eax >> > imulq $429497139, %rax, %rax >> > shrq $32, %rax >> > >> > The smaller multiplier fits imulq's 32-bit signed immediate, saving >> > the separate movl, and the post-shift drops from 35 to 32. Similar >> > effects for /3, /7, /100. For signed /7, codegen also shifts from >> > the long path (addl-fixup) to the short path, reducing seven >> > instructions to five. >> > >> > Functions without VRP range information pass -1 and are unaffected. >> > >> > gcc.target/i386/pr115910.c is updated. The test has two functions >> > both doing x / 3U: foo with __builtin_unreachable when x < 0, and >> > bar with no range information. Previously both produced identical >> > codegen; the test asserted 2 imulq + 2 shrq $33. With this patch foo's >> > __builtin_unreachable narrows x to [0, INT_MAX] giving VRP-derived >> > precision 31, so choose_multiplier picks a smaller M (1431655766, >> > fits imulq's 32-bit immediate) and post-shift 32 instead of 33. >> > bar is unchanged. The dg-final assertions are split: still 2 imulq, >> > plus 1 shrq $32 (foo) and 1 shrq $33 (bar). >> > >> > One drawback of this patch is some code bloat: the same >> > divisor at two call sites with different ranges produces different >> > assembly instructions. Unmodified GCC produces the same assembly >> > instructions regardless of range, allowing the linker's identical- >> > code-folding (ICF) pass to merge the two function bodies into a >> > single copy in the final binary. With this patch the bodies differ, >> > so ICF cannot fold them. Clang has the same drawback: it also >> > produces different assembly instructions for the same divisor at >> > different ranges, which ICF cannot fold. >> > >> > For example, consider /3 at two different ranges: >> > >> > unsigned int f3 (unsigned int a) { >> > if (a > 65535) __builtin_unreachable (); >> > return a / 3; >> > } >> > unsigned int f (unsigned int a) { >> > if (a > 655356) __builtin_unreachable (); >> > return a / 3; >> > } >> > >> > Unmodified GCC produces byte-identical assembly instructions for >> > both functions, which ICF folds into a single copy: >> > f3: >> > mov eax, edi >> > mov edx, 2863311531 >> > imul rax, rdx >> > shr rax, 33 >> > ret >> > f: >> > mov eax, edi >> > mov edx, 2863311531 >> > imul rax, rdx >> > shr rax, 33 >> > ret >> > >> > Clang produces different assembly instructions; ICF cannot fold them: >> > f3: >> > imul eax, edi, 43691 >> > shr eax, 17 >> > ret >> > f: >> > mov ecx, edi >> > mov eax, 2863311531 >> > imul rax, rcx >> > shr rax, 33 >> > ret >> > >> > Patched GCC produces different assembly instructions; ICF cannot >> > fold them: >> > f3: >> > mov eax, edi >> > imul rax, rax, 1431677610 >> > shr rax, 32 >> > ret >> > f: >> > mov eax, edi >> > imul rax, rax, 1431657130 >> > shr rax, 32 >> > ret >> > >> > A live demonstration: https://godbolt.org/z/75r6axnvG, compiled with >> > >> > -O2 -ffunction-sections -fuse-ld=gold -Wl,--icf=all >> > >> > The flags: >> > -ffunction-sections put each function into its own section, so >> > the linker can compare and fold them >> > individually >> > -fuse-ld=gold use the gold linker (the default linker >> > does not support ICF) >> > -Wl,--icf=all enable ICF in the linker >> > >> > Under unmodified GCC, f3 and f end up at the same address. Under >> > clang and patched GCC, they end up at different addresses. >> > >> > Under default compiler/linker flags this patch likely has no meaningful >> > downside: ICF is not enabled by default, so unmodified and patched >> > both keep two copies of the body in .text. We consider this an >> > acceptable tradeoff: each individual division is at least as small >> > as before, and clang accepts the same one. >> > >> > Bootstrapped on x86_64-pc-linux-gnu with the default language set >> > (c, c++, fortran, objc, obj-c++, lto): "make bootstrap" followed by >> > "make -k check" from the top of the build tree. Stage-2 vs. stage-3 >> > binary comparison successful (the two compilers GCC produces of itself >> > are byte-identical). No new regressions in the testsuite, and all >> > dg-final assertions across the seven new gcc.target/i386/divmod-range-*.c >> > tests pass. >> > >> > gcc/ChangeLog: >> > >> > PR target/91883 >> > * explow.cc (round_push, align_dynamic_address): Pass -1 for >> > dividend_prec. >> > * expmed.cc (choose_multiplier): Prepend comment stating the >> > multiply-shift formula it computes. >> > (adjusted_dividend_prec): New helper. >> > (expand_divmod): Add dividend_prec parameter. Forward to >> > adjusted_dividend_prec at the unsigned TRUNC, signed TRUNC, >> > and signed FLOOR XOR-sign-flip call sites. Pass -1 at the >> > recursive calls. >> > * expmed.h (expand_divmod): Add dividend_prec parameter. >> > * expr.cc (force_operand): Pass -1 for dividend_prec. >> > (determine_value_range): New helper, queries range_query >> > during pass_expand. >> > (expand_expr_divmod): Compute dividend_prec from treeop0's >> > range and forward to expand_divmod. >> > * optabs.cc (expand_doubleword_mod, expand_doubleword_divmod): >> > Pass -1 for dividend_prec. >> > >> > gcc/testsuite/ChangeLog: >> > >> > PR target/91883 >> > * gcc.target/i386/pr115910.c: Adjust shrq assertions for >> > VRP-narrowed multiplier on foo. >> > * gcc.target/i386/divmod-range-1.c: New test: VRP-narrowed >> > multiplier for unsigned /10. >> > * gcc.target/i386/divmod-range-2.c: New test: VRP-narrowed >> > multiplier for unsigned %10. >> > * gcc.target/i386/divmod-range-3.c: New test: VRP-narrowed >> > multiplier for signed /10. >> > * gcc.target/i386/divmod-range-4.c: New test: VRP shifts >> > signed /7 from the long (addl-fixup) path to the short path. >> > * gcc.target/i386/divmod-range-5.c: New test: regression guard, >> > no narrower range preserves baseline unsigned divmod codegen. >> > * gcc.target/i386/divmod-range-6.c: New test: regression guard, >> > no narrower range preserves signed /7 long-path codegen. >> > * gcc.target/i386/divmod-range-7.c: New test: VRP-narrowed >> > multiplier for signed FLOOR_DIV (uses GIMPLE Front End since >> > FLOOR_DIV_EXPR cannot be produced from C source). >> > >> > Signed-off-by: Nguyen Tran <[email protected]> >> > --- >> > This is v2 of the patch originally posted here: >> > https://gcc.gnu.org/pipermail/gcc-patches/2026-May/717904.html >> > Jeff's review: >> > https://gcc.gnu.org/pipermail/gcc-patches/2026-June/719238.html >> > Thanks Jeff for the review. v2 addresses all the points raised: >> > 1. DCO / sign-off: Added a Signed-off-by: tag to put the contribution >> > under the DCO. >> > 2. "hint" -> "range"/"precision": Reworded throughout. You're right >> > that "hint" implies getting it wrong is harmless, which is not the >> > case here: a wrong value produces wrong codegen. All comments and >> > the new parameter are now phrased in terms of the dividend's range / >> > precision. >> > 3. No default argument: Dropped the default value on the new >> > dividend_prec parameter. It is now a required parameter and all call >> > sites are updated explicitly to pass -1 (which selects the dividend's >> > full type precision), except expand_expr_divmod which passes the >> > VRP-derived value. >> > 4. Overflow in -min - 1: Reworked the signed negative-range >> > computation to use -(min + 1), which yields the same value but cannot >> > overflow at any intermediate step (because min is guaranteed to be >> > < 0 here, so min + 1 cannot overflow, and negating a value in >> > [INT_MIN+1 .. 0] is always representable). >> > 5. Bootstrap + regression, platform stated: Bootstrapped and >> > regression-tested on x86_64-pc-linux-gnu (languages: c, c++, fortran, >> > objc, obj-c++, lto). "make bootstrap" + "make -k check" from the top >> > of the build tree; stage-2 vs. stage-3 comparison is clean >> > (byte-identical), no new regressions. >> > --- >> > gcc/explow.cc | 4 +- >> > gcc/expmed.cc | 51 ++++++++++-- >> > gcc/expmed.h | 3 +- >> > gcc/expr.cc | 77 +++++++++++++++++-- >> > gcc/optabs.cc | 6 +- >> > .../gcc.target/i386/divmod-range-1.c | 15 ++++ >> > .../gcc.target/i386/divmod-range-2.c | 14 ++++ >> > .../gcc.target/i386/divmod-range-3.c | 14 ++++ >> > .../gcc.target/i386/divmod-range-4.c | 14 ++++ >> > .../gcc.target/i386/divmod-range-5.c | 13 ++++ >> > .../gcc.target/i386/divmod-range-6.c | 13 ++++ >> > .../gcc.target/i386/divmod-range-7.c | 34 ++++++++ >> > gcc/testsuite/gcc.target/i386/pr115910.c | 3 +- >> > 13 files changed, 241 insertions(+), 20 deletions(-) >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-1.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-2.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-3.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-4.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-5.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-6.c >> > create mode 100644 gcc/testsuite/gcc.target/i386/divmod-range-7.c >> > >> > diff --git a/gcc/explow.cc b/gcc/explow.cc >> > index 755c1d2e6dc..750c155467f 100644 >> > --- a/gcc/explow.cc >> > +++ b/gcc/explow.cc >> > @@ -1114,7 +1114,7 @@ round_push (rtx size) >> > size = expand_binop (Pmode, add_optab, size, alignm1_rtx, >> > NULL_RTX, 1, OPTAB_LIB_WIDEN); >> > size = expand_divmod (0, TRUNC_DIV_EXPR, Pmode, size, align_rtx, >> > - NULL_RTX, 1); >> > + NULL_RTX, 1, -1); >> > size = expand_mult (Pmode, size, align_rtx, NULL_RTX, 1); >> > >> > return size; >> > @@ -1285,7 +1285,7 @@ align_dynamic_address (rtx target, unsigned >> > required_align) >> > target = expand_divmod (0, TRUNC_DIV_EXPR, Pmode, target, >> > gen_int_mode (required_align / BITS_PER_UNIT, >> > Pmode), >> > - NULL_RTX, 1); >> > + NULL_RTX, 1, -1); >> > target = expand_mult (Pmode, target, >> > gen_int_mode (required_align / BITS_PER_UNIT, >> > Pmode), >> > diff --git a/gcc/expmed.cc b/gcc/expmed.cc >> > index fe39506b6bd..95b2a52cae2 100644 >> > --- a/gcc/expmed.cc >> > +++ b/gcc/expmed.cc >> > @@ -3733,6 +3733,15 @@ expand_widening_mult (machine_mode mode, rtx op0, >> > rtx op1, rtx target, >> > unsignedp, OPTAB_LIB_WIDEN); >> > } >> > >> > +/* For any x satisfying 0 <= x < 2^PRECISION, this function picks >> > + integers M and K satisfying >> > + >> > + floor (x / D) = floor ((x * M) / 2^K) >> > + >> > + so an unsigned divide by the constant D can be replaced by one >> > + multiply and one right shift. K is internal, recovered from the >> > outputs as >> > + K = N + (*POST_SHIFT_PTR). */ >> > + >> > /* Choose a minimal N + 1 bit approximation to 2**K / D that can be used >> > to >> > replace division by D, put the least significant N bits of the result >> > in >> > *MULTIPLIER_PTR, the value K - N in *POST_SHIFT_PTR, and return the >> > most >> > @@ -4243,6 +4252,21 @@ expand_sdiv_pow2 (scalar_int_mode mode, rtx op0, >> > HOST_WIDE_INT d) >> > emit_label (label); >> > return expand_shift (RSHIFT_EXPR, mode, temp, logd, NULL_RTX, 0); >> > } >> > + >> > +/* Return the precision to pass to choose_multiplier. When DIVIDEND_PREC >> > + is non-negative, use it as the value-range-derived precision, raised >> > + to at least LGUP since choose_multiplier asserts lgup <= precision. >> > + When -1, the value range was not used; return DEFAULT_PREC. By >> > + construction in the caller, DIVIDEND_PREC never exceeds DEFAULT_PREC. >> > */ >> > + >> > +static inline int >> > +adjusted_dividend_prec (int dividend_prec, int default_prec, int lgup) >> > +{ >> > + if (dividend_prec >= 0) >> > + return dividend_prec < lgup ? lgup : dividend_prec; >> > + return default_prec; >> > +} >> > + >> > >> > /* Emit the code to divide OP0 by OP1, putting the result in TARGET >> > if that is convenient, and returning where the result is. >> > @@ -4285,7 +4309,7 @@ expand_sdiv_pow2 (scalar_int_mode mode, rtx op0, >> > HOST_WIDE_INT d) >> > rtx >> > expand_divmod (int rem_flag, enum tree_code code, machine_mode mode, >> > rtx op0, rtx op1, rtx target, int unsignedp, >> > - enum optab_methods methods) >> > + int dividend_prec, enum optab_methods methods) >> > { >> > machine_mode compute_mode; >> > rtx tquotient; >> > @@ -4528,7 +4552,12 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > { >> > /* Find a suitable multiplier and right shift count >> > instead of directly dividing by D. */ >> > - mh = choose_multiplier (d, size, size, >> > + >> > + /* Use reduced precision if range info available. >> > */ >> > + int prec = adjusted_dividend_prec (dividend_prec, >> > size, >> > + ceil_log2 (d)); >> > + >> > + mh = choose_multiplier (d, size, prec, >> > &ml, &post_shift); >> > >> > /* If the suggested multiplier is more than SIZE >> > bits, >> > @@ -4537,6 +4566,9 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > if (mh != 0 && (d & 1) == 0) >> > { >> > pre_shift = ctz_or_zero (d); >> > + >> > + /* Only reached when the prior call used full >> > SIZE, >> > + so DIVIDEND_PREC here is saturated. */ >> > mh = choose_multiplier (d >> pre_shift, size, >> > size - pre_shift, >> > &ml, &post_shift); >> > @@ -4681,7 +4713,7 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > int_mode, op0, >> > gen_int_mode (abs_d, >> > int_mode), >> > - NULL_RTX, 0); >> > + NULL_RTX, 0, -1); >> > else >> > quotient = expand_sdiv_pow2 (int_mode, op0, abs_d); >> > >> > @@ -4706,8 +4738,12 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > } >> > else if (size <= HOST_BITS_PER_WIDE_INT) >> > { >> > - choose_multiplier (abs_d, size, size - 1, >> > + int prec = adjusted_dividend_prec (dividend_prec, size >> > - 1, >> > + ceil_log2 (abs_d)); >> > + >> > + choose_multiplier (abs_d, size, prec, >> > &ml, &post_shift); >> > + >> > if (ml < HOST_WIDE_INT_1U << (size - 1)) >> > { >> > rtx t1, t2, t3; >> > @@ -4827,7 +4863,10 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > { >> > rtx t1, t2, t3, t4; >> > >> > - mh = choose_multiplier (d, size, size - 1, >> > + int prec = adjusted_dividend_prec (dividend_prec, size >> > - 1, >> > + ceil_log2 (d)); >> > + >> > + mh = choose_multiplier (d, size, prec, >> > &ml, &post_shift); >> > gcc_assert (!mh); >> > >> > @@ -4869,7 +4908,7 @@ expand_divmod (int rem_flag, enum tree_code code, >> > machine_mode mode, >> > t3 = force_operand (gen_rtx_MINUS (int_mode, t1, nsign), >> > NULL_RTX); >> > t4 = expand_divmod (0, TRUNC_DIV_EXPR, int_mode, t3, op1, >> > - NULL_RTX, 0); >> > + NULL_RTX, 0, -1); >> > if (t4) >> > { >> > rtx t5; >> > diff --git a/gcc/expmed.h b/gcc/expmed.h >> > index 557b58392ba..2678d2f1ad4 100644 >> > --- a/gcc/expmed.h >> > +++ b/gcc/expmed.h >> > @@ -711,7 +711,8 @@ extern rtx maybe_expand_shift (enum tree_code, >> > machine_mode, rtx, int, rtx, >> > int); >> > #ifdef GCC_OPTABS_H >> > extern rtx expand_divmod (int, enum tree_code, machine_mode, rtx, rtx, >> > - rtx, int, enum optab_methods = OPTAB_LIB_WIDEN); >> > + rtx, int, int, >> > + enum optab_methods = OPTAB_LIB_WIDEN); >> > #endif >> > #endif >> > >> > diff --git a/gcc/expr.cc b/gcc/expr.cc >> > index 7ca5dd84ce9..d9158ab4dce 100644 >> > --- a/gcc/expr.cc >> > +++ b/gcc/expr.cc >> > @@ -66,6 +66,8 @@ along with GCC; see the file COPYING3. If not see >> > #include "tree-pretty-print.h" >> > #include "flags.h" >> > #include "internal-fn.h" >> > +#include "value-query.h" >> > +#include "value-range.h" >> > >> > >> > /* If this is nonzero, we do not bother generating VOLATILE >> > @@ -8811,16 +8813,16 @@ force_operand (rtx value, rtx target) >> > return expand_divmod (0, >> > FLOAT_MODE_P (GET_MODE (value)) >> > ? RDIV_EXPR : TRUNC_DIV_EXPR, >> > - GET_MODE (value), op1, op2, target, 0); >> > + GET_MODE (value), op1, op2, target, 0, >> > -1); >> > case MOD: >> > return expand_divmod (1, TRUNC_MOD_EXPR, GET_MODE (value), op1, >> > op2, >> > - target, 0); >> > + target, 0, -1); >> > case UDIV: >> > return expand_divmod (0, TRUNC_DIV_EXPR, GET_MODE (value), op1, >> > op2, >> > - target, 1); >> > + target, 1, -1); >> > case UMOD: >> > return expand_divmod (1, TRUNC_MOD_EXPR, GET_MODE (value), op1, >> > op2, >> > - target, 1); >> > + target, 1, -1); >> > case ASHIFTRT: >> > return expand_simple_binop (GET_MODE (value), code, op1, op2, >> > target, 0, OPTAB_LIB_WIDEN); >> > @@ -9765,6 +9767,35 @@ expand_misaligned_mem_ref (rtx temp, machine_mode >> > mode, int unsignedp, >> > return temp; >> > } >> > >> > +/* Determine the value range of OP at the current statement. >> > + Returns true if range is known, stores bounds in MIN_VAL and MAX_VAL. >> > */ >> > + >> > +static bool >> > +determine_value_range (tree op, wide_int *min_val, wide_int *max_val) >> > +{ >> > + if (!currently_expanding_gimple_stmt) >> > + return false; >> > + >> > + if (TREE_CODE (op) != SSA_NAME) >> > + return false; >> > + >> > + tree type = TREE_TYPE (op); >> > + if (!INTEGRAL_TYPE_P (type)) >> > + return false; >> > + >> > + int_range_max r; >> > + if (!get_range_query (cfun)->range_of_expr (r, op, >> > + >> > currently_expanding_gimple_stmt)) >> > + return false; >> > + >> > + if (r.undefined_p () || r.varying_p ()) >> > + return false; >> > + >> > + *min_val = r.lower_bound (); >> > + *max_val = r.upper_bound (); >> > + return true; >> > +} >> > + >> > /* Helper function of expand_expr_2, expand a division or modulo. >> > op0 and op1 should be already expanded treeop0 and treeop1, using >> > expand_operands. */ >> > @@ -9775,6 +9806,35 @@ expand_expr_divmod (tree_code code, machine_mode >> > mode, tree treeop0, >> > { >> > bool mod_p = (code == TRUNC_MOD_EXPR || code == FLOOR_MOD_EXPR >> > || code == CEIL_MOD_EXPR || code == ROUND_MOD_EXPR); >> > + >> > + /* Calculate dividend precision from value range if available. */ >> > + int dividend_prec = -1; >> > + >> > + if (SCALAR_INT_MODE_P (mode) >> > + && optimize >= 2 >> > + && TREE_CODE (treeop1) == INTEGER_CST) >> > + { >> > + wide_int min_val, max_val; >> > + if (determine_value_range (treeop0, &min_val, &max_val)) >> >> can you use ->get_precision () on the int_range instead? >> >> I do wonder whether passing an actual value_range object to >> expand_divmod would be more useful (see the most recent >> patch to attempt to special case division by {1, power-of-two}? >> >> > + { >> > + if (unsignedp || wi::ges_p (min_val, 0)) >> > + { >> > + /* Unsigned or known non-negative: precision from upper >> > bound. */ >> > + dividend_prec = wi::min_precision (max_val, UNSIGNED); >> > + } >> > + else >> > + { >> > + /* Signed with possible negative values. Take the unsigned >> > + precision of the larger of -(min + 1) and max (or just >> > + -(min + 1) when max < 0, since -(min + 1) dominates >> > then). */ >> > + wide_int neg_side = -(min_val + 1); >> > + wide_int pos_side = wi::ges_p (max_val, 0) ? max_val : >> > neg_side; >> > + wide_int worst = wi::umax (neg_side, pos_side); >> > + dividend_prec = wi::min_precision (worst, UNSIGNED); >> > + } >> > + } >> > + } >> > + >> > if (SCALAR_INT_MODE_P (mode) >> > && optimize >= 2 >> > && get_range_pos_neg (treeop0, currently_expanding_gimple_stmt) == 1 >> > @@ -9786,10 +9846,12 @@ expand_expr_divmod (tree_code code, machine_mode >> > mode, tree treeop0, >> > bool speed_p = optimize_insn_for_speed_p (); >> > do_pending_stack_adjust (); >> > start_sequence (); >> > - rtx uns_ret = expand_divmod (mod_p, code, mode, op0, op1, target, >> > 1); >> > + rtx uns_ret = expand_divmod (mod_p, code, mode, op0, op1, target, 1, >> > + dividend_prec); >> > rtx_insn *uns_insns = end_sequence (); >> > start_sequence (); >> > - rtx sgn_ret = expand_divmod (mod_p, code, mode, op0, op1, target, >> > 0); >> > + rtx sgn_ret = expand_divmod (mod_p, code, mode, op0, op1, target, 0, >> > + dividend_prec); >> > rtx_insn *sgn_insns = end_sequence (); >> > unsigned uns_cost = seq_cost (uns_insns, speed_p); >> > unsigned sgn_cost = seq_cost (sgn_insns, speed_p); >> > @@ -9817,7 +9879,8 @@ expand_expr_divmod (tree_code code, machine_mode >> > mode, tree treeop0, >> > emit_insn (sgn_insns); >> > return sgn_ret; >> > } >> > - return expand_divmod (mod_p, code, mode, op0, op1, target, unsignedp); >> > + return expand_divmod (mod_p, code, mode, op0, op1, target, unsignedp, >> > + dividend_prec); >> > } >> > >> > /* Return true if EXP has a range of values [0..1], false >> > diff --git a/gcc/optabs.cc b/gcc/optabs.cc >> > index d17fbcf12ed..9aaa37aacd8 100644 >> > --- a/gcc/optabs.cc >> > +++ b/gcc/optabs.cc >> > @@ -1143,7 +1143,7 @@ expand_doubleword_mod (machine_mode mode, rtx op0, >> > rtx op1, bool unsignedp) >> > } >> > rtx remainder = expand_divmod (1, TRUNC_MOD_EXPR, word_mode, sum, >> > gen_int_mode (INTVAL (op1), >> > word_mode), >> > - NULL_RTX, 1, OPTAB_DIRECT); >> > + NULL_RTX, 1, -1, OPTAB_DIRECT); >> > if (remainder == NULL_RTX) >> > return NULL_RTX; >> > >> > @@ -1246,7 +1246,7 @@ expand_doubleword_divmod (machine_mode mode, rtx >> > op0, rtx op1, rtx *rem, >> > if (op11 != const1_rtx) >> > { >> > rtx rem2 = expand_divmod (1, TRUNC_MOD_EXPR, mode, quot1, op11, >> > - NULL_RTX, unsignedp, OPTAB_DIRECT); >> > + NULL_RTX, unsignedp, -1, OPTAB_DIRECT); >> > if (rem2 == NULL_RTX) >> > return NULL_RTX; >> > >> > @@ -1261,7 +1261,7 @@ expand_doubleword_divmod (machine_mode mode, rtx >> > op0, rtx op1, rtx *rem, >> > return NULL_RTX; >> > >> > rtx quot2 = expand_divmod (0, TRUNC_DIV_EXPR, mode, quot1, op11, >> > - NULL_RTX, unsignedp, OPTAB_DIRECT); >> > + NULL_RTX, unsignedp, -1, OPTAB_DIRECT); >> > if (quot2 == NULL_RTX) >> > return NULL_RTX; >> > >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-1.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-1.c >> > new file mode 100644 >> > index 00000000000..57ab5149374 >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-1.c >> > @@ -0,0 +1,15 @@ >> > +/* Verify that VRP-derived range info reduces the unsigned divmod >> > multiplier. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +unsigned int >> > +foo (unsigned int a) >> > +{ >> > + if (a > 1000000) __builtin_unreachable (); >> > + return a / 10; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\timulq\t\$429497139,} } } */ >> > +/* { dg-final { scan-assembler {\tshrq\t\$32,} } } */ >> > +/* { dg-final { scan-assembler-not {\$3435973837} } } */ >> > +/* { dg-final { scan-assembler-not {\tshrq\t\$35,} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-2.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-2.c >> > new file mode 100644 >> > index 00000000000..ac5e9edc7aa >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-2.c >> > @@ -0,0 +1,14 @@ >> > +/* Verify that VRP-derived range info reduces the unsigned modulo >> > multiplier. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +unsigned int >> > +foo (unsigned int a) >> > +{ >> > + if (a > 1000000) __builtin_unreachable (); >> > + return a % 10; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\timulq\t\$429497139,} } } */ >> > +/* { dg-final { scan-assembler {\tshrq\t\$32,} } } */ >> > +/* { dg-final { scan-assembler-not {\$3435973837} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-3.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-3.c >> > new file mode 100644 >> > index 00000000000..bbca89b42c0 >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-3.c >> > @@ -0,0 +1,14 @@ >> > +/* Verify that VRP-derived range info reduces the signed-divide-by-10 >> > multiplier. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +int >> > +foo (int a) >> > +{ >> > + if (a > 1000000 || a < -1000000) __builtin_unreachable (); >> > + return a / 10; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\timulq\t\$429497139,} } } */ >> > +/* { dg-final { scan-assembler-not {\$1717986919} } } */ >> > +/* { dg-final { scan-assembler-not {\tsarq\t\$34,} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-4.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-4.c >> > new file mode 100644 >> > index 00000000000..ea9b7926862 >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-4.c >> > @@ -0,0 +1,14 @@ >> > +/* Verify that VRP-derived range info shifts signed /7 from LONG to SHORT >> > codegen. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +int >> > +foo (int a) >> > +{ >> > + if (a > 1000000 || a < -1000000) __builtin_unreachable (); >> > + return a / 7; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\timulq\t\$613567341,} } } */ >> > +/* { dg-final { scan-assembler-not {\$-1840700269} } } */ >> > +/* { dg-final { scan-assembler-not {\taddl\t%edi, %eax} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-5.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-5.c >> > new file mode 100644 >> > index 00000000000..195f41bca9b >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-5.c >> > @@ -0,0 +1,13 @@ >> > +/* Regression test: dividend with no narrower range must preserve baseline >> > + codegen. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +unsigned int >> > +foo (unsigned int a) >> > +{ >> > + return a / 10; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\$3435973837} } } */ >> > +/* { dg-final { scan-assembler {\tshrq\t\$35,} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-6.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-6.c >> > new file mode 100644 >> > index 00000000000..b1a178bca9f >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-6.c >> > @@ -0,0 +1,13 @@ >> > +/* Regression test: dividend with no narrower range must preserve >> > LONG-path >> > + codegen. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +int >> > +foo (int a) >> > +{ >> > + return a / 7; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\$-1840700269} } } */ >> > +/* { dg-final { scan-assembler {\taddl\t%edi, %eax} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/divmod-range-7.c >> > b/gcc/testsuite/gcc.target/i386/divmod-range-7.c >> > new file mode 100644 >> > index 00000000000..524b14b824c >> > --- /dev/null >> > +++ b/gcc/testsuite/gcc.target/i386/divmod-range-7.c >> > @@ -0,0 +1,34 @@ >> > +/* Verify that VRP-derived range info reduces the signed FLOOR_DIV >> > multiplier. >> > + Uses __GIMPLE FE because FLOOR_DIV_EXPR cannot be produced from C >> > source. */ >> > +/* { dg-do compile { target { ! ia32 } } } */ >> > +/* { dg-options "-O2 -fgimple -march=x86-64 -mtune=generic -masm=att" } */ >> > + >> > +int __GIMPLE (ssa) >> > +foo (int a) >> > +{ >> > + int t_2; >> > + unsigned int au_4; >> > + unsigned int range_5; >> > + >> > +__BB(2): >> > + au_4 = (unsigned int) a_3(D); >> > + range_5 = au_4 + 1000000u; >> > + if (range_5 > 2000000u) >> > + goto __BB3; >> > + else >> > + goto __BB4; >> > + >> > +__BB(3): >> > + __builtin_unreachable (); >> > + >> > +__BB(4): >> > + t_2 = a_3(D) __FLOOR_DIV 7; >> > + return t_2; >> > +} >> > + >> > +/* { dg-final { scan-assembler {\timulq\t\$613567341,} } } */ >> > +/* { dg-final { scan-assembler {\tshrq\t\$32,} } } */ >> > +/* { dg-final { scan-assembler {\tsarl\t\$31,} } } */ >> > +/* { dg-final { scan-assembler-times {\txorl\t} 2 } } */ >> > +/* { dg-final { scan-assembler-not {\$2454267027} } } */ >> > +/* { dg-final { scan-assembler-not {\tshrq\t\$34,} } } */ >> > diff --git a/gcc/testsuite/gcc.target/i386/pr115910.c >> > b/gcc/testsuite/gcc.target/i386/pr115910.c >> > index 5f1cd9aa010..41e94142392 100644 >> > --- a/gcc/testsuite/gcc.target/i386/pr115910.c >> > +++ b/gcc/testsuite/gcc.target/i386/pr115910.c >> > @@ -2,7 +2,8 @@ >> > /* { dg-do compile { target { ! ia32 } } } */ >> > /* { dg-options "-O2 -march=x86-64 -mtune=generic -masm=att" } */ >> > /* { dg-final { scan-assembler-times {\timulq\t} 2 } } */ >> > -/* { dg-final { scan-assembler-times {\tshrq\t\$33,} 2 } } */ >> > +/* { dg-final { scan-assembler-times {\tshrq\t\$32,} 1 } } */ >> > +/* { dg-final { scan-assembler-times {\tshrq\t\$33,} 1 } } */ >> > /* { dg-final { scan-assembler-not {\tsarl\t} } } */ >> > >> > int >> > -- >> > 2.43.0 >> >
