Thanks, I had only checked scalar code generation initially.  In an
isolated scalar test on x86_64, the new form is 7 instructions versus  10
for the existing form, but I can reproduce that benchmark results depend on
optimization context.  With GCC 17 -O3 generic, the loop favors the
existing form, while with -fno-tree-vectorize or -march=native the new form
is faster on my machine.

Given that, I have to do more benchmarking before claiming this is a
performance improvement, focus was mostly on number of instructions


Em qui., 30 de jul. de 2026 às 12:11, Jonathan Wakely <[email protected]>
escreveu:

> On Thu, 30 Jul 2026 at 01:58, Francisco Muniz wrote:
> >
> > Use Falk Hueffner's leap-year test for year::is_leap after shifting
> > the valid std::chrono::year range by a multiple of 400. The shift
> > preserves divisibility by 4, 100, and 400, and the unsigned conversion
> > gives the intended modulo 2^32 arithmetic.
> >
> > Idea by Cassio Neri: add 32800, which is 82 * 400, to shift the signed
> > year range into the supported non-negative range.
>
> This is interesting, but your new code is slower than Cassio's code
> when I benchmark it. Have you done your own benchmarking?
>
> >
> > Tested on x86_64-pc-linux-gnu:
> >   make -j$(nproc) all-gcc
> >   make -j$(nproc) all-target-libstdc++-v3
> >   make check RUNTESTFLAGS='conformance.exp=std/time/year/1.cc'
> >   make check RUNTESTFLAGS='conformance.exp=std/time/year/2.cc'
> >
> > libstdc++-v3/ChangeLog:
> >
> >         * include/std/chrono (year::is_leap): Use Hueffner leap-year
> >         test after biasing the year by a multiple of 400.
> >
> > Signed-off-by: Francisco Muniz <[email protected]>
> > ---
> >  libstdc++-v3/include/std/chrono | 31 ++++++++-----------------------
> >  1 file changed, 8 insertions(+), 23 deletions(-)
> >
> > diff --git a/libstdc++-v3/include/std/chrono
> b/libstdc++-v3/include/std/chrono
> > index 692fd6025e7..4483914c08b 100644
> > --- a/libstdc++-v3/include/std/chrono
> > +++ b/libstdc++-v3/include/std/chrono
> > @@ -904,29 +904,14 @@ _GLIBCXX_BEGIN_NAMESPACE_VERSION
> >        constexpr bool
> >        is_leap() const noexcept
> >        {
> > -       // Testing divisibility by 100 first gives better performance
> [1], i.e.,
> > -       //     return _M_y % 100 == 0 ? _M_y % 400 == 0 : _M_y % 16 == 0;
> > -       // Furthermore, if _M_y % 100 == 0, then _M_y % 400 == 0 is
> equivalent
> > -       // to _M_y % 16 == 0, so we can simplify it to
> > -       //     return _M_y % 100 == 0 ? _M_y % 16 == 0 : _M_y % 4 == 0.
> // #1
> > -       // Similarly, we can replace 100 with 25 (which is good since
> > -       // _M_y % 25 == 0 requires one fewer instruction than _M_y % 100
> == 0
> > -       // [2]):
> > -       //     return _M_y % 25 == 0 ? _M_y % 16 == 0 : _M_y % 4 == 0.
> // #2
> > -       // Indeed, first assume _M_y % 4 != 0.  Then _M_y % 16 != 0 and
> hence,
> > -       // _M_y % 4 == 0 and _M_y % 16 == 0 are both false.  Therefore,
> #2
> > -       // returns false as it should (regardless of _M_y % 25.) Now
> assume
> > -       // _M_y % 4 == 0.  In this case, _M_y % 25 == 0 if, and only if,
> > -       // _M_y % 100 == 0, that is, #1 and #2 are equivalent.  Finally,
> #2 is
> > -       // equivalent to
> > -       //     return (_M_y & (_M_y % 25 == 0 ? 15 : 3)) == 0.
> > -
> > -       // References:
> > -       // [1] https://github.com/cassioneri/calendar
> > -       // [2] https://godbolt.org/z/55G8rn77e
> > -       // [3]
> https://gcc.gnu.org/pipermail/libstdc++/2021-June/052815.html
> > -
> > -       return (_M_y & (_M_y % 25 == 0 ? 15 : 3)) == 0;
> > +       // Shift into the range supported by Falk Hueffner's leap-year
> test:
> > +       //
> hueffner.de/falk/blog/a-leap-year-check-in-three-instructions.html
> > +       // Adding a multiple of 400 preserves divisibility by 4, 100,
> and 400.
> > +       // Idea by Cassio Neri: add 32800 (82 * 400).
> > +       // The conversion to uint32_t gives the algorithm's intended
> modulo 2^32
> > +       // arithmetic.
> > +       const auto __y = static_cast<uint32_t>(_M_y) + 32800u;
> > +       return ((__y * 1073750999u) & 3221352463u) <= 126976u;
> >        }
> >
> >        explicit constexpr
> > --
> > 2.47.3
> >
>
>

Reply via email to