On Thu, 20 Aug 2026 at 12:59, Yuao Ma <[email protected]> wrote:
>
> On Thu, Aug 20, 2026 at 6:31 PM Jonathan Wakely <[email protected]> wrote:
> >
> > On Thu, 20 Aug 2026 at 11:30, Jonathan Wakely <[email protected]> wrote:
> > >
> > > On Wed, 19 Aug 2026 at 17:23, Yuao Ma <[email protected]> wrote:
> > > >
> > > > Hi!
> > > >
> > > > Similar to std::for_each, this patch optimizes ranges::for_each for
> > > > segmented iterators.
> > >
> > > If I understand correctly, this will break cases that require
> > > std::invoke to invoke the function object, e.g.
> > >
> > > ranges::for_each(r, &T::f);
> >
> > A more concrete example:
> >
> > struct T { void f() { } };
> > std::deque<T> d;
> > ranges::for_each(d, &T::f);
> >
> > deque's _S_for_each_segment just uses __func without std::invoke, doesn't 
> > it?
> >
>
> Actually this will compile and run without error, and my local check
> verifies this. I think the reason is that what we passed to the __func
> is the internal lambda of the std::__for_each_segmented, rather than
> the &T::f. The only place which will be called with member function is
> correctly handled with std::invoke.

Ah yes! When ranges::__for_each stops recursing and calls the 'else'
branch it uses std::__invoke. Nice.

Is there any benefit to passing __f and __proj separately, using two
parameter slots?

ranges::__for_each could take a single __f with no proj, and then just
call __f(*__first) in its else branch. And ranges::for_each could pass
it a lambda which invokes proj and f. That would mean an additional
indirection, but only passing one parameter. Maybe it's not an
improvement.


>
> > >
> > > >
> > > > Fully tested on x86_64-linux with no regressions.
> > > >
> > > > Using the newly added benchmark, it shows a 3x improvement when using
> > > > ranges::for_each with std::deque.
> > > >
> > > > === Wed Aug 19 03:28:22 PM UTC 2026 ===
> > > > for_each.cc               std::for_each vector<int>   2r    1u    0s
> > > >       0mem    0pf
> > > > for_each.cc               std::for_each deque<int>   2r    2u    0s
> > > >      0mem    0pf
> > > > for_each.cc               std::for_each list<int>    13r   14u    0s
> > > >       0mem    0pf
> > > > for_each.cc               std::ranges::for_each vector<int>   2r    2u
> > > >    0s         0mem    0pf
> > > > for_each.cc               std::ranges::for_each deque<int>   6r    5u
> > > >   0s         0mem    0pf
> > > > for_each.cc               std::ranges::for_each list<int>  13r   14u
> > > >  0s         0mem    0pf
> > > > === Wed Aug 19 04:09:51 PM UTC 2026 ===
> > > > for_each.cc               std::for_each vector<int>   2r    1u    0s
> > > >       0mem    0pf
> > > > for_each.cc               std::for_each deque<int>   2r    2u    0s
> > > >      0mem    0pf
> > > > for_each.cc               std::for_each list<int>    13r   14u    0s
> > > >       0mem    0pf
> > > > for_each.cc               std::ranges::for_each vector<int>   2r    2u
> > > >    0s         0mem    0pf
> > > > for_each.cc               std::ranges::for_each deque<int>   2r    1u
> > > >   0s         0mem    0pf
> > > > for_each.cc               std::ranges::for_each list<int>  13r   14u
> > > >  0s         0mem    0pf
> > > >
> > > > Please take a look when you are available, thanks!
> > > >
> > > > Note: after preparing this patch I found the -std=gnu++11 in the check
> > > > performance script based on Jonathan's guidance. I can prepare a patch
> > > > for this tomorrow and get rid of the STD in the benchmark.
> > > >
> > > > Yuao
> >
>

Reply via email to