> From: Konstantin Ananyev [mailto:[email protected]]
> 
> > About replacing rte_memcpy with memcpy()...
> >
> > > From: Stephen Hemminger [mailto:[email protected]]
> > > Sent: Thursday, 20 August 2026 07.12
> > >
> > > The DPDK function rte_memcpy() only exists as an optimization
> > > for shortcomings in performance of libc memcpy() on some platforms.
> >
> > Yes, and those platforms should benefit from it.
> >
> > E.g. the vhost performance improvements for Haswell and Broadwell
> [1].
> > Where similar performance improvements implemented in the relevant
> > compilers (GCC, Clang, MSVC)?
> >
> > [1]:
> >
> https://github.com/DPDK/dpdk/commit/4b42e90ef0e421dc777f2b2e377eb237cd
> > 3675fa
> >
> > IMO, performance should remain a high priority for DPDK.
> 
> As I can read the series, good few of them do remove rte_memcpy from
> the CP,
> where it is clearly irrelevant.

Agree!

> For those on the DP, at least for some of them we can run perf tests:
> let say for hash we do have perf_autotest which can be used to measure
> the
> perf diff. If there is none, or neglectable - then no point to keep
> rte_memcpy here.

Unless that perf test is run on all platforms, the result only shows perf diff 
on the tested platforms.

> 
> >
> > > Many platforms have no special rte_memcpy() and just use memcpy().
> > >
> > > But many analysis and test tools know that memcpy() is a special
> > > case and check for overwrite, bounds errors etc. Therefore memcpy()
> > > should be preferred wherever possible.
> >
> > I think this is the only substantial benefit of replacing
> rte_memcpy() with
> > memcpy()!
> > Could we reap this benefit by having special builds for such tools,
> where
> > rte_memcpy() is modified to use memcpy() instead?
> > Then we wouldn't have to compromise on performance.
> >
> > Also, rte_memcpy() used to have a pragma disabling bounds checks due
> to some
> > Intel drivers using [0] instead of []; the pragma was removed from
> rte_memcpy()
> > when the Intel drivers were fixed.
> > I'm not sufficiently familiar with analysis/test tools to say what
> they can detect
> > when using memcpy() instead of the copy methods used by rte_memcpy().
> >
> > >
> > > This patch series introduces a coccinelle script to find
> > > calls to rte_memcpy() where size is fixed, and change them to
> > > regular memcpy(). This was the starting point for this cleanup.
> > >
> > > There is also some cleanups to include rte_memcpy.h and string.h
> > > where needed. Often the includes were happening by some other
> > > header. And also removal of rte_memcpy.h where no longer needed.
> > >
> > > The result is a 46% reduction in use of rte_memcpy.
> > > The remaining rte_memcpy can be cleaned up later:
> > >  - drivers with active maintenance (like mlx5);
> > >  - changes to rte_memcpy which need benchmarking;
> > >  - test code for rte_memcpy can be removed as last step.
> > >
> > > No functional change, no warnings in all compilers including LTO.
> >
> > memcpy() does not always use inline vector instructions for fixed
> size copy [2].
> >
> > [2]:
> > https://inbox.dpdk.org/dev/98CBD80474FA8B44BF855DF32C47DC35F659B8@sma
> > rtserver.smartshare.dk/
> >
> >
> > Another disadvantage of rte_memcpy() is the lack of developer
> guidance.
> > It is not well documented when to use rte_memcpy() and when to use
> memcpy().
> > We discussed something similar on the Tech Board meeting yesterday;
> it is not
> > well documented when to use which type of "ring" (normal, RTS, HTS),
> so maybe
> > we could remove one of them.
> > But removing an option is not an improvement, if the removed option
> would
> > have been the better choice for some use cases.
> >
> > PS: The general guidance for rte_memcpy() usage is something like:
> > rte_memcpy() only in fast path,
> > memcpy() everywhere else,
> > assignment "=" when copying fixed size structures.
> 
> I suppose for te_memcpy() we can be even more strict:
> Use it only for DP, and only after measurement, that shows
> clear perf improvement over ordinal memcpy().
> Alnd also ask contributors to document it (in the comments), i.e.:
> /* on <platform testsed> rte_memcpy() gives X% perf boost when doing
> ...*/
> rte_memcpy(...);

Disagree!
DPDK has performance optimized libs and functions.
Developers should not need to document that using a DPDK function is faster 
than using a libc function.
We don't require perf measurements for using DPDK rte_hash instead of libc 
hashmap.

I agree about not using rte_memcpy() in the control plane.
And I support Stephen's effort to clean this up.

But why the eagerness to avoid using rte_memcpy() in the fast path?

Reply via email to