> About replacing rte_memcpy with memcpy()...
> 
> > From: Stephen Hemminger [mailto:[email protected]]
> > Sent: Thursday, 20 August 2026 07.12
> >
> > The DPDK function rte_memcpy() only exists as an optimization
> > for shortcomings in performance of libc memcpy() on some platforms.
> 
> Yes, and those platforms should benefit from it.
> 
> E.g. the vhost performance improvements for Haswell and Broadwell [1].
> Where similar performance improvements implemented in the relevant
> compilers (GCC, Clang, MSVC)?
> 
> [1]:
> https://github.com/DPDK/dpdk/commit/4b42e90ef0e421dc777f2b2e377eb237cd
> 3675fa
> 
> IMO, performance should remain a high priority for DPDK.

As I can read the series, good few of them do remove rte_memcpy from the CP,
where it is clearly irrelevant.
For those on the DP, at least for some of them we can run perf tests:
let say for hash we do have perf_autotest which can be used to measure the 
perf diff. If there is none, or neglectable - then no point to keep rte_memcpy 
here.

> 
> > Many platforms have no special rte_memcpy() and just use memcpy().
> >
> > But many analysis and test tools know that memcpy() is a special
> > case and check for overwrite, bounds errors etc. Therefore memcpy()
> > should be preferred wherever possible.
> 
> I think this is the only substantial benefit of replacing rte_memcpy() with
> memcpy()!
> Could we reap this benefit by having special builds for such tools, where
> rte_memcpy() is modified to use memcpy() instead?
> Then we wouldn't have to compromise on performance.
> 
> Also, rte_memcpy() used to have a pragma disabling bounds checks due to some
> Intel drivers using [0] instead of []; the pragma was removed from 
> rte_memcpy()
> when the Intel drivers were fixed.
> I'm not sufficiently familiar with analysis/test tools to say what they can 
> detect
> when using memcpy() instead of the copy methods used by rte_memcpy().
> 
> >
> > This patch series introduces a coccinelle script to find
> > calls to rte_memcpy() where size is fixed, and change them to
> > regular memcpy(). This was the starting point for this cleanup.
> >
> > There is also some cleanups to include rte_memcpy.h and string.h
> > where needed. Often the includes were happening by some other
> > header. And also removal of rte_memcpy.h where no longer needed.
> >
> > The result is a 46% reduction in use of rte_memcpy.
> > The remaining rte_memcpy can be cleaned up later:
> >  - drivers with active maintenance (like mlx5);
> >  - changes to rte_memcpy which need benchmarking;
> >  - test code for rte_memcpy can be removed as last step.
> >
> > No functional change, no warnings in all compilers including LTO.
> 
> memcpy() does not always use inline vector instructions for fixed size copy 
> [2].
> 
> [2]:
> https://inbox.dpdk.org/dev/98CBD80474FA8B44BF855DF32C47DC35F659B8@sma
> rtserver.smartshare.dk/
> 
> 
> Another disadvantage of rte_memcpy() is the lack of developer guidance.
> It is not well documented when to use rte_memcpy() and when to use memcpy().
> We discussed something similar on the Tech Board meeting yesterday; it is not
> well documented when to use which type of "ring" (normal, RTS, HTS), so maybe
> we could remove one of them.
> But removing an option is not an improvement, if the removed option would
> have been the better choice for some use cases.
> 
> PS: The general guidance for rte_memcpy() usage is something like:
> rte_memcpy() only in fast path,
> memcpy() everywhere else,
> assignment "=" when copying fixed size structures.

I suppose for te_memcpy() we can be even more strict:
Use it only for DP, and only after measurement, that shows
clear perf improvement over ordinal memcpy().
Alnd also ask contributors to document it (in the comments), i.e.:
/* on <platform testsed> rte_memcpy() gives X% perf boost when doing ...*/
rte_memcpy(...);
 
   
 


Reply via email to