On x86 hardware manages the cache coherency. As far as I understood, DMA sync operations are no-ops on x86. But the confusion arose when I realised that there are no arch specific userlibs. So essentially as Roland pointed out, on a non cache coherent architecture userspace applications will break. This will happen for IB as well.
Can PCIe root complex can read dirty data from CPU caches? I am seeing almost 100% LLC misses (L3 on Nehalem, which are inclusive for L1, and L2) for memcpy on RDMA buffers after transmission. Depending upon your CPU, LLC miss penalty could be very high, and shadow performance gains. Has anyone ever tried using RDMA on non cache-coherent systems? I think, cache line flushing is not a privileged instruction and can be called without going to kernel for memory-cache synchronization. Thought? Thanks, -- Animesh Or Gerlitz <[email protected]> wrote on 10/11/2012 11:04:02 PM: > > On Thu, Oct 11, 2012 at 10:44 PM, Roland Dreier > <[email protected]> wrote: > > > No one has really ever tried to deal with the issue of userspace RDMA on > > a cache-incoherent architecture. Basically if you try the currentstack, the > > in-kernel users (IPoIB etc) should be OK but libibverbs etc. will > be completely broken. > > I think the question might refer even to cache-coherent systems, e.g > in the kernel IB core and ULPs all buffers are dma mapped to/from the > device before/after they are touched by the CPU and vise versa, wheres > in user space, after the buffers are registered once, they are > repeatedly touched by the CPUs and provided to the HW for DMA, e.g all > user space buffers are treated like kernel DMA coherent ones. > > Or. > -- To unsubscribe from this list: send the line "unsubscribe linux-rdma" in the body of a message to [email protected] More majordomo info at http://vger.kernel.org/majordomo-info.html
