Hey Sean, Thanks for your answer.
Im my case, SGE == 0 so I don't think I can go any lighter. It really looks like I am hitting some kind of barrier (I am still in PCIe 2.0, I should be able to run my test on PCIe 3.0 -enabled system next week, I'll send an update then) and from my analysis it looks like the barrier is the HW. I would like to understand and make sure that I am exploiting the best parallelism out of the controller. Thanks, - Xavier On Jun 1, 2012, at 7:48 PM, "Hefty, Sean" <[email protected]> wrote: >> In my test case to evaluate fetch-and-add, I spawn multiple threads, each >> owning its own QP inside the same PD and context and sending fetch-and-add >> requests without any inner contention (no lock, etc.). I quickly reach a >> ceiling of about 900KOPS with 5/6 threads, and I have a hard time figuring >> out >> why. > > If you're setting inline data, try reducing it to 16 bytes or less. We've > seen small message rates drop drastically when using a larger inline data > size. > > - Sean -- To unsubscribe from this list: send the line "unsubscribe linux-rdma" in the body of a message to [email protected] More majordomo info at http://vger.kernel.org/majordomo-info.html
