Hey Sean,

Thanks for your answer.

Im my case, SGE == 0 so I don't think I can go any lighter. It really looks 
like I am hitting some kind of barrier (I am still in PCIe 2.0, I should be 
able to run my test on PCIe 3.0 -enabled system next week, I'll send an update 
then) and from my analysis it looks like the barrier is the HW.

I would like to understand and make sure that I am exploiting the best 
parallelism out of the controller.

Thanks,

- Xavier

On Jun 1, 2012, at 7:48 PM, "Hefty, Sean" <[email protected]> wrote:

>> In my test case to evaluate fetch-and-add, I spawn multiple threads, each
>> owning its own QP inside the same PD and context and sending fetch-and-add
>> requests without any inner contention (no lock, etc.). I quickly reach a
>> ceiling of about 900KOPS with 5/6 threads, and I have a hard time figuring 
>> out
>> why.
> 
> If you're setting inline data, try reducing it to 16 bytes or less.  We've 
> seen small message rates drop drastically when using a larger inline data 
> size.
> 
> - Sean
--
To unsubscribe from this list: send the line "unsubscribe linux-rdma" in
the body of a message to [email protected]
More majordomo info at  http://vger.kernel.org/majordomo-info.html

Reply via email to