alxrxs opened a new pull request, #3569:
URL: https://github.com/apache/brpc/pull/3569

   ### What problem does this PR solve?
   
   Issue Number: none
   
   Problem Summary:
   
   `RdmaEndpoint::CutFromIOBufList()` posts every message *by reference*: the 
send WR carries SGEs that point into IOBuf blocks, so the NIC has to DMA-read 
the message from host memory before it can transmit it. The QP asks for no 
inline data (`MAX_INLINE_DATA = 64` is in the file but was never used, and was 
commented out in #2876 as unused). For small RPC requests and responses, that 
read is on the latency path of every call.
   
   ### What is changed and the side effects?
   
   Changed:
   
   - `AllocateQp()` asks for `MAX_INLINE_DATA` (now 128) bytes of inline data 
and, if the device refuses, creates the QP without it, as before. The inline 
size the device granted (`ibv_create_qp()` writes it back into `attr`), capped 
at `MAX_INLINE_DATA`, is kept in `RdmaResource::max_inline_data`.
   - `CutFromIOBufList()` sets `IBV_SEND_INLINE` on a message that is no larger 
than that and whose SGEs all point into the RDMA block pool. The message is 
then copied into the send WQE and the NIC does not read it from memory. User 
registered memory (`RegisterMemoryForRdma` / user-data meta) is never inlined, 
because it may be device memory that the CPU cannot copy from. Larger messages 
are posted as before.
   
   Each message is already posted on its own, one WR per `ibv_post_send()`, so 
the inlined WQE can go through the doorbell. 128 bytes keeps the WQE (control 
segment, inline header, payload) within the 256-byte doorbell (BlueFlame) 
buffer of mlx5 devices; past that size inlining no longer saves the read.
   
   Side effects:
   
   - Performance effects: not measured on brpc. The change removes one DMA read 
from host memory before each small message is sent. The same kind of change was 
measured in other RDMA systems, for example in Valkey's RDMA transport (two 
Intel Xeon Platinum 8360Y hosts with ConnectX-6 Dx at RoCEv2 100 Gb/s, 
`valkey-benchmark --rdma -c 1 -P 1 -t get`, 30 paired repetitions), where 
inlining the server's reply lowered the median GET latency by 0.448 us [0.348, 
0.550] (95% interval). What it is worth in brpc has not been measured. On mlx5 
the QP already gets a large inline size without asking, because `max_send_sge` 
is the device maximum, so the send queue does not grow there.
   
   - Breaking backward compatibility: no. `cut_into_sglist_and_iobuf()` 
(private) takes one more argument.
   
   ### Tests
   
   No RDMA hardware was available, so this was tested functionally on soft-RoCE 
(rdma_rxe) in a VM:
   
   - `libbrpc.a` built with `WITH_RDMA=ON` (CMake, gcc 15, dependencies from 
conda-forge), no new warnings.
   - The `rdma_performance` example server, and a small client that sends 
`PerfTestService.Test` with `echo_attachment` and compares the echoed 
attachment with what it sent: attachments of 0 to 300000 bytes (2000 calls per 
size up to 1000 bytes, 50 above), all equal. Requests of 58 to 111 bytes and 
responses of 37 to 85 bytes went inline (2 or 3 SGEs), larger messages by 
reference, checked with a temporary log line that is not part of this PR.
   - The retry without inline data was not exercised (rxe grants inline data). 
Unit tests were not run.
   
   ### Check List:
   - Please make sure your changes are compilable.
   - When providing us with a new feature, it is best to add related tests.
   - Please follow [Contributor Covenant Code of 
Conduct](https://github.com/apache/brpc/blob/master/CODE_OF_CONDUCT.md).
   
   This change was prepared with Claude Code (Claude Opus 5.5) and is marked 
`Generated-by:` in the commit message, as the ASF's generative tooling guidance 
suggests. It is a draft; the submitter will review it before marking it ready.
   
   Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to