alxrxs opened a new pull request, #3569: URL: https://github.com/apache/brpc/pull/3569
### What problem does this PR solve? Issue Number: none Problem Summary: `RdmaEndpoint::CutFromIOBufList()` posts every message *by reference*: the send WR carries SGEs that point into IOBuf blocks, so the NIC has to DMA-read the message from host memory before it can transmit it. The QP asks for no inline data (`MAX_INLINE_DATA = 64` is in the file but was never used, and was commented out in #2876 as unused). For small RPC requests and responses, that read is on the latency path of every call. ### What is changed and the side effects? Changed: - `AllocateQp()` asks for `MAX_INLINE_DATA` (now 128) bytes of inline data and, if the device refuses, creates the QP without it, as before. The inline size the device granted (`ibv_create_qp()` writes it back into `attr`), capped at `MAX_INLINE_DATA`, is kept in `RdmaResource::max_inline_data`. - `CutFromIOBufList()` sets `IBV_SEND_INLINE` on a message that is no larger than that and whose SGEs all point into the RDMA block pool. The message is then copied into the send WQE and the NIC does not read it from memory. User registered memory (`RegisterMemoryForRdma` / user-data meta) is never inlined, because it may be device memory that the CPU cannot copy from. Larger messages are posted as before. Each message is already posted on its own, one WR per `ibv_post_send()`, so the inlined WQE can go through the doorbell. 128 bytes keeps the WQE (control segment, inline header, payload) within the 256-byte doorbell (BlueFlame) buffer of mlx5 devices; past that size inlining no longer saves the read. Side effects: - Performance effects: not measured on brpc. The change removes one DMA read from host memory before each small message is sent. The same kind of change was measured in other RDMA systems, for example in Valkey's RDMA transport (two Intel Xeon Platinum 8360Y hosts with ConnectX-6 Dx at RoCEv2 100 Gb/s, `valkey-benchmark --rdma -c 1 -P 1 -t get`, 30 paired repetitions), where inlining the server's reply lowered the median GET latency by 0.448 us [0.348, 0.550] (95% interval). What it is worth in brpc has not been measured. On mlx5 the QP already gets a large inline size without asking, because `max_send_sge` is the device maximum, so the send queue does not grow there. - Breaking backward compatibility: no. `cut_into_sglist_and_iobuf()` (private) takes one more argument. ### Tests No RDMA hardware was available, so this was tested functionally on soft-RoCE (rdma_rxe) in a VM: - `libbrpc.a` built with `WITH_RDMA=ON` (CMake, gcc 15, dependencies from conda-forge), no new warnings. - The `rdma_performance` example server, and a small client that sends `PerfTestService.Test` with `echo_attachment` and compares the echoed attachment with what it sent: attachments of 0 to 300000 bytes (2000 calls per size up to 1000 bytes, 50 above), all equal. Requests of 58 to 111 bytes and responses of 37 to 85 bytes went inline (2 or 3 SGEs), larger messages by reference, checked with a temporary log line that is not part of this PR. - The retry without inline data was not exercised (rxe grants inline data). Unit tests were not run. ### Check List: - Please make sure your changes are compilable. - When providing us with a new feature, it is best to add related tests. - Please follow [Contributor Covenant Code of Conduct](https://github.com/apache/brpc/blob/master/CODE_OF_CONDUCT.md). This change was prepared with Claude Code (Claude Opus 5.5) and is marked `Generated-by:` in the commit message, as the ASF's generative tooling guidance suggests. It is a draft; the submitter will review it before marking it ready. Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
