Felix-Gong opened a new pull request, #3571:
URL: https://github.com/apache/brpc/pull/3571

   ### Summary
   
   `PR3332` introduced a RISC-V Zvbc vector CRC32C path, but its implementation 
kept per-lane **scalar round-trips** (vector→scalar xor→vector) in the hot loop 
and was **hardcoded VLEN=128** (`vsetvl_e64m1(2)`), so VLEN≥256 cores used only 
half the vector width. On Spacemit X100 the Zvbc path measured ~450–600 MB/s, 
**slower than the scalar Zbc path (~1.5 GB/s)** — making the existing "Zvbc 
preferred" dispatch a ~2.5× regression there.
   
   This PR **rewrites `rv_crc32c_vclmul`**:
   
   - **Adaptive VLEN**: one `e64m1` pair at VLEN≥256 (4 lanes), two pairs at 
VLEN=128 (4 lanes carried on 2 vectors) — the vector path now runs at full 
width on 128-bit cores too.
   - **Pure-vector main loop**: `vlseg2e64` de-interleaved loads + `vclmul_vx` 
broadcast constants (k1..k4), **zero scalar extraction**.
   - **All-vector reduction**: single-element vector `vclmul` for the 4-to-1 
lane merge and Barrett reduction — removes the scalar `clmul` dependency, so 
the path is **safe on cores with Zvbc but no scalar Zbc (e.g. SG2044)**.
   - Same fold constants and math as the scalar core → **bit-exact** with the 
existing table/scalar implementations.
   
   Note: the CRC32C dispatch order (scalar-first vs vector-first) is 
deliberately **left untouched** here and will be addressed in a separate, 
focused PR.
   
   ### Correctness
   
   k3 (Spacemit X100, QEMU-free):
   - RFC3720 vectors (4 standard cases) — bit-consistent table/clmul/vclmul
   - Extend composition (`CRC("hello ")+"world" == CRC("hello world")`)
   - 70 boundary/misalign/large combos (63 B … 1 MiB × misalign 0/1/3/7/15, 
non-zero init)
   - **Official `CRC.*` gtest suite: 5/5 PASSED** (full `test_butil` build with 
`-DWITH_RISCV_ZVBC=ON -DBUILD_UNIT_TESTS=ON`)
   
   ### Performance
   
   k3 (X100, VLEN=256, `-O3 -march=rv64gc_zbc_zbb_zvbc`, warmup 10 + 100-round 
median — same harness/methodology as PR3312):
   
   | len | PR3332 vclmul (MB/s) | this (MB/s) | speedup |
   |-----|----------------------|-------------|---------|
   | 64 B   | 454 | 714 | 1.6× |
   | 1 KiB  | 578 | 5895 | **10.2×** |
   | 4 KiB  | 590 | 9400 | **15.9×** |
   | 64 KiB | 600 | 11502 | **19.2×** |
   | 1 MiB  | 508 | 11708 | **23.1×** |
   
   ### Test plan
   
   - [x] `test_butil` full build (unit tests on, Zvbc on)
   - [x] `CRC.*` official gtest 5/5 PASSED
   - [x] RFC vectors + 70 boundary combos + KAT bit-consistent across 
table/clmul/vclmul
   - [ ] VLEN=128 (SG2044-class) hardware run — pending platform access


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to