Rewrite RISC-V Zvbc CRC32C with adaptive VLEN and pure-vector folding - #3571
Open
Felix-Gong wants to merge 1 commit into
Open
Felix-Gong wants to merge 1 commit into
Felix-Gong wants to merge 1 commit into
Conversation
PR3332 introduced a Zvbc vector CRC32C path, but its implementation
kept per-lane scalar round-trips (vector->scalar xor->vector) in the
hot loop and fixed VLEN=128 (vsetvl_e64m1(2)), so VLEN>=256 cores
only used half the width. On Spacemit X100 the Zvbc path measured
~450-600 MB/s, slower than the scalar Zbc path (~1.5 GB/s), making
the "Zvbc preferred" dispatch a ~2.5x regression there.
Rewrite rv_crc32c_vclmul:
- adaptive VLEN: one e64m1 pair at VLEN>=256 (4 lanes) or two
pairs at VLEN=128 (4 lanes carried on 2 vectors), so 128-bit
cores run the vector path at full width;
- pure-vector main loop: vlseg2e64 de-interleaved loads and
vclmul_vx broadcast constants (k1..k4), zero scalar extraction;
- all-vector reduction: single-element vector vclmul for the
4-to-1 lane merge and Barrett reduction, removing scalar-clmul
dependencies (safe on cores with Zvbc but no scalar Zbc, e.g.
SG2044);
- same fold constants and math as the scalar core -> bit-exact.
Correctness on k3: RFC3720 vectors (4 standard PASS), Extend
composition, 70 boundary/misalign/large combos, and the official
CRC.* gtest suite (5/5 PASSED) all bit-consistent across
table/clmul/vclmul.
Performance on k3 (X100, VLEN=256, 100-round median, same harness as
PR3312):
len PR3332 vclmul (MB/s) this (MB/s) speedup
1 KiB 578 5895 10.2x
4 KiB 590 9400 15.9x
64 KiB 600 11502 19.2x
1 MiB 508 11708 23.1x
Signed-off-by: Xiaofei Gong <gongxiaofei24@iscas.ac.cn>
Signed-off-by: YuanSheng <yuansheng@isrc.iscas.ac.cn>
Contributor
|
LGTM |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
PR3332introduced a RISC-V Zvbc vector CRC32C path, but its implementation kept per-lane scalar round-trips (vector→scalar xor→vector) in the hot loop and was hardcoded VLEN=128 (vsetvl_e64m1(2)), so VLEN≥256 cores used only half the vector width. On Spacemit X100 the Zvbc path measured ~450–600 MB/s, slower than the scalar Zbc path (~1.5 GB/s) — making the existing "Zvbc preferred" dispatch a ~2.5× regression there.This PR rewrites
rv_crc32c_vclmul:e64m1pair at VLEN≥256 (4 lanes), two pairs at VLEN=128 (4 lanes carried on 2 vectors) — the vector path now runs at full width on 128-bit cores too.vlseg2e64de-interleaved loads +vclmul_vxbroadcast constants (k1..k4), zero scalar extraction.vclmulfor the 4-to-1 lane merge and Barrett reduction — removes the scalarclmuldependency, so the path is safe on cores with Zvbc but no scalar Zbc (e.g. SG2044).Correctness
k3 (Spacemit X100, QEMU-free):
CRC.*gtest suite: 5/5 PASSED (fulltest_butilbuild with-DWITH_RISCV_ZVBC=ON -DBUILD_UNIT_TESTS=ON)Performance
k3 (X100, VLEN=256,
-O3 -march=rv64gc_zbc_zbb_zvbc, warmup 10 + 100-round median — same harness/methodology as PR3312):Test plan
test_butilfull build (unit tests on, Zvbc on)CRC.*official gtest 5/5 PASSED