Skip to content

Rewrite RISC-V Zvbc CRC32C with adaptive VLEN and pure-vector folding - #3571

Open
Felix-Gong wants to merge 1 commit into
apache:masterfrom
Felix-Gong:riscv-crc32c-zvbc-rewrite
Open

Felix-Gong wants to merge 1 commit into
apache:masterfrom
Felix-Gong:riscv-crc32c-zvbc-rewrite

Conversation

@Felix-Gong

@Felix-Gong Felix-Gong commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

PR3332 introduced a RISC-V Zvbc vector CRC32C path, but its implementation kept per-lane scalar round-trips (vector→scalar xor→vector) in the hot loop and was hardcoded VLEN=128 (vsetvl_e64m1(2)), so VLEN≥256 cores used only half the vector width. On Spacemit X100 the Zvbc path measured ~450–600 MB/s, slower than the scalar Zbc path (~1.5 GB/s) — making the existing "Zvbc preferred" dispatch a ~2.5× regression there.

This PR rewrites rv_crc32c_vclmul:

  • Adaptive VLEN: one e64m1 pair at VLEN≥256 (4 lanes), two pairs at VLEN=128 (4 lanes carried on 2 vectors) — the vector path now runs at full width on 128-bit cores too.
  • Pure-vector main loop: vlseg2e64 de-interleaved loads + vclmul_vx broadcast constants (k1..k4), zero scalar extraction.
  • All-vector reduction: single-element vector vclmul for the 4-to-1 lane merge and Barrett reduction — removes the scalar clmul dependency, so the path is safe on cores with Zvbc but no scalar Zbc (e.g. SG2044).
  • Same fold constants and math as the scalar core → bit-exact with the existing table/scalar implementations.

Correctness

k3 (Spacemit X100, QEMU-free):

  • RFC3720 vectors (4 standard cases) — bit-consistent table/clmul/vclmul
  • Extend composition
  • 70 boundary/misalign/large combos (63 B … 1 MiB × misalign 0/1/3/7/15, non-zero init)
  • Official CRC.* gtest suite: 5/5 PASSED (full test_butil build with -DWITH_RISCV_ZVBC=ON -DBUILD_UNIT_TESTS=ON)

Performance

k3 (X100, VLEN=256, -O3 -march=rv64gc_zbc_zbb_zvbc, warmup 10 + 100-round median — same harness/methodology as PR3312):

len PR3332 vclmul (MB/s) this (MB/s) speedup
64 B 454 714 1.6×
1 KiB 578 5895 10.2×
4 KiB 590 9400 15.9×
64 KiB 600 11502 19.2×
1 MiB 508 11708 23.1×

Test plan

  • test_butil full build (unit tests on, Zvbc on)
  • CRC.* official gtest 5/5 PASSED
  • RFC vectors + 70 boundary combos + KAT bit-consistent across table/clmul/vclmul
  • VLEN=128 (SG2044-class) hardware run — pending platform access

PR3332 introduced a Zvbc vector CRC32C path, but its implementation
kept per-lane scalar round-trips (vector->scalar xor->vector) in the
hot loop and fixed VLEN=128 (vsetvl_e64m1(2)), so VLEN>=256 cores
only used half the width. On Spacemit X100 the Zvbc path measured
~450-600 MB/s, slower than the scalar Zbc path (~1.5 GB/s), making
the "Zvbc preferred" dispatch a ~2.5x regression there.

Rewrite rv_crc32c_vclmul:

  - adaptive VLEN: one e64m1 pair at VLEN>=256 (4 lanes) or two
    pairs at VLEN=128 (4 lanes carried on 2 vectors), so 128-bit
    cores run the vector path at full width;
  - pure-vector main loop: vlseg2e64 de-interleaved loads and
    vclmul_vx broadcast constants (k1..k4), zero scalar extraction;
  - all-vector reduction: single-element vector vclmul for the
    4-to-1 lane merge and Barrett reduction, removing scalar-clmul
    dependencies (safe on cores with Zvbc but no scalar Zbc, e.g.
    SG2044);
  - same fold constants and math as the scalar core -> bit-exact.

Correctness on k3: RFC3720 vectors (4 standard PASS), Extend
composition, 70 boundary/misalign/large combos, and the official
CRC.* gtest suite (5/5 PASSED) all bit-consistent across
table/clmul/vclmul.

Performance on k3 (X100, VLEN=256, 100-round median, same harness as
PR3312):

    len      PR3332 vclmul (MB/s)   this (MB/s)   speedup
    1 KiB           578                 5895         10.2x
    4 KiB           590                 9400         15.9x
    64 KiB          600                11502         19.2x
    1 MiB           508                11708         23.1x

Signed-off-by: Xiaofei Gong <gongxiaofei24@iscas.ac.cn>
Signed-off-by: YuanSheng <yuansheng@isrc.iscas.ac.cn>
@chenBright
chenBright requested a lite review from Copilot September 29, 2026 09:02

This comment was marked as low quality.

@wwbmmm

wwbmmm commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

LGTM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants