Skip to content

Add FMA_LZA to choose an LZA or a leading zero counter for FMA normalization - #1966

Open
davidharrishmc wants to merge 2 commits into
openhwfoundation:rva23from
davidharrishmc:dh/fma-lza-param
Open

davidharrishmc wants to merge 2 commits into
openhwfoundation:rva23from
davidharrishmc:dh/fma-lza-param

Conversation

@davidharrishmc

@davidharrishmc davidharrishmc commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

FMA_LZA chooses how the FMA computes its normalization count: 1 (the default, set in every config) keeps the leading zero anticipator, which works in parallel with the adder; 0 counts the leading zeros of the sum instead, which is smaller but in series after the adder. The counter's count is exact, so the one-bit LZA correction in shiftcorrection never applies. fdqh_lzc_rv64gc and fdqh_lzc_rv32gc (FMA_LZA = 0) pass the testfloat add, sub, mul and fma vectors and are added to the nightly testfloat runs; the standard regression passes with the default. The PR also fixes fpu.sv's fmv.x sign mux, which indexed XE[LLEN-1] and so ran past XE once Zacas made LLEN 128, failing synthesis of rva23's rv64gc.

Impact in TSMC28 (tt 0.9 V, synthDC, rv64gc; the core is syn_sram_rv64gc, caches in SRAM macros):

LZA (FMA_LZA = 1) LZC (FMA_LZA = 0)
FMA alone, fastest achievable delay 0.556 ns 0.646 ns (+16%)
Normalization-count logic area (FMA alone, 300-1800 MHz targets) 550-840 µm² 180-590 µm²
Core at 850 MHz: worst slack (HPTW PTE register to D$ SRAM, both runs) -0.065 ns -0.057 ns
Core at 850 MHz: slack of the FMA path ending at SCnt -0.005 ns -0.048 ns
Core at 850 MHz: FPU / core area 36,037 / 480,956 µm² 35,414 / 480,515 µm²

The FMA's Execute path, which starts at the FP register address and goes through the operand forwarding before the FMA, is among the core's most critical paths. The LZC saves about 600 µm² of FPU (1.7%, 0.1% of the core) but uses up almost all of that path's margin, so it only pays off where the FMA has timing slack; LZA stays the default. Reproduce with make -C synthDC synth DESIGN=wallypipelinedcore CONFIG=syn_sram_rv64gc MOD=lzc TECH=tsmc28 FREQ=850 (MOD=orig for LZA); TSMC28 synthesis needs #1964 on machines whose DC setup does not define SYN_pdk.

🤖 Generated with Claude Code

davidharrishmc and others added 2 commits October 10, 2026 16:28
The fourth input of the sign mux selected XE[LLEN-1].  LLEN used to equal FLEN
wherever FPSIZES was 3 or 4, but Zacas makes LLEN 128 on RV64 for amocas.q, so
the index ran past the 64-bit XE; synthesis rejected it (ELAB-298).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: David Harris <David_Harris@hmc.edu>
…nter

FMA_LZA = 1 (the default) keeps the leading zero anticipator, which computes
the normalization count in parallel with the adder.  FMA_LZA = 0 counts the
leading zeros of the sum instead: smaller, but in series after the adder.  The
counter's count is exact, so the one-bit LZA correction never applies.  The
fdqh_lzc_rv64gc and fdqh_lzc_rv32gc derivatives run the testfloat add, sub,
mul and fma vectors in the nightly regression.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: David Harris <David_Harris@hmc.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant