Repository navigation
Add FMA_LZA to choose an LZA or a leading zero counter for FMA normalization - #1966
Open
davidharrishmc wants to merge 2 commits into
Open
davidharrishmc wants to merge 2 commits into
davidharrishmc wants to merge 2 commits into
Conversation
The fourth input of the sign mux selected XE[LLEN-1]. LLEN used to equal FLEN wherever FPSIZES was 3 or 4, but Zacas makes LLEN 128 on RV64 for amocas.q, so the index ran past the 64-bit XE; synthesis rejected it (ELAB-298). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
…nter FMA_LZA = 1 (the default) keeps the leading zero anticipator, which computes the normalization count in parallel with the adder. FMA_LZA = 0 counts the leading zeros of the sum instead: smaller, but in series after the adder. The counter's count is exact, so the one-bit LZA correction never applies. The fdqh_lzc_rv64gc and fdqh_lzc_rv32gc derivatives run the testfloat add, sub, mul and fma vectors in the nightly regression. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
FMA_LZAchooses how the FMA computes its normalization count: 1 (the default, set in every config) keeps the leading zero anticipator, which works in parallel with the adder; 0 counts the leading zeros of the sum instead, which is smaller but in series after the adder. The counter's count is exact, so the one-bit LZA correction inshiftcorrectionnever applies.fdqh_lzc_rv64gcandfdqh_lzc_rv32gc(FMA_LZA = 0) pass the testfloat add, sub, mul and fma vectors and are added to the nightly testfloat runs; the standard regression passes with the default. The PR also fixesfpu.sv's fmv.x sign mux, which indexedXE[LLEN-1]and so ran pastXEonce Zacas made LLEN 128, failing synthesis of rva23's rv64gc.Impact in TSMC28 (tt 0.9 V, synthDC, rv64gc; the core is
syn_sram_rv64gc, caches in SRAM macros):SCntThe FMA's Execute path, which starts at the FP register address and goes through the operand forwarding before the FMA, is among the core's most critical paths. The LZC saves about 600 µm² of FPU (1.7%, 0.1% of the core) but uses up almost all of that path's margin, so it only pays off where the FMA has timing slack; LZA stays the default. Reproduce with
make -C synthDC synth DESIGN=wallypipelinedcore CONFIG=syn_sram_rv64gc MOD=lzc TECH=tsmc28 FREQ=850(MOD=orig for LZA); TSMC28 synthesis needs #1964 on machines whose DC setup does not defineSYN_pdk.🤖 Generated with Claude Code