Repository navigation
Add a cycle-accounting performance monitor and sweep tools - #1960
Open
davidharrishmc wants to merge 3 commits into
Open
davidharrishmc wants to merge 3 commits into
davidharrishmc wants to merge 3 commits into
Conversation
testbench/common/perfmon.sv (enabled with wsim --params PERF_MONITOR=1) charges every non-retiring cycle of each ELF to one cause: LSU or IFU stall subcause, trap, WFI, or the bubble cause carried down a shadow tag pipeline (mispredict, CSR write, fence, load-use, MDU, divide, ...). It also records stall episodes, cache fills, structural stalls without a true integer dependency, and the HPM events over the same window, and writes per-ELF and per-PC CSVs plus PERFANOM lines for anything it cannot explain. bin/perfsweep runs the arch-test ELF directories (and optionally CoreMark and Embench) with the monitor; bin/perfanalyze.py checks the accounting, cross-checks the HPM counters, and reports lost cycles by cause and site. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
…stall cause is active An integer instruction whose rs field matches an FP load's rd usually also has a real FP dependency on it (FPUStallD), so the bubble was not avoidable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds an optional testbench performance monitor (
wsim --params PERF_MONITOR=1, off by default) that charges every non-retiring cycle of each ELF to one cause, using LSU/IFU stall subcauses, traps, WFI, and a shadow tag pipeline for bubble causes (mispredicts, CSR writes, fences, load-use, MDU, divide). It also records cache fills and stall episodes, finds structural stalls without a true integer dependency, and samples the HPM events over the same window.bin/perfsweepruns the arch-test ELF directories (plus CoreMark and Embench with-b) through it, andbin/perfanalyze.pychecks the accounting, cross-checks the HPM counters, and ranks lost cycles by cause and site, so stalls that point to RTL defects stand out. On rv64gc and rv32gc, every cycle of all 2,426 arch-test and benchmark ELFs is accounted for. The sweep found that the D$ miss counter also counts D$ flushes and CMOs (fix in a separate PR) and thatMatchDEstalls on register fields the decode-stage instruction does not read.🤖 Generated with Claude Code