Skip to content

feat(loads): support PMA-gated misaligned integer loads - #14

Merged
RossComputerGuy merged 2 commits into
LilithSemi:masterfrom
murdoa:perf/misaligned-integer-loads
Oct 9, 2026
Merged

RossComputerGuy merged 2 commits into
LilithSemi:masterfrom
murdoa:perf/misaligned-integer-loads

Conversation

@murdoa

@murdoa murdoa commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Misaligned integer loads currently trap even when their bytes are in ordinary RAM. Add an opt-in hardware path that reads one aligned beat, or two beats when necessary, instead of requiring software emulation.

Both in-order executors use the same sequencer. Each beat is translated independently, and the MMU permits the widened read only when the complete physical beat belongs to an explicitly described, readable, idempotent RAM PMA supporting that access width and misaligned accesses. Partial data stays private until success; failures preserve the destination and report the faulting portion's virtual address.

The path supports uncached, non-H in-order RV32/RV64 Bare and RV64 Sv39 configurations, without configured PMP or PBMT. Aligned transactions stay unchanged, unsupported configurations retain their alignment traps, and stores, atomics and FP loads are not expanded. Split loads do not gain an atomicity guarantee.

Validation:

  • 515 identical baseline/patched cases: 371 passes plus 144 new-feature expectation failures upstream → 515 passes. All baseline passes preserved; upstream's alignment traps are legal.
  • 19 additional sequencer/configuration tests cover cancellation, reset, held responses, back-to-back requests, wrapping addresses and invalid PMAs.
  • Coverage includes every byte offset and integer load width, signed/unsigned results, compressed loads, x0, MMIO rejection, PMA boundaries, noncontiguous Sv39 pages, first/second faults and recovery.
  • Bare halfword-load simulation versus the same program's trap handler: static execution drops from 260/263 to 45/54 cycles for same/split-beat loads; microcoded execution drops from 805–883 to 105/114 cycles. These are simulation measurements, not FPGA timing claims.
  • SystemVerilog emitted for all six supported width/paging/executor combinations; no new analyzer diagnostics.

@RossComputerGuy
RossComputerGuy merged commit 1edb95f into LilithSemi:master Oct 9, 2026
1 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants