Open source RTL simulation acceleration on commodity hardware
-
Updated
Apr 13, 2023 - Python
Open source RTL simulation acceleration on commodity hardware
Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
KV260 integration lane for PCCX™ v002 LLM IP-core bring-up, validation, and board/runtime evidence.
NPUWattch: ML-based Power, Area, and Timing Modeling for Neural Accelerators
INT8 CNN inference datapath for real-time video frame interpolation (Open-Frame-gen, JCSSE 2026). Hand-written Verilog, bit-exact cocotb verification, synthesized on Artix-7 XC7A100T with Vivado 2026.1.
INT8 16x16 GEMM tile accelerator for LLM FFN inference (FPGA → MPW ready).
Hardware-agnostic AI compiler suite. Compile GGUF, ONNX, PyTorch, and SafeTensors models onto FPGAs, analog circuits, MCU swarms, photonic MZI meshes, neuromorphic chips, and CIM accelerators — not GPUs. Includes SiL emulator, real-time dashboard, federated learning, and carbon-aware compilation.
Neuromorphic SNN Accelerator: 4-Stage LIF PE, On-Chip STDP, Spike-Driven FlashAttention, Banked BRAM Arbiter, and 6.3x Simulated EDP Reduction (Kria KV260)
Silicon as a Distributed System: Closing the Cross-Layer Loop from Spatial Manufacturing Defects to Wafer-Level Economics
HW/SW co-design bridge from CIFAR-10 CNN training pipeline to cycle-accurate ASIC simulation in pure Python via Amaranth HDL. Pytorch · Amaranth · CIFAR-10 · Tiny-ImageNet
Reproducible simulator, hardware models, paper figures, and reference RTL for CAMformer.
Hand-written SystemVerilog SIMD core (Fall 2025) followed by an Allo DSL/HLS redesign adding INT8 GEMM, ReLU, and fusion kernels (Spring 2026), targeting the Nexys A7.
Hardware accelerator for Capsule Neural Networks
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Defensive publication: executable reference for conflict-free streaming stencils with single-port banked SRAM, multicast delivery, and ping-pong buffering.
To associate your repository with the hardware-accelerator topic, visit your repo's landing page and select "manage topics."