An agentic system that auto-optimizes LLM workloads on AMD GPUs.
-
Updated
Sep 4, 2026 - Python
An agentic system that auto-optimizes LLM workloads on AMD GPUs.
How-to guides and AMD/ROCm optimization recipes for open AI-for-science models.
Inference scaling benchmark of Qwen3.5-2B on AMD Instinct MI300X using ROCm and Hugging Face Transformers.
One sentence in, trained computer vision model out, zero human labels. An autonomous agent swarm (FLUX.2 - SAM 3 - Gemma 4) running entirely on one AMD Instinct MI300X (AMD Compute used). AMD Developer Hackathon ACT II.
Agentic AI store operations on AMD Instinct MI300X - Gemma tool-calling agents with enforced calculator grounding, human-in-the-loop approval, and a chat-first ops console. AMD ACT II Hackathon, Track 3.
LLM inference benchmarking dashboard: Python FastAPI backend with async orchestration, WebSocket live TTFT/TBT/throughput comparison across configs (512/128 to 4096/1024 tokens), Grafana + Docker Compose stack, GitHub Actions CI; 21/21 pytest passing.
On-prem AI coding agent + LLM cost-router for regulated enterprises — 100% AMD, open-weight, air-gapped, Ed25519 tamper-evident audit. AMD ACT I 1st-place winner.
Automated CUDA-to-ROCm GPU kernel translation via semantic BridgeIR. 316 HIP API mappings, wavefront-aware optimization, MFMA targeting, 6 multi-target backends. Break free from CUDA vendor lock-in.
Progressive CDNA kernel curriculum on MI300X — wavefronts, LDS bank conflicts, MFMA register mechanics (VGPR/SGPR/AGPR), tiled GEMM, XCD awareness, generated assembly
An AI-powered command center that correlates AWS FinOps data with SecOps vulnerabilities using Fireworks AI and AMD Instinct MI300X.
Standalone AMD ROCm/PyTorch tools for pruning, expanding, and continued-pretraining an LLM checkpoint on a single MI300X GPU
Zero-config LLM benchmarking on AMD GPUs with ROCm. Auto-detect MI300X/MI250/Radeon, CUDA and CPU fallback.
White paper & reproducible benchmark suite for LLM inference optimization on AMD MI300X using ROCm 6.1
Perseus Vault x AMD Instinct — encrypted, local-first agent memory kept off the GPU. AMD Developer Hackathon Act II (Unicorn Track).
To associate your repository with the mi300x topic, visit your repo's landing page and select "manage topics."