Arrakis is a library to conduct, track and visualize mechanistic interpretability experiments.
-
Updated
Jul 8, 2026 - Jupyter Notebook
Arrakis is a library to conduct, track and visualize mechanistic interpretability experiments.
[NeurIPS 2025 MechInterp Workshop - Spotlight] Official implementation of the paper "RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching"
Lightweight representation engineering dataflow operations for agent developers.
Real-time 3D visualisation of SAE feature activations inside GPT-2, token by token
A minimal decoder probing GPT-2's residual stream activations for mechanistic interpretability research .
Mechanistic interpretability study: identifying the 4 attention heads (out of 144) that control 83% of subject-verb agreement in GPT-2 Small
Open-source EU AI Act Annex IV documentation toolkit. Mechanistic interpretability + circuit discovery for transformers. One function call generates a structured, hash-chained evidence package.
Investigating whether language models encode anticipated social consequences in their activations. Uses a 2x2 factorial design crossing truth × social valence to show that models are more sensitive to expected approval/disapproval than to truth itself.
Training and exploration of linear probes into Othello-GPT by Li et al. (2022)
Implementation and analysis of Sparse Autoencoders for neural network interpretability research. Features interactive visualization dashboard and W&B integration.
Evaluating how a model 'knowing what it knows' changes from base to instruct
Testing role-based pathways on small LLMs
Browser-based mechanistic interpretability toolkit for GPT-2. Visualize attention patterns, extract hidden states, and experiment with steering vectors — all running locally via WebGPU.
Does Quantization Kill Interpretability? Scaling study across 5 models (124M-2.8B): RTN destroys induction heads in small models, GPTQ preserves them at all scales.
CLI and MCP server that traces which attention heads and neurons drive an LLM's output, built on TransformerLens.
Mechanistic interpretability toolkit for comparing transformer activations, token shifts, and activation patching behaviour.
Probing whether induction head universality is truly architecture-agnostic or whether model family leaves a structural fingerprint on the circuits that enable in-context learning.
Mechanistic interpretability and safety auditing for Decision Transformers. Features circuit mapping, causal patching, TopK SAEs, and behavioral steering.
A Flax-based library for examining transformers, based on TransformerLens.
Mechanistic study of the refusal direction across base, instruction-tuned, and reasoning-distilled Qwen2.5-1.5B variants: extraction, ablation, transplant, and phase-aware analysis.
Add a description, image, and links to the transformerlens topic page so that developers can more easily learn about it.
To associate your repository with the transformerlens topic, visit your repo's landing page and select "manage topics."