A focused, source-checked guide to methods that extract internal directions, subspaces, or features for reading and steering neural models, with an emphasis on large language models.
The central question is practical: given a behavior or concept, how did each paper turn examples or model states into something that can be measured or intervened on?
Included methods must extract a representation from model activations and explain how it is obtained. The main list covers vectors, but also includes closely related subspaces, soft projection matrices, sparse features, and transport maps when they are useful alternatives to forcing a concept into one direction.
Prompt engineering, weight-space model merging, soft prompts, and ordinary fine-tuning are out of scope unless they directly learn or evaluate an activation-space intervention.
The machine-readable extraction catalog is in data/methods.csv. A separate catalog of reliability and failure studies is in data/reliability.csv. Selection rules and a compact evaluation protocol are in docs/methodology.md.
| If you have... | Start with... | Why |
|---|---|---|
| One or a few contrastive prompt pairs | ActAdd | Fastest training-free directional check. |
| A matched positive/negative dataset | CAA or difference-in-means | Strong, transparent baseline that is hard to accidentally overcomplicate. |
| Labeled activations | TCAV or ITI-style probe selection | Tests linear separability and can identify useful layers or heads. |
| In-context task demonstrations | ICV, task vectors, or function vectors | Extracts task information already induced by demonstrations. |
| A behavior that is not well described by one line | Conceptors, SAEs, or Activation Transport | Represents a region, feature dictionary, or distributional map instead of one direction. |
| Enough data to optimize an intervention | BiPO or ReFT-r1 | Learns the intervention against an explicit behavioral objective. |
| A verbal label but no contrastive dataset | Jacobian Lens / J-space | Builds a direction for each vocabulary token from averaged local causal effects. |
| A broad recurring feature in an unlabeled corpus | Sparse-autoencoder features | Learns a feature dictionary, then uses selected decoder directions for analysis or steering. |
| Year | Method | How the representation is extracted | Paper | Official code |
|---|---|---|---|---|
| 2018 | TCAV | Train a linear classifier to separate user-supplied concept examples from random examples at a chosen layer, then use the classifier normal as the concept direction. | Paper | tensorflow/tcav |
| 2022 | Latent steering vectors | Optimize a continuous latent vector while keeping the language model frozen so that injecting the vector reconstructs a target sentence. | Paper | nishantsubramani/steering_vectors |
| 2023 | ActAdd | Subtract the intermediate activations produced by a negative prompt from those produced by a positive prompt at a chosen layer and token position. | Paper | montemac/activation_additions |
| 2023 | Representation Engineering / LAT | Collect activations from contrastive positive and negative statements, form paired differences, and extract a layerwise direction with PCA or another linear readout. | Paper | andyzoujm/representation-engineering |
| 2023 | Mean-centred steering | Average activations for the target set and subtract the average activation over the broader training set. | Paper | Not located |
| 2024 | Contrastive Activation Addition (CAA) | Average residual-stream activation differences across matched positive and negative behavioral examples at each layer. | Paper | nrimsky/CAA |
| 2024 | Refusal direction | Subtract the mean residual activation for harmless instructions from the mean for harmful instructions, then causally test addition and removal of that direction. | Paper | andyrdt/refusal_direction |
| 2025 | Instruction steering | Compute the activation difference between paired inputs with and without a particular instruction, then select the layer and scale on held-out development data. | Paper | microsoft/llm-steer-instruct |
| Year | Method | How the representation is extracted | Paper | Official code |
|---|---|---|---|---|
| 2023 | Inference-Time Intervention (ITI) | Rank attention heads by held-out linear-probe accuracy for truthfulness and derive a truthful class-separation direction at the selected heads. | Paper | likenneth/honest_llama |
| 2023 | In-context task vectors | Take the hidden state induced by an in-context training set at a causally validated layer and position, then inject it into a query-only run. | Paper | roeehendel/icl_task_vectors |
| 2024 | Function vectors | Use causal mediation to select attention heads, average each selected head's task-conditioned activation over demonstrations, and combine their outputs into one task vector. | Paper | ericwtodd/function_vectors |
| 2024 | In-context vectors (ICV) | Run paired demonstration inputs and outputs separately and use the leading principal component of their hidden-state differences as the task direction. | Paper | shengliu66/ICV |
| Year | Method | How the representation is extracted | Paper | Official code |
|---|---|---|---|---|
| 2024 | Scaling Monosemanticity / Golden Gate Claude | Train a sparse autoencoder on model activations and use a selected decoder column as a feature direction that can be clamped during generation. | Paper | Not located |
| 2025 | Anthropic concept-injection vectors | Subtract residual activations for a concept prompt and a matched control prompt, then inject the difference and ask whether the model can identify it. | Paper | Not located |
| 2025 | Persona vectors | Average activations from responses that exhibit a personality trait and subtract the mean activations from responses that oppose it. | Paper | safety-research/persona_vectors |
| 2026 | Assistant Axis | Subtract the mean activations of many role-conditioned responses from default-assistant activations to define a layerwise assistant-identity axis. | Paper | safety-research/assistant-axis |
| 2026 | Emotion concept vectors | Generate stories expressing named emotions and compare their internal activations to identify the layerwise patterns associated with each emotion. | Paper | Not located |
| 2026 | Pre-CoT answer directions | Train linear probes on residual states immediately before chain-of-thought begins and use the answer-separating directions for steering. | Paper | Not located |
| 2026 | Jacobian Lens / J-space | Average each token logit's Jacobian with respect to residual activations across contexts to obtain one verbalizable direction per token. | Paper | anthropics/jacobian-lens |
| 2026 | Source selection and tail subtraction | Read a contrastive direction at a causally relevant execution boundary and subtract shared continuation activity to reduce prompt and position contamination. | Paper | Not located |
| Year | Method | How the representation is extracted | Paper | Official code |
|---|---|---|---|---|
| 2024 | BiPO | Treat the steering direction as a trainable parameter and optimize a bidirectional preference loss that rewards desired behavior under positive steering and the reverse under negative steering. | Paper | CaoYuanpu/BiPO |
| 2024 | Conditional Activation Steering (CAST) | Extract behavior and condition directions from contrastive examples, then apply the behavior direction only when the input projects past a learned condition threshold. | Paper | IBM/activation-steering |
| 2024 | SADI | Use activation differences from contrastive pairs to locate important heads, states, and neurons, then build an input-dependent element-wise intervention from the current semantics. | Paper | weixuan-wang123/SADI |
| 2024 | Conceptor steering | Fit a regularized soft projection matrix to a collection of concept activations so the intervention represents an ellipsoidal activation region rather than one vector. | Paper | jorispos/ConceptorSteering |
| 2025 | Activation Transport (AcT) | Estimate an optimal-transport map between source and target activation distributions and use the resulting map to move new activations toward the target distribution. | Paper | apple/ml-act |
| 2025 | SAE-Targeted Steering | Learn a linear model of how residual-stream interventions change sparse-autoencoder features, then optimize a vector that increases chosen features while limiting off-target effects. | Paper | slavachalnev/SAE-TS |
| 2025 | ReFT-r1 / AxBench | Learn a rank-one representation intervention from weakly supervised concept examples and evaluate it jointly on concept detection and generation steering. | Paper | stanfordnlp/axbench |
| 2025 | GCAV | Compress layerwise CAVs, align them across layers with contrastive learning, and fuse them with attention into one global concept representation. | Paper | Zhenghao-He/GCAV |
- steering-vectors/steering-vectors — a small PyTorch/Hugging Face library for training and applying steering vectors.
- IBM/activation-steering — extraction and conditional steering with contrastive examples.
- vgel/repeng — a practical representation-engineering pipeline for Hugging Face models.
- decoderesearch/SAELens — training and analysis tools for sparse autoencoders on language-model activations.
- stanfordnlp/pyreft — representation fine-tuning and intervention tooling.
- stanfordnlp/axbench — common datasets and evaluation code for detection and steering methods.
- anthropics/jacobian-lens — constructs and analyzes verbalizable residual-stream directions using averaged Jacobians.
- stanfordnlp/pyvene — reusable hooks for static, trainable, and interchange interventions.
- atticusg/Interchange — reference implementation for causal-abstraction and interchange-intervention experiments.
- aypan17/latentqa — trains question-answering decoders over hidden activations; useful as an external-readout positive control.
- kmeng01/rome — causal-tracing and state-restoration patterns for localizing factual behavior.
- acsresearch/latent-introspection-code and elyhahami18/llama-introspection-new — open implementations for testing reports about injected internal concepts.
Extraction, behavioral steering, selectivity, and self-report are different questions. These papers are especially useful for deciding whether an apparent concept vector is reliable; the same entries are available in data/reliability.csv.
| Year | Study | Main lesson | Paper | Resource |
|---|---|---|---|---|
| 2019 | Probe control tasks | Random-label controls reveal when a probe is memorizing rather than exposing a meaningful concept. | Paper | Not located |
| 2021 | Amnesic Probing | Decodability and causal use can diverge, while erasure may damage correlated information. | Paper | Not located |
| 2025 | Steering off Course | Many method/model combinations do not improve behavior or make it worse. | Paper | Code |
| 2025 | Causal probing reliability | Causal probes trade off completeness and selectivity, so a changed output does not identify the intended variable by itself. | Paper | Project |
| 2025 | Anthropic feature-steering evaluation | Strong sparse-feature steering can cause off-target behavior and capability loss. | Study | Not located |
| 2025 | Detecting the Disturbance | Models may detect that an activation was perturbed without reliably identifying the injected concept. | Paper | Code |
| 2026 | Latent Introspection | Better elicitation can substantially improve injected-concept reports, so a negative report may be an interface failure. | Paper | Code |
| 2026 | Introspection reality check | Input-only and anomaly controls weaken some apparent privileged-access effects. | Paper | Not located |
| 2026 | Steering-induced misalignment | Targeted vectors can produce broad persona and safety side effects that must not be mistaken for concept-specific control. | Paper | Not located |
- Extraction: how was the vector, feature, subspace, or map obtained?
- Behavior: does intervening on it change the intended output?
- Selectivity: does the change survive random, shuffled, opposite-direction, and damage controls?
- Reportability: can the model itself or an external decoder identify the internal content, and under what access conditions?
A paper can answer one of these without answering the others. In particular, probe accuracy is not causal use, behavioral steering is not specific steering, and external decoding is not the same as native verbal report.
A direction is not reliable merely because a probe can decode a label from it. Extract on one split, choose layer and strength on a second split, and report final effects on a third; lock the sign before evaluation; compare with zero, random, shuffled-label, and opposite directions; measure both the intended behavior and off-target damage; verify that the intervention actually changed the intended internal state; and reproduce a small fixed shard on another device or process. See the full reliability checklist.
These projects influenced the simple taxonomy-and-summary organization used here:
- Awesome Activation Engineering
- Awesome Inference-Time Steering
- Awesome Mechanistic Interpretability LM Papers
Corrections and additions are welcome. Please read CONTRIBUTING.md; every row needs a primary source and an independently written one-sentence summary, with official code or project links when they can be verified.
The original summaries and repository code are available under the MIT License. Papers, linked repositories, names, and trademarks remain the property of their respective owners.