Skip to content

About

A focused guide to extracting and validating concept and steering vectors from neural model activations.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Concept Vector Extraction

A focused, source-checked guide to methods that extract internal directions, subspaces, or features for reading and steering neural models, with an emphasis on large language models.

The central question is practical: given a behavior or concept, how did each paper turn examples or model states into something that can be measured or intervened on?

Scope

Included methods must extract a representation from model activations and explain how it is obtained. The main list covers vectors, but also includes closely related subspaces, soft projection matrices, sparse features, and transport maps when they are useful alternatives to forcing a concept into one direction.

Prompt engineering, weight-space model merging, soft prompts, and ordinary fine-tuning are out of scope unless they directly learn or evaluate an activation-space intervention.

The machine-readable extraction catalog is in data/methods.csv. A separate catalog of reliability and failure studies is in data/reliability.csv. Selection rules and a compact evaluation protocol are in docs/methodology.md.

Quick map

If you have... Start with... Why
One or a few contrastive prompt pairs ActAdd Fastest training-free directional check.
A matched positive/negative dataset CAA or difference-in-means Strong, transparent baseline that is hard to accidentally overcomplicate.
Labeled activations TCAV or ITI-style probe selection Tests linear separability and can identify useful layers or heads.
In-context task demonstrations ICV, task vectors, or function vectors Extracts task information already induced by demonstrations.
A behavior that is not well described by one line Conceptors, SAEs, or Activation Transport Represents a region, feature dictionary, or distributional map instead of one direction.
Enough data to optimize an intervention BiPO or ReFT-r1 Learns the intervention against an explicit behavioral objective.
A verbal label but no contrastive dataset Jacobian Lens / J-space Builds a direction for each vocabulary token from averaged local causal effects.
A broad recurring feature in an unlabeled corpus Sparse-autoencoder features Learns a feature dictionary, then uses selected decoder directions for analysis or steering.

Methods landscape

Foundational concept directions

Year Method How the representation is extracted Paper Official code
2018 TCAV Train a linear classifier to separate user-supplied concept examples from random examples at a chosen layer, then use the classifier normal as the concept direction. Paper tensorflow/tcav
2022 Latent steering vectors Optimize a continuous latent vector while keeping the language model frozen so that injecting the vector reconstructs a target sentence. Paper nishantsubramani/steering_vectors
2023 ActAdd Subtract the intermediate activations produced by a negative prompt from those produced by a positive prompt at a chosen layer and token position. Paper montemac/activation_additions
2023 Representation Engineering / LAT Collect activations from contrastive positive and negative statements, form paired differences, and extract a layerwise direction with PCA or another linear readout. Paper andyzoujm/representation-engineering
2023 Mean-centred steering Average activations for the target set and subtract the average activation over the broader training set. Paper Not located
2024 Contrastive Activation Addition (CAA) Average residual-stream activation differences across matched positive and negative behavioral examples at each layer. Paper nrimsky/CAA
2024 Refusal direction Subtract the mean residual activation for harmless instructions from the mean for harmful instructions, then causally test addition and removal of that direction. Paper andyrdt/refusal_direction
2025 Instruction steering Compute the activation difference between paired inputs with and without a particular instruction, then select the layer and scale on held-out development data. Paper microsoft/llm-steer-instruct

Probes and causally localized task vectors

Year Method How the representation is extracted Paper Official code
2023 Inference-Time Intervention (ITI) Rank attention heads by held-out linear-probe accuracy for truthfulness and derive a truthful class-separation direction at the selected heads. Paper likenneth/honest_llama
2023 In-context task vectors Take the hidden state induced by an in-context training set at a causally validated layer and position, then inject it into a query-only run. Paper roeehendel/icl_task_vectors
2024 Function vectors Use causal mediation to select attention heads, average each selected head's task-conditioned activation over demonstrations, and combine their outputs into one task vector. Paper ericwtodd/function_vectors
2024 In-context vectors (ICV) Run paired demonstration inputs and outputs separately and use the leading principal component of their hidden-state differences as the task direction. Paper shengliu66/ICV

Anthropic methods and newer extraction choices

Year Method How the representation is extracted Paper Official code
2024 Scaling Monosemanticity / Golden Gate Claude Train a sparse autoencoder on model activations and use a selected decoder column as a feature direction that can be clamped during generation. Paper Not located
2025 Anthropic concept-injection vectors Subtract residual activations for a concept prompt and a matched control prompt, then inject the difference and ask whether the model can identify it. Paper Not located
2025 Persona vectors Average activations from responses that exhibit a personality trait and subtract the mean activations from responses that oppose it. Paper safety-research/persona_vectors
2026 Assistant Axis Subtract the mean activations of many role-conditioned responses from default-assistant activations to define a layerwise assistant-identity axis. Paper safety-research/assistant-axis
2026 Emotion concept vectors Generate stories expressing named emotions and compare their internal activations to identify the layerwise patterns associated with each emotion. Paper Not located
2026 Pre-CoT answer directions Train linear probes on residual states immediately before chain-of-thought begins and use the answer-separating directions for steering. Paper Not located
2026 Jacobian Lens / J-space Average each token logit's Jacobian with respect to residual activations across contexts to obtain one verbalizable direction per token. Paper anthropics/jacobian-lens
2026 Source selection and tail subtraction Read a contrastive direction at a causally relevant execution boundary and subtract shared continuation activity to reduce prompt and position contamination. Paper Not located

Learned, adaptive, and non-vector alternatives

Year Method How the representation is extracted Paper Official code
2024 BiPO Treat the steering direction as a trainable parameter and optimize a bidirectional preference loss that rewards desired behavior under positive steering and the reverse under negative steering. Paper CaoYuanpu/BiPO
2024 Conditional Activation Steering (CAST) Extract behavior and condition directions from contrastive examples, then apply the behavior direction only when the input projects past a learned condition threshold. Paper IBM/activation-steering
2024 SADI Use activation differences from contrastive pairs to locate important heads, states, and neurons, then build an input-dependent element-wise intervention from the current semantics. Paper weixuan-wang123/SADI
2024 Conceptor steering Fit a regularized soft projection matrix to a collection of concept activations so the intervention represents an ellipsoidal activation region rather than one vector. Paper jorispos/ConceptorSteering
2025 Activation Transport (AcT) Estimate an optimal-transport map between source and target activation distributions and use the resulting map to move new activations toward the target distribution. Paper apple/ml-act
2025 SAE-Targeted Steering Learn a linear model of how residual-stream interventions change sparse-autoencoder features, then optimize a vector that increases chosen features while limiting off-target effects. Paper slavachalnev/SAE-TS
2025 ReFT-r1 / AxBench Learn a rank-one representation intervention from weakly supervised concept examples and evaluate it jointly on concept detection and generation steering. Paper stanfordnlp/axbench
2025 GCAV Compress layerwise CAVs, align them across layers with contrastive learning, and fuse them with attention into one global concept representation. Paper Zhenghao-He/GCAV

Practical libraries

Reliability and failure studies

Extraction, behavioral steering, selectivity, and self-report are different questions. These papers are especially useful for deciding whether an apparent concept vector is reliable; the same entries are available in data/reliability.csv.

Year Study Main lesson Paper Resource
2019 Probe control tasks Random-label controls reveal when a probe is memorizing rather than exposing a meaningful concept. Paper Not located
2021 Amnesic Probing Decodability and causal use can diverge, while erasure may damage correlated information. Paper Not located
2025 Steering off Course Many method/model combinations do not improve behavior or make it worse. Paper Code
2025 Causal probing reliability Causal probes trade off completeness and selectivity, so a changed output does not identify the intended variable by itself. Paper Project
2025 Anthropic feature-steering evaluation Strong sparse-feature steering can cause off-target behavior and capability loss. Study Not located
2025 Detecting the Disturbance Models may detect that an activation was perturbed without reliably identifying the injected concept. Paper Code
2026 Latent Introspection Better elicitation can substantially improve injected-concept reports, so a negative report may be an interface failure. Paper Code
2026 Introspection reality check Input-only and anomaly controls weaken some apparent privileged-access effects. Paper Not located
2026 Steering-induced misalignment Targeted vectors can produce broad persona and safety side effects that must not be mistaken for concept-specific control. Paper Not located

Four questions to keep separate

  1. Extraction: how was the vector, feature, subspace, or map obtained?
  2. Behavior: does intervening on it change the intended output?
  3. Selectivity: does the change survive random, shuffled, opposite-direction, and damage controls?
  4. Reportability: can the model itself or an external decoder identify the internal content, and under what access conditions?

A paper can answer one of these without answering the others. In particular, probe accuracy is not causal use, behavioral steering is not specific steering, and external decoding is not the same as native verbal report.

Reliability in one paragraph

A direction is not reliable merely because a probe can decode a label from it. Extract on one split, choose layer and strength on a second split, and report final effects on a third; lock the sign before evaluation; compare with zero, random, shuffled-label, and opposite directions; measure both the intended behavior and off-target damage; verify that the intervention actually changed the intended internal state; and reproduce a small fixed shard on another device or process. See the full reliability checklist.

Related curated lists

These projects influenced the simple taxonomy-and-summary organization used here:

Contributing

Corrections and additions are welcome. Please read CONTRIBUTING.md; every row needs a primary source and an independently written one-sentence summary, with official code or project links when they can be verified.

License

The original summaries and repository code are available under the MIT License. Papers, linked repositories, names, and trademarks remain the property of their respective owners.

About

A focused guide to extracting and validating concept and steering vectors from neural model activations.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages