Skip to content
#

self-report

Here are 10 public repositories matching this topic...

Measures whether an LLM's self-reports agree with what its internals show. Injected concepts are perfectly recoverable by probes (AUC 1.000) and absent from the model's verbal report - report-state agreement fails in a graded way. Gemma-3-4B, Jacobian lens, ~3,750 trials.

  • Updated Aug 11, 2026
  • Python

Anchoring vignettes applied to language models for the first time. Placing reference exemplars before a self-assessment transfers their rating onto it; reversing the turn order removes the effect. 5 models, 900 observations, execution-verified ground truth, all thresholds preregistered.

  • Updated Aug 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the self-report topic, visit your repo's landing page and select "manage topics."

Learn more