Confidence calibration toolkit for LLM verbalized-probability outputs. Real benchmark on 998 BoolQ questions with Llama-3.1-8B: ECE 0.148 -> 0.030, log-loss 3.9 -> 0.41.
-
Updated
May 21, 2026 - Python
Confidence calibration toolkit for LLM verbalized-probability outputs. Real benchmark on 998 BoolQ questions with Llama-3.1-8B: ECE 0.148 -> 0.030, log-loss 3.9 -> 0.41.
One-pass residual readouts vs fair closed-set AR scoring (BoolQ, RuleTaker, ARC)
Add a description, image, and links to the boolq topic page so that developers can more easily learn about it.
To associate your repository with the boolq topic, visit your repo's landing page and select "manage topics."