Confidence calibration toolkit for LLM verbalized-probability outputs. Real benchmark on 998 BoolQ questions with Llama-3.1-8B: ECE 0.148 -> 0.030, log-loss 3.9 -> 0.41.
-
Updated
May 21, 2026 - Python
Confidence calibration toolkit for LLM verbalized-probability outputs. Real benchmark on 998 BoolQ questions with Llama-3.1-8B: ECE 0.148 -> 0.030, log-loss 3.9 -> 0.41.
Catch language while it’s still becoming culture. Hermes skill for memetic emergence — slang lineage, receipts, settled Brier only.
Add a description, image, and links to the brier topic page so that developers can more easily learn about it.
To associate your repository with the brier topic, visit your repo's landing page and select "manage topics."