Skip to content
#

mt-bench

Here are 8 public repositories matching this topic...

Language: All
Filter by language
judge-calibration

Do LLM judges know when they're wrong? Three open-weight judges (Qwen2.5-7B, kev-8b, auto-j-13b) graded against MT-Bench human votes: calibration, position and padding attacks, and what auto-accepting confident verdicts would cost. Interactive site included.

  • Updated Sep 25, 2026
  • Python

Add this topic to your repo

To associate your repository with the mt-bench topic, visit your repo's landing page and select "manage topics."

Learn more