Modified mt_bench with API and HF scripts for LFMs.
-
Updated
Jul 9, 2025 - Python
Modified mt_bench with API and HF scripts for LFMs.
Do LLM judges know when they're wrong? Three open-weight judges (Qwen2.5-7B, kev-8b, auto-j-13b) graded against MT-Bench human votes: calibration, position and padding attacks, and what auto-accepting confident verdicts would cost. Interactive site included.
Production-grade LLM-as-Judge evaluation framework with position, verbosity & self-enhancement bias mitigation. FastAPI + Streamlit + Python SDK.
⚖️ Dual-Judge: 让AI测试结果真正有说服力 | 双LLM交叉验证消除单模型偏见 | 独立于具体Agent的通用评估框架 | Making AI Evaluation Trustworthy
Reference Flutter chatbot + benchmark harness comparing on-device (Gemma) vs cloud (Claude) LLMs — MSc thesis artifact
Korean MT-Bench 문항, 모델 응답, LLM-as-a-Judge 판정 기록 및 논문 결과 재현 코드
Reproducible LLM-as-a-judge reliability lab: chance-corrected agreement (Cohen's kappa, Krippendorff's alpha) with bootstrap CIs, computed keyless from a committed MT-Bench snapshot and re-derived in CI as a drift gate.
Score an LLM judge against human labels: chance-corrected agreement with cluster-bootstrap intervals, the human-human ceiling it has to be read against, and a written drift protocol. Stdlib only, no API key.
To associate your repository with the mt-bench topic, visit your repo's landing page and select "manage topics."