LLM ベンチマークの汚染を、検定・信頼区間・検出力・多重比較補正つきで測る。JMMLU での実測と「言えないこと」の記録つき(Python 標準ライブラリのみ)/ Measuring LLM benchmark contamination with hypothesis tests, CIs, power analysis and multiplicity correction (JMMLU).
benchmark statistics japanese hypothesis-testing preregistration data-contamination statistical-power llm llm-evaluation jmmlu
-
Updated
Sep 15, 2026 - Python