composo-ai
Popular repositories Loading
-
llm-judge-criteria-ensembling
llm-judge-criteria-ensembling PublicAn Empirical Investigation of Practical LLM-as-a-Judge Improvement Techniques on RewardBench 2
TeX 11
-
omission-bench
omission-bench PublicOmissionBench harness: code behind the two companion papers on omission blindness in LLM judges of AI clinical notes (data on Hugging Face: ComposoAI/OmissionBench)
Python 3
-
judge-uncertainty-decomposition
judge-uncertainty-decomposition PublicCode and data for "Decomposing LLM-Judge Uncertainty to Target Expert Labels" (arXiv:2609.06444)
Python 1
Repositories
- judge-uncertainty-decomposition Public
Code and data for "Decomposing LLM-Judge Uncertainty to Target Expert Labels" (arXiv:2609.06444)
- omission-bench Public
OmissionBench harness: code behind the two companion papers on omission blindness in LLM judges of AI clinical notes (data on Hugging Face: ComposoAI/OmissionBench)
- PRIME Public
PrimeBench: an open benchmark for evaluating evaluators. 400 paired responses across finance, news, biomedical and technical support, edited along conflicting criteria to test whether a reward model or LLM judge can tell better from plausible-but-worse.
- llm-judge-criteria-ensembling Public
An Empirical Investigation of Practical LLM-as-a-Judge Improvement Techniques on RewardBench 2
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…