OmissionBench harness: code behind the two companion papers on omission blindness in LLM judges of AI clinical notes (data on Hugging Face: ComposoAI/OmissionBench)
benchmark clinical-nlp medical-ai llm-evaluation hallucination-detection llm-as-a-judge ambient-scribe omission-detection
-
Updated
Sep 28, 2026 - Python