JAMO — OCR-mediated Hangul rendering benchmark. Code companion to huggingface.co/datasets/Nasser4963/jamo-gold
-
Updated
Aug 17, 2026 - Python
JAMO — OCR-mediated Hangul rendering benchmark. Code companion to huggingface.co/datasets/Nasser4963/jamo-gold
Code, data, and raw model outputs for "Failure Modes in Perturbation-Based Measurement of Language Model Reliability". evaluation/reproduce.py re-derives every reported number offline, no API key required.
Audit of territorial access to public services across 24 organisation types and 64,105 entries in France's national administrative directory, using descriptive statistics, outlier sensitivity analysis and population-normalised indicators.
A controlled benchmark of eight React state management configurations, with the full replication package: cross-hardware, cross-engine, and a row sweep from 10 to 1000 rows.
Measurement validity, construct validity, and unsupported claims derived from AI agent telemetry.
To associate your repository with the measurement-validity topic, visit your repo's landing page and select "manage topics."