Your model is graded against an answer key that contains mistakes, and nobody has counted them.
Labels come from people, and people make mistakes at a rate almost no team measures. This project measures that rate, tests whether a detector can find the bad labels, and answers the question a working team actually has to decide: is fixing the data worth more than improving the model?
Confident learning is a published method with a maintained library, and running it is easy. Three things here are not, and they are where the project earns its place:
- A hand audit with three verdicts. You read the top flagged examples yourself and record wrong, correct, or ambiguous. Precision is reported as a range, and a blind random control gives it something to be compared against.
- Labels against models, as an exchange rate. Not "which is better", but how many labels you must fix to match what a model upgrade bought.
- Three noise models, reported separately. A detection F1 quoted without saying which kind of noise produced it is close to meaningless.
The method is confident learning (Northcutt et al.), and the code is commented
with the reasoning rather than the mechanics, so labelaudit/detect.py is the
place to start reading.
pip install -r requirements.txt
# no download needed, whole pipeline in under a minute
python run_experiment.py --dataset synthetic --quick
# the real corpus
python run_experiment.py --dataset banking77
# then read the flagged examples yourself
python audit.py
python audit.py --scoreEach run writes results.md, results.json, a blind audit_queue.csv and five
charts into --out (default results/, which is gitignored).
audit.py does not know where the experiment wrote its output. Its defaults
point at results/, so if you pass --out something_else you must also pass
--queue, --key and --verdicts.
The committed results_banking77/ is the exact run every number in Findings
comes from, kept in the repo so the claims can be checked without running
anything.
CPU only. No GPU, no API keys. The only downloads are Banking77 and, if you pass
--upgrade minilm, the MiniLM weights you probably already have locally.
python -m pytest tests/ -q # 51 tests, about 30 secondscorpus.py synthetic corpus with per-class difficulty, and Banking77
noise.py symmetric, asymmetric and instance-dependent corruption
models.py TF-IDF and MiniLM pipelines, K-fold out-of-sample probabilities
detect.py confident learning, plus four baselines
evaluate.py detection precision, recall and F1 against known corruption
experiments.py relabelling budget curves, labels-versus-models
audit.py the hand audit, blind, three verdicts
Everything in Findings below is mode 1. If you point this at your own data
you are in mode 2, and no noise is injected anywhere.
| mode 1: validate the method | mode 2: audit your data | |
|---|---|---|
| what runs | run_experiment.py |
run_experiment.py, then audit.py |
| labels corrupted | yes, 15% of the training labels | no, never |
| is there an answer key | yes, because we made the errors | no, your errors are real and unknown |
| what you get | precision, recall and F1 per detector | a ranked list, a prize estimate, a confusion map |
| the question | does this method work at all | is my data worth fixing, and which rows |
Mode 1 exists because you cannot score a detector without knowing which labels are wrong. Banking77's own mistakes have never been counted, so there would be nothing to measure against. Corrupting 15% of the labels yourself creates an answer key, on real text.
Noise is injected into the training labels only. The test labels stay clean, so accuracy is measured against an honest ruler. That is only possible because we did the corrupting, which is the entire reason the injected path exists.
In mode 2 none of that applies. Your data already contains real errors. The detector runs on your labels as they are, and the output is a queue for a human, not a corrected dataset. The tool narrows the rows worth reading; it does not overwrite anybody's label.
Out-of-sample predictions are required, not tidier. A model asked about data
it trained on reports what it memorised, so a mislabelled example it fitted
perfectly comes back looking confidently correct. Detection then finds almost
nothing, and it fails silently: you get a short flag list and conclude the
dataset is clean. in_sample_probabilities exists only so a test can measure the
cost of skipping cross-validation instead of asserting it.
A fold can be missing a class. predict_proba then returns fewer columns
than there are classes, and assuming column j means class j misaligns every
probability from that point on. That corrupts the detector rather than crashing
it, so _expand_to_full_classes handles it and a test pins it.
Loss ranking and self-confidence ranking are the same ordering, because
-log is monotonic. They look like two independent methods and are one method
with two names. orderings_identical asserts it, and the report says so.
Measured on Banking77, 13,083 real customer service messages across 77
intents, with 15% of the training labels corrupted so that the answer key
exists. Full output in results_banking77/. The synthetic generator is kept as
a contrast, because on two of these questions it gave the wrong answer.
symmetric asymmetric instance-dependent
self-confidence 1.000 1.000 0.290
loss 1.000 1.000 0.290
margin 0.970 0.980 0.890
entropy 0.170 0.160 0.240
confident-learning 1.000 1.000 0.760
precision@100. Under symmetric and asymmetric noise, three detectors tie at a perfect 1.000. That is a ceiling, not a result: the experiment measures nothing and every ranking it produces is arbitrary.
Under instance-dependent noise, where corruption lands on genuinely
confusable examples the way real annotation errors do, the ordering reverses.
Self-confidence, the most common detector, falls from 1.000 to 0.290, and
margin, third on the easy models, wins.
The mechanism is specific. self-confidence is 1 - P[given label], which asks
whether the model is unsure. On confusable pairs it is unsure about everything,
so it cannot separate hard-but-correct from wrong. margin is
P[best] - P[given], which asks whether the model prefers one specific
alternative, and that is what a real mislabelling looks like.
Most published evaluations use symmetric noise. On this corpus symmetric noise ranks self-confidence first and margin third, and the realistic noise model reverses both.
strategy flagged precision recall F1
per-class (confident learning) 993 0.744 0.538 0.624
global 0.3 663 0.830 0.400 0.540
global 0.5 231 0.957 0.161 0.275
global 0.7 58 0.966 0.041 0.078
global 0.9 2 1.000 0.001 0.003
On the synthetic corpus a global cutoff of 0.3 beat per-class thresholds on F1, 0.750 to 0.657, and this README used to report that. On Banking77 per-class wins outright, 0.624 to 0.540, and the robustness argument survives as well: the global cutoff still collapses from 0.540 to 0.003 across plausible values, and picking 0.3 requires the answer key you do not have.
The correction is not that the old claim was too modest. It is that the generator ranked two methods in the wrong order, and only real text exposed it.
noisy baseline 0.8512
clean ceiling 0.8652
total prize +0.0140 fixing every one of the 1,374 corrupted labels
model upgrade +0.0148 word n-grams plus character n-grams, labels untouched
Swapping the model beat a perfect hand-cleaning of the entire corpus. No labels changed, no annotator time, and a larger gain.
That is the intended use of this tool. It is a seven-minute check on whether relabelling is worth a week, and on this corpus the answer was no. Banking77 is a curated benchmark, so a small prize is what a clean dataset should look like; a production support inbox is where the number would be different.
1.1x to 1.6x random across the three noise models, against 5.1x to 6.7x for the confidence-based detectors. It ranks hard examples, which is a different set from wrong ones, and it is the natural thing to reach for.
The synthetic corpus originally defaulted to a difficulty of 0.5, which gave nearly disjoint class vocabularies, 99.8% accuracy, and a model that shrugged off 25% label noise entirely. The clean ceiling equalled the noisy baseline, so there was no prize and every experiment measured zero. The default is now 0.8, which lands near 0.93 clean with roughly 5 points lost to that noise.
Finding 1 is the same lesson at the next level up: a task can be hard enough to produce a number and still be too easy to rank the methods.
An exchange rate needs a guard. When the upgraded model scores below the baseline, the naive computation reports "0 labels matched the upgrade", because the untouched baseline already beats the regressed upgrade. That is arithmetically true and meaningless, so the result now comes back undefined with a stated reason.
- The upgrade arm is only meaningful on real text. On the synthetic corpus
the vocabulary is generated tokens, so character n-grams have nothing real to
find and the upgrade regresses. Use
--dataset banking77. - Relabelling is assumed perfect, meaning a human asked to recheck returns the correct answer. Real annotators do not, so every relabelling number is an upper bound.
- Banking77's own test labels are noisy too, so accuracy there is measured with an imperfect ruler. The injected-noise path keeps a clean test set by construction; on the real corpus, hand-clean a small test slice before quoting anything.
- One corpus. Every real number here is Banking77. The findings may not transfer to a corpus with a different class structure or a different error rate.
- The model is TF-IDF plus logistic regression, not a transformer. Better probabilities would probably mean better detection, which is a confound the project does not control for.
- The gain from relabelling is not steady. Accuracy wobbles a few tenths of a point as labels are fixed, because with 77 classes a few hundred changed rows move the decision boundary. Partway through an audit you cannot tell whether it is working.
- Per-class threshold values are not saved to
results.json, only the aggregate metrics, so their spread across classes cannot be inspected after a run.