Skip to content

Repository files navigation

Label error detection

Your model is graded against an answer key that contains mistakes, and nobody has counted them.

Labels come from people, and people make mistakes at a rate almost no team measures. This project measures that rate, tests whether a detector can find the bad labels, and answers the question a working team actually has to decide: is fixing the data worth more than improving the model?

Confident learning is a published method with a maintained library, and running it is easy. Three things here are not, and they are where the project earns its place:

  1. A hand audit with three verdicts. You read the top flagged examples yourself and record wrong, correct, or ambiguous. Precision is reported as a range, and a blind random control gives it something to be compared against.
  2. Labels against models, as an exchange rate. Not "which is better", but how many labels you must fix to match what a model upgrade bought.
  3. Three noise models, reported separately. A detection F1 quoted without saying which kind of noise produced it is close to meaningless.

The method is confident learning (Northcutt et al.), and the code is commented with the reasoning rather than the mechanics, so labelaudit/detect.py is the place to start reading.


Install and run

pip install -r requirements.txt

# no download needed, whole pipeline in under a minute
python run_experiment.py --dataset synthetic --quick

# the real corpus
python run_experiment.py --dataset banking77

# then read the flagged examples yourself
python audit.py
python audit.py --score

Each run writes results.md, results.json, a blind audit_queue.csv and five charts into --out (default results/, which is gitignored).

audit.py does not know where the experiment wrote its output. Its defaults point at results/, so if you pass --out something_else you must also pass --queue, --key and --verdicts.

The committed results_banking77/ is the exact run every number in Findings comes from, kept in the repo so the claims can be checked without running anything.

CPU only. No GPU, no API keys. The only downloads are Banking77 and, if you pass --upgrade minilm, the MiniLM weights you probably already have locally.

python -m pytest tests/ -q      # 51 tests, about 30 seconds

What the pipeline does

corpus.py        synthetic corpus with per-class difficulty, and Banking77
noise.py         symmetric, asymmetric and instance-dependent corruption
models.py        TF-IDF and MiniLM pipelines, K-fold out-of-sample probabilities
detect.py        confident learning, plus four baselines
evaluate.py      detection precision, recall and F1 against known corruption
experiments.py   relabelling budget curves, labels-versus-models
audit.py         the hand audit, blind, three verdicts

Two modes, and it matters which one you are reading about

Everything in Findings below is mode 1. If you point this at your own data you are in mode 2, and no noise is injected anywhere.

mode 1: validate the method mode 2: audit your data
what runs run_experiment.py run_experiment.py, then audit.py
labels corrupted yes, 15% of the training labels no, never
is there an answer key yes, because we made the errors no, your errors are real and unknown
what you get precision, recall and F1 per detector a ranked list, a prize estimate, a confusion map
the question does this method work at all is my data worth fixing, and which rows

Mode 1 exists because you cannot score a detector without knowing which labels are wrong. Banking77's own mistakes have never been counted, so there would be nothing to measure against. Corrupting 15% of the labels yourself creates an answer key, on real text.

Noise is injected into the training labels only. The test labels stay clean, so accuracy is measured against an honest ruler. That is only possible because we did the corrupting, which is the entire reason the injected path exists.

In mode 2 none of that applies. Your data already contains real errors. The detector runs on your labels as they are, and the output is a queue for a human, not a corrected dataset. The tool narrows the rows worth reading; it does not overwrite anybody's label.


Three things that are easy to get wrong

Out-of-sample predictions are required, not tidier. A model asked about data it trained on reports what it memorised, so a mislabelled example it fitted perfectly comes back looking confidently correct. Detection then finds almost nothing, and it fails silently: you get a short flag list and conclude the dataset is clean. in_sample_probabilities exists only so a test can measure the cost of skipping cross-validation instead of asserting it.

A fold can be missing a class. predict_proba then returns fewer columns than there are classes, and assuming column j means class j misaligns every probability from that point on. That corrupts the detector rather than crashing it, so _expand_to_full_classes handles it and a test pins it.

Loss ranking and self-confidence ranking are the same ordering, because -log is monotonic. They look like two independent methods and are one method with two names. orderings_identical asserts it, and the report says so.


Findings

Measured on Banking77, 13,083 real customer service messages across 77 intents, with 15% of the training labels corrupted so that the answer key exists. Full output in results_banking77/. The synthetic generator is kept as a contrast, because on two of these questions it gave the wrong answer.

1. The standard noise models cannot tell the detectors apart

                       symmetric   asymmetric   instance-dependent
self-confidence          1.000       1.000          0.290
loss                     1.000       1.000          0.290
margin                   0.970       0.980          0.890
entropy                  0.170       0.160          0.240
confident-learning       1.000       1.000          0.760

precision@100. Under symmetric and asymmetric noise, three detectors tie at a perfect 1.000. That is a ceiling, not a result: the experiment measures nothing and every ranking it produces is arbitrary.

Under instance-dependent noise, where corruption lands on genuinely confusable examples the way real annotation errors do, the ordering reverses. Self-confidence, the most common detector, falls from 1.000 to 0.290, and margin, third on the easy models, wins.

The mechanism is specific. self-confidence is 1 - P[given label], which asks whether the model is unsure. On confusable pairs it is unsure about everything, so it cannot separate hard-but-correct from wrong. margin is P[best] - P[given], which asks whether the model prefers one specific alternative, and that is what a real mislabelling looks like.

Most published evaluations use symmetric noise. On this corpus symmetric noise ranks self-confidence first and margin third, and the realistic noise model reverses both.

2. Per-class thresholds win on real data, and the generator said otherwise

strategy                       flagged  precision  recall     F1
per-class (confident learning)     993      0.744   0.538   0.624
global 0.3                         663      0.830   0.400   0.540
global 0.5                         231      0.957   0.161   0.275
global 0.7                          58      0.966   0.041   0.078
global 0.9                           2      1.000   0.001   0.003

On the synthetic corpus a global cutoff of 0.3 beat per-class thresholds on F1, 0.750 to 0.657, and this README used to report that. On Banking77 per-class wins outright, 0.624 to 0.540, and the robustness argument survives as well: the global cutoff still collapses from 0.540 to 0.003 across plausible values, and picking 0.3 requires the answer key you do not have.

The correction is not that the old claim was too modest. It is that the generator ranked two methods in the wrong order, and only real text exposed it.

3. The tool's answer here was "do not bother", which is the point

noisy baseline    0.8512
clean ceiling     0.8652
total prize      +0.0140        fixing every one of the 1,374 corrupted labels

model upgrade    +0.0148        word n-grams plus character n-grams, labels untouched

Swapping the model beat a perfect hand-cleaning of the entire corpus. No labels changed, no annotator time, and a larger gain.

That is the intended use of this tool. It is a seven-minute check on whether relabelling is worth a week, and on this corpus the answer was no. Banking77 is a curated benchmark, so a small prize is what a clean dataset should look like; a production support inbox is where the number would be different.

4. Entropy loses, as designed

1.1x to 1.6x random across the three noise models, against 5.1x to 6.7x for the confidence-based detectors. It ranks hard examples, which is a different set from wrong ones, and it is the natural thing to reach for.

5. A too-easy task measures nothing

The synthetic corpus originally defaulted to a difficulty of 0.5, which gave nearly disjoint class vocabularies, 99.8% accuracy, and a model that shrugged off 25% label noise entirely. The clean ceiling equalled the noisy baseline, so there was no prize and every experiment measured zero. The default is now 0.8, which lands near 0.93 clean with roughly 5 points lost to that noise.

Finding 1 is the same lesson at the next level up: a task can be hard enough to produce a number and still be too easy to rank the methods.

An exchange rate needs a guard. When the upgraded model scores below the baseline, the naive computation reports "0 labels matched the upgrade", because the untouched baseline already beats the regressed upgrade. That is arithmetically true and meaningless, so the result now comes back undefined with a stated reason.


Honest limitations

  • The upgrade arm is only meaningful on real text. On the synthetic corpus the vocabulary is generated tokens, so character n-grams have nothing real to find and the upgrade regresses. Use --dataset banking77.
  • Relabelling is assumed perfect, meaning a human asked to recheck returns the correct answer. Real annotators do not, so every relabelling number is an upper bound.
  • Banking77's own test labels are noisy too, so accuracy there is measured with an imperfect ruler. The injected-noise path keeps a clean test set by construction; on the real corpus, hand-clean a small test slice before quoting anything.
  • One corpus. Every real number here is Banking77. The findings may not transfer to a corpus with a different class structure or a different error rate.
  • The model is TF-IDF plus logistic regression, not a transformer. Better probabilities would probably mean better detection, which is a confound the project does not control for.
  • The gain from relabelling is not steady. Accuracy wobbles a few tenths of a point as labels are fixed, because with 77 classes a few hundred changed rows move the decision boundary. Partway through an audit you cannot tell whether it is working.
  • Per-class threshold values are not saved to results.json, only the aggregate metrics, so their spread across classes cannot be inspected after a run.

About

Finds mislabelled training data, and measures whether fixing it beats improving the model

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages