Skip to content

fix: make predictive_maintenance hold up on the real AI4I CSV - #182

Merged
danmcleran merged 1 commit into
masterfrom
fix/predictive-maintenance-real-data
Sep 23, 2026
Merged

danmcleran merged 1 commit into
masterfrom
fix/predictive-maintenance-real-data

Conversation

@danmcleran

Copy link
Copy Markdown
Owner

The example reported F1 0.84, but only on the synthetic fallback. On the real ai4i2020.csv it scored precision 0.29 and F1 0.43 (10-seed mean). The synthetic generator produced 16% failures against the real 3.4%, so the imbalance the docs describe was never exercised.

Changes

  • Training target: learn the three modes the sensor readings determine (HDF, PWF, OSF); score against the full Machine failure label. TWF and RNF are random by construction.
  • Sampling: 20% positives (was 50%), 300k iterations (was 80k).
  • Synthetic generator fitted to the real CSV: rpm/torque correlation, variant mix, tool-replacement window. Base rate now ~3.75%.
  • README, docs page and plot report only measured figures, labelled by data source.

Results (seeds 1-10)

Data Precision Recall F1
Real CSV, before 0.29 0.91 0.43
Real CSV, after 0.66 0.80 0.72
Synthetic, after 0.72 0.80 0.75

A full-batch Adam float MLP of the same shape reaches ~0.77 on the real CSV, so this is close to the architecture's ceiling. Recall is capped near 0.85: 52 of 339 real failures are TWF/RNF only.

Test plan

  • make release with g++ and clang++, -Werror clean
  • Synthetic run (seed 7): P 0.742 R 0.880 F1 0.805
  • Real CSV run (seed 7): P 0.606 R 0.717 F1 0.657
  • 10-seed sweeps on both data sources

🤖 Generated with Claude Code

https://claude.ai/code/session_015jDhsfLHi6epg599E5dD1h

The example reported F1 0.84, but that figure came from the synthetic
fallback only. On the real ai4i2020.csv it scored precision 0.29 and
F1 0.43 (10-seed mean): the synthetic generator drew rpm and torque
independently and let tool wear run to 253 min, inflating failures to
16% against the real 3.4%, so the real imbalance was never exercised.

- Train on the three modes the readings determine (HDF, PWF, OSF) and
  score against the full Machine failure label. TWF (random tool
  replacement) and RNF (0.1% random) cannot be predicted, and as
  positives they taught the net to alarm on high tool wear.
- Draw 20% of training samples from the failure pool instead of 50%,
  and train for 300k iterations instead of 80k. At 50/50 the same
  target produces about 1.55x the false alarms.
- Fit the synthetic generator to the real CSV: log-log rpm/torque
  relation (corr -0.88), 60/30/10 variant mix, tool replaced between
  200 and 240 min. It now yields ~3.75% failures.

Real CSV, seeds 1-10: precision 0.66, recall 0.80, F1 0.72 (was 0.29,
0.91, 0.43). Synthetic: F1 0.75. README and docs now report only
measured figures, labelled by data source; plot regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015jDhsfLHi6epg599E5dD1h
@danmcleran
danmcleran merged commit 77915d5 into master Sep 23, 2026
24 checks passed
@danmcleran
danmcleran deleted the fix/predictive-maintenance-real-data branch September 23, 2026 12:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant