A TensorFlow LSTM scores sensor telemetry; a language model turns each prediction into a work order a technician can act on. Served over a REST API with a dashboard, containerised, and tested end to end.
Quick start · Documentation · API · Model card · Results · Roadmap
Scope, stated up front. The dataset is synthetic, with a degradation pattern deliberately designed to be detectable. The pipeline transfers to real telemetry; these numbers would not. There is no authentication, so this belongs on localhost or a trusted network. Full accounting in
docs/model.md#limitationsandSECURITY.md.
Machine 51 failed at 2024-10-31 12:00. This is what the model said about it at
every hour of the two days beforehand — each point scored only on the
evidence available at that hour, via the API's as_of parameter, so nothing
downstream of the moment leaks in.
It is flat at zero for thirty hours, first flickers at -17h (0.0005), and crosses the alert threshold at -16h — sixteen hours of warning on a machine that gave no earlier sign. Then it saturates and stays there.
The flat part on the left matters as much as the climb. The model was trained to see 24 hours ahead and no further, so a day before the failure it is silent and should be. That is the horizon, drawn.
Regenerate it yourself. The chart is committed, but so is the script that
draws it — an image nobody can reproduce is the same problem as a status badge
that cannot go red. data/ and models/ are gitignored, so a fresh clone
builds them first:
python scripts/generate_data.py # ~1 min
python scripts/run_preprocessing.py # ~3 min
python scripts/train_model.py # ~81 min on CPU, seeded, reproducible
make docker-up-d # or: make run-api, in another shell
pip install -r requirements-dev.txt # the script needs matplotlib
python scripts/plot_horizon.py # machine 51, the chart above
python scripts/plot_horizon.py --machine 96 --failure 2024-11-14T00:00:00Machine 96 crosses at -23h, machine 51 at -16h. Warning time varies by how early a machine's sensors start drifting. Across all 8 failure events in the held-out period the model catches 8, with a median of 23.5 hours of warning and a worst case of 16 — machine 51, the one charted above. The 24 hours in the headline is the ceiling the model was trained to, not a promise about any one machine.
Or drive it yourself: open the dashboard at localhost:8501, turn on Rewind
in the sidebar, and set the date to 2024-10-31 hour 6.
- See it work
- Overview
- Architecture
- Features
- Tech Stack
- Getting Started
- Usage
- Results
- Project Structure
- Documentation
- Testing
- Project Status
- Roadmap
- Screenshots
- FAQ
- Troubleshooting
- Contributing
- Citation
- Acknowledgements
- License
At a glance: an LSTM predicts equipment failure 24 hours ahead from four sensors, and an LLM turns each prediction into a work order a technician can act on. On held-out data it caught 8 of 8 failure events with a median 23.5 hours of warning, at 21 false alarms across 172,800 hourly readings. Predictions serve in 137 ms; the whole thing runs with no API key.
Full numbers, with their caveats, in docs/RESULTS.md.
Industrial equipment failures cost the global manufacturing industry $50 billion per year in unplanned downtime. This project implements a Predictive Maintenance system that:
- Predicts failures — Uses an LSTM (Long Short-Term Memory) neural network trained on sensor telemetry data to predict when equipment will fail.
- Explains failures — Uses LangChain + LLM to convert ML predictions into plain-English maintenance reports.
- Enables proactive maintenance — Provides a REST API and interactive dashboard for maintenance teams.
| Traditional Approach | This Project |
|---|---|
| Fix equipment after it breaks | Predict failures before they happen |
| ML model outputs a number (0.87) | GenAI explains: "Bearing temperature rising 3°C/hr, recommend immediate inspection" |
| Requires ML expertise to interpret | Maintenance managers can read plain-English reports |
| Isolated scripts | Production-ready API + Dashboard |
5 raw sensor tables (telemetry · machines · errors · maintenance · failures)
│
▼
DataIngestion → DataValidator → DataPreprocessor
│ merge · 63 engineered features · 24h labels
│ temporal split · scale on train only · 24-step windows
▼
LSTM (128 → 64 → 32 → 1) class-weighted, {0: 0.50, 1: 364.89}
│
▼
Predictor raw tables in → probability + risk band + evidence out
│ reuses the SAME feature code as training (verified: 100%
│ alert agreement over 172,800 sequences)
▼
ReportGenerator / MaintenanceAssistant LangChain · OpenAI|Gemini|Ollama
│ every figure it cites is supplied; nothing is invented
▼
FastAPI ── 137 ms predictions ─────────────────────┐
│ /report isolated so a 21 s LLM call │
│ can never delay or break a prediction │
▼ ▼
Streamlit dashboard ── pure HTTP client, holds no model ──
Each layer depends only on the ones above it. src/data/ imports no
TensorFlow; src/models/ imports no pandas; the dashboard imports neither.
See docs/architecture.md for the detailed architecture diagram.
- 🤖 LSTM Predictive Model — TensorFlow-based time-series model for failure prediction
- 🧠 GenAI Insights — LangChain-powered natural language maintenance reports
- 🔌 REST API — FastAPI with auto-generated OpenAPI documentation
- 📊 Interactive Dashboard — Real-time equipment health monitoring (Streamlit)
- ⏪ Point-in-time assessment — rewind the fleet to any hour and score it on the evidence available then, with everything after it hidden
- 🐳 Dockerized — One-command deployment with Docker
- ✅ Tested — Unit + integration tests with pytest
- 📝 Well-Documented — Comprehensive code documentation and architecture docs
| Component | Technology |
|---|---|
| ML Framework | TensorFlow / Keras |
| GenAI | LangChain + OpenAI / Ollama |
| API | FastAPI + Uvicorn |
| Dashboard | Streamlit |
| Data Processing | Pandas, NumPy, Scikit-learn |
| Configuration | Pydantic Settings |
| Logging | Loguru |
| Testing | Pytest |
| Code Quality | Black, Flake8, MyPy |
| Containerization | Docker |
| CI/CD | GitHub Actions |
- Python 3.12 — required. TensorFlow publishes no wheels for 3.13+, so a newer
system Python (e.g. 3.14) will fail at install time. On macOS:
brew install python@3.12 - Git
- ~6 GB free disk space (the generated dataset and its processed tensors)
- (Optional) An OpenAI / Google API key for the GenAI features — a local Ollama model
works with no key at all. The provider packages are optional extras, imported
lazily, so install the one you want:
pip install -e ".[ollama]"(keyless) orpip install -e ".[google]" - (Optional) Docker for containerized deployment
# 1. Clone
git clone https://github.com/Vanshcloud/Predictive-Maintenance-GenAI.git
cd Predictive-Maintenance-GenAI
# 2a. Automated setup (finds a compatible Python for you)
chmod +x scripts/setup.sh
./scripts/setup.sh
# 2b. …or manually. Note python3.12 explicitly — plain `python3` may be
# a version TensorFlow does not support.
python3.12 -m venv venv # macOS: /opt/homebrew/bin/python3.12
source venv/bin/activate
pip install --upgrade pip
pip install -r requirements-dev.txt
cp .env.example .env # then add an LLM API key if you want reportsThe raw and processed datasets are not in this repository — together they are about 6 GB. They are fully reproducible from a fixed seed:
source venv/bin/activate
python scripts/generate_data.py # -> data/raw/ (883,231 rows, seed=42)
python scripts/run_preprocessing.py # -> data/processed/*.npy (~5.2 GB)data/sample/ is committed and is what the test suite runs against, so make test
works immediately after install — but note it deliberately contains zero failure
events, so it cannot be used to train or evaluate a model.
# Activate virtual environment
source venv/bin/activate
# Run smoke tests
make test
# See all available commands
make helpsource venv/bin/activate
# Data
python scripts/generate_data.py # full dataset (100 machines x 365 days)
python scripts/generate_data.py --sample # small fixture -> data/sample/
python scripts/eda_analysis.py # 8-dimension exploratory report
python scripts/run_preprocessing.py # raw CSVs -> LSTM tensors + scaler
# Model
python scripts/train_model.py # train, evaluate, write metrics
python scripts/train_model.py --epochs 50 --monitor val_f1 --resume
python scripts/evaluate_model.py # threshold sweep + curves
# Predict
python scripts/predict.py # current fleet status, most urgent first
python scripts/predict.py --machine 47 # one machine, as JSON
python scripts/predict.py --alerts-only -o alerts.csv
# AI maintenance reports
python scripts/generate_report.py --machine 51 --dry-run # prompt only — no API key needed
python scripts/generate_report.py --machine 51 # full report
python scripts/generate_report.py --fleet
# Ask follow-up questions about one machine
python scripts/ask.py --machine 51
python scripts/ask.py --machine 51 --ask "Has vibration been rising?"
# REST API
make run-api # http://localhost:8000/docs
curl localhost:8000/health
curl localhost:8000/machines/51/predict
curl localhost:8000/fleet?alerts_only=true
# Dashboard (needs the API running)
make run-dashboard # http://localhost:8501
# Containers (API + dashboard together)
make docker-build
make docker-up # api :8000 · dashboard :8501
# Quality
make test # 246 unit tests
make quality # lint + format-check + typecheckTrained on a three-way chronological split — 567,000 train / 129,000 validation /
172,800 test sequences of shape (24, 63). Early stopping at epoch 28 of 30, best
weights from epoch 23. Wall clock ~81 minutes, CPU-only (Apple Silicon).
Training is seeded (--seed 42), so these numbers are reproducible rather
than merely reported: python scripts/train_model.py re-derives them.
The alert threshold (0.3415) was chosen by sweeping the precision-recall curve on the validation split; the test set is scored once, at that threshold.
| Metric | Value |
|---|---|
| ROC-AUC | 0.9999 |
| Precision | 0.8976 |
| Recall | 0.9200 |
| F1 | 0.9086 |
| Single-sequence inference | 54 ms median |
Confusion matrix over 172,800 held-out sequences:
predicted 0 predicted 1
actual 0 172,579 21 <- false alarms
actual 1 16 184 <- caught 184 of 200 hourly labels
Every hour in the 24 hours before a failure carries a positive label, so the hourly figures above count hours, not failures. Measured over failure events — where catching any hour means the technician was warned:
| Metric | Value |
|---|---|
| Failure events in the test period | 8 |
| Events warned about | 8 (100%) |
| Lead time (median / min / max) | 23.5h / 16h / 24h |
Eight events is a small sample: this says the model warned in 8 of 8 cases, not that it never misses. Reported alongside precision, never instead of it — event recall says nothing about alert fatigue.
Accuracy is deliberately not reported. With a 1:864 positive rate a model that predicts "no failure" every time scores 99.88%, so accuracy would be actively misleading. Quality is judged on AUC, precision, recall, and F1 only.
Reports are grounded in the engineered features the model actually consumed — never in the model's imagination. Each sensor line carries an explicit verdict, and causal explanations are attached only when a reading deviates in the direction that matters:
pressure: 65.89 PSI (24h baseline 93.47 PSI, 1.91 sigma below; -32.06 PSI over 24h)
-> ABNORMAL; typically indicates a leak or a failing seal
rotation: 400.04 RPM (24h baseline 418.31 RPM, 0.33 sigma below)
-> within normal variation; no action indicated by this sensor
Run it with no API key at all: --dry-run prints the grounded facts, and a
local Ollama model generates the full report keyless.
One caveat, stated plainly. The dataset is synthetic, with a degradation pattern deliberately designed to be detectable. These metrics reflect this dataset's difficulty, not that of real industrial equipment. The pipeline transfers; the numbers would not. Full accounting in
docs/devlog/day-05.md.
Predictive-Maintenance-GenAI/
├── README.md # You are here
├── CONTRIBUTING.md # Dev setup, quality gates, and the invariants that must hold
├── SECURITY.md # Threat model and vulnerability reporting
├── CODE_OF_CONDUCT.md # Contributor Covenant 2.1
├── CHANGELOG.md # Keep a Changelog format
├── LICENSE # MIT
├── Makefile # make test / lint / format / typecheck / quality
├── pyproject.toml # PEP 621 metadata + Black/isort/pytest/mypy/coverage config
│
├── .github/
│ ├── workflows/ci.yml # Lint, types, tests, image builds, dependency audit
│ ├── ISSUE_TEMPLATE/ # Bug report + feature request forms
│ └── dependabot.yml # Grouped monthly dependency updates
│
├── config/ # Centralized configuration (pydantic-settings)
├── src/
│ ├── utils/ # Logging, exception hierarchy
│ ├── data/ # Ingestion, validation, preprocessing + feature engineering
│ ├── models/ # LSTM architecture, training loop, evaluator
│ ├── prediction/ # Predictor: raw tables -> ranked predictions
│ ├── genai/ # LangChain prompts + report chains
│ └── api/ # FastAPI REST API — 9 endpoints, /docs
├── dashboard/ # Streamlit UI — pure HTTP client of the API
├── docker/ # Dockerfiles + compose (API and dashboard)
│
├── scripts/ # Entry points: generate_data, eda, preprocessing,
│ # train_model, evaluate_model, predict
├── tests/
│ ├── conftest.py # Session bootstrap (import order matters — see the file)
│ ├── unit/ # 246 tests, ~27s
│ └── integration/ # 13 tests — parity, grounding, time travel; `make test-integration`
├── docs/
│ ├── README.md # Documentation index
│ ├── IMPLEMENTATION_PLAN.md # Full engineering spec: scope, risks, milestones
│ ├── architecture.md # Layers, module responsibilities, invariants
│ ├── RESULTS.md # Every metric, with its caveats
│ └── devlog/ # Build journal — one entry per milestone
│
├── data/
│ ├── raw/ # gitignored — regenerate with generate_data.py
│ ├── processed/ # gitignored — regenerate with run_preprocessing.py
│ └── sample/ # committed test fixture (no failure events — by design)
├── models/ # .keras gitignored; metrics.json + history committed
├── notebooks/ # Jupyter scratch space
└── logs/ # gitignored, rotated daily
The src/ packages form a strict dependency chain — each may only import from those to
its left, which is what keeps the layers independently testable:
config/ -> src/utils/ -> src/data/ -> src/models/ -> src/prediction/ -> src/genai/ -> src/api/ -> dashboard/
| Document | What it is |
|---|---|
docs/ |
Documentation index — start here |
docs/architecture.md |
Layer diagram, module responsibilities, and the correctness invariants |
docs/RESULTS.md |
Every metric in one place, each with the caveat it needs |
docs/IMPLEMENTATION_PLAN.md |
Full engineering specification: scope, requirements, dataset, model, deployment plan, risk register, milestones |
docs/devlog/ |
Build journal — one entry per milestone, including what went wrong and why |
CONTRIBUTING.md |
Development setup, quality gates, code style, and the correctness invariants a change must not break |
SECURITY.md |
Threat model, what is and is not hardened, and how to report a vulnerability |
CHANGELOG.md |
Release history |
make test # 246 unit tests
make test-cov # with coverage report
make lint # flake8
make format # black + isort (writes)
make quality # lint + format-check + typecheck (no writes)Current status: 246 unit + 13 integration tests passing, 0 flake8 issues, mypy clean, Black and isort clean.
Tests run against the committed data/sample/ fixture, so they need no generated data.
The first run pays roughly 90 seconds for TensorFlow's initial import on ARM64 macOS;
subsequent runs finish in about 4 seconds.
tests/conftest.pyimports TensorFlow before anything else, and that is load-bearing — not a stray import. TensorFlow and Apache Arrow (pulled in by pandas/scikit-learn) each bundle their own copy of abseil, and whichever loads first claims the shared symbol for the whole process. Get the order wrong and the suite deadlocks at 0% CPU with no traceback. The file explains it in full.
Status: complete. The project has reached its planned scope, and every component below is built, tested, and documented.
It was built over 12 days, followed by three post-project milestones. The build
journal — one entry per milestone, including the bugs and the dead ends — is in
docs/devlog/.
| Data pipeline | Synthetic generator (883K rows, five related tables) → 63 engineered features → chronological three-way split → (N, 24, 63) sequences |
| Model | 149,825-parameter LSTM, seeded and reproducible — F1 0.9086 at t=0.3415, 8 of 8 failure events caught |
| Prediction | Training/serving parity verified; point-in-time assessment rewinds the fleet to any hour |
| GenAI | LangChain reports grounded in real sensor evidence; multi-turn Q&A that declines what the data cannot answer; runs keyless on local Ollama |
| API | FastAPI — 9 endpoints, 137 ms predictions, degrades gracefully when no model or LLM is available |
| Dashboard | Streamlit — a pure HTTP client of the API, holding no model of its own |
| Deployment | Two Docker images (2.87 GB API, 803 MB dashboard), compose-verified |
| CI | GitHub Actions — lint, format, types, 246 unit + 13 integration tests, CodeQL, link checking, release automation |
| Documentation | 39 documents, every metric stated with its caveats |
Known limitations, stated rather than discovered. The full list — with status
markers and what is explicitly not planned — is in
docs/roadmap.md.
The short version:
| 🔴 Before an untrusted network | No authentication · no rate limiting · unbounded request bodies |
| 🔴 Highest value overall | Validate against a real dataset — the numbers here describe a synthetic one |
| 🟡 Known inefficiencies | scripts/predict.py --machine N scores the whole fleet (~1000× slower than it needs to be) · explain_machine() engineers features twice |
| 🟡 Model | Not calibrated · no drift detection · never benchmarked against gradient boosting |
Every image below is real output — no mockups. All are rewound to 2024-10-31 06:00, six hours before machine 51's recorded failure, and all are regenerated by committed scripts rather than trusted.
Fleet overview — 100 machines scored, 2 alerting, sorted most urgent first.
Machine detail — the prediction next to the sensor evidence behind it.
AI report — written by a local model from that evidence and nothing else; every figure it quotes traces back to the table above.
Regenerate them against your own stack:
make docker-up-d # http://localhost:8501
playwright install chromium # once
python scripts/capture_screenshots.pyA demo GIF is still missing — docs/images/README.md
records what it should show, and is a good first contribution.
Is this usable on real equipment?
The pipeline is. The numbers are not transferable — the dataset is synthetic with a degradation pattern designed to be detectable. Expect substantially worse performance on real telemetry, which has sensor dropout, calibration drift, mislabelled failures, and failure modes with no sensor signature at all. Validating against a real dataset is the top item on the roadmap.
Do I need an OpenAI API key?
No. Prediction endpoints never call a language model. Reports work keyless
through a local Ollama model (pip install -e ".[ollama]"), and with no
provider at all /report returns a 502 that still carries the prediction.
Why 0.3415 and not 0.5?
0.5 is inherited from balanced problems. At a 1:864 positive rate it is
almost never the right operating point. This threshold was chosen by sweeping
the precision–recall curve on the validation split; the test set was scored
once, at that point. See docs/model.md.
Why is accuracy never reported?
Because it would be misleading. At a 1:864 positive rate, predicting "no
failure" every time scores 99.88%. ModelEvaluator does not compute it at
all — quality is judged on AUC, precision, recall, and F1.
Why an LSTM rather than XGBoost?
The signal is degradation over time, and flattening a window discards ordering unless you hand-engineer it back. Honest caveat: gradient boosting was not benchmarked head to head, so "better here" is not a claim this project has earned. It is on the roadmap.
Why is training a hand-written loop instead of model.fit()?
History, and it is documented. During Day 4 fit() appeared to hang; the real
cause turned out to be an abseil symbol collision between TensorFlow and Apache
Arrow, not Keras. The manual loop is kept because it works, is tested, and makes
class weighting and early stopping explicit — but fit() has not been
re-benchmarked since, and the docs say so rather than claiming it is broken.
See docs/devlog/day-04.md.
The dashboard shows "0 alerting". Is it broken?
Almost certainly not. Most hours are quiet, and the dataset's final hour has no
machine inside a pre-failure window. Turn on Rewind and set
2024-10-31 hour 6 — machine 51 goes critical, six hours before it failed.
Why does the first test run take 90 seconds?
TensorFlow's initial import on ARM64 macOS. Subsequent runs take ~26 s.
Can I use Python 3.13?
No — TensorFlow publishes no wheels for it. Use 3.10, 3.11, or 3.12. This will follow when upstream does.
Most reported problems are already covered, with the fix, in
docs/troubleshooting.md.
The three that come up most:
| Symptom | Cause |
|---|---|
Could not find a version that satisfies the requirement tensorflow |
Python 3.13+. Use 3.12 |
| Training hangs at 0% CPU with no traceback | Import order — TensorFlow must load before pandas. Full diagnosis |
503 from every endpoint |
No trained model. curl localhost:8000/health will confirm |
Contributions are welcome. CONTRIBUTING.md covers
development setup, the quality gates, and — most importantly — the correctness
invariants in this codebase that fail silently if broken.
make setup # venv + dependencies
make quality # lint + format-check + typecheck
make test # unit testsPlease read SECURITY.md before reporting anything with a
security dimension, and note the Code of Conduct.
SUPPORT.md sets out what to expect from a bug report or a
question — this is a one-person project, and it says so rather than leaving you
guessing.
If this project is useful in your work, citing it is appreciated. GitHub also
renders a "Cite this repository" panel from CITATION.cff.
@software{tomar_predictive_maintenance_genai_2026,
author = {Tomar, Vansh},
title = {Predictive Maintenance + GenAI Insight Generator},
year = {2026},
version = {1.0.0},
url = {https://github.com/Vanshcloud/Predictive-Maintenance-GenAI},
license = {MIT}
}This is a software citation. There is no associated paper, and none is implied.
- The dataset schema follows the shape of Microsoft's Azure Predictive Maintenance sample — five related tables of telemetry, errors, maintenance, machines, and failures. The data here is generated, not that dataset.
- Built on TensorFlow, LangChain, FastAPI, Streamlit, pandas, and scikit-learn.
- Keep a Changelog and Semantic Versioning for release conventions, and the Contributor Covenant for the code of conduct.
MIT — see LICENSE. Use it, fork it, learn from it.
Vansh Tomar — GitHub
Built to demonstrate an end-to-end ML system: not just a model, but the preprocessing that does not leak, the evaluation that does not flatter, the service that degrades instead of failing, and the documentation that says what it does not know.



