A reproducible ML monitoring workbench: detect changes in incoming data and measure their effect on a trained classifier.
Open the interactive dashboard → · Executed analysis notebook · Methodology · Model card · Interview guide
A model's test score is only one part of its lifecycle. ShiftWatch asks what happens when new data no longer resembles the data used for training. It combines a reusable Python drift library, a command-line workflow, a real classification experiment, and an interactive TypeScript report viewer.
The bundled experiment uses UCI Wine. Shifted scenarios are explicitly synthetic; this is an educational monitoring workbench, with no claim of production deployment or business impact.
- Trains and selects models: compares a dummy baseline, logistic regression, and a random forest with five-fold stratified cross-validation. Imputation and scaling are fitted inside each fold.
- Evaluates a separate test set: reports accuracy, macro F1, log loss, a confusion matrix, and bootstrap accuracy intervals.
- Detects numerical data drift: computes Population Stability Index, two-sample KS tests, Benjamini–Hochberg adjusted p-values, Wasserstein distance, and missingness changes.
- Works with your data: compares two numerical CSVs and exports a validated JSON report. An optional nonzero exit code supports CI data gates.
- Explains the result: the dashboard offers scenario switching, feature search, distribution inspection, exploratory thresholds, model comparisons, report import, and export.
- Reproduces the evidence: committed results, dataset fingerprints, split indices, fixed seeds, dependency locks, and automated tests make the experiment inspectable.
The dashboard reads files locally. It does not send uploaded reports to a server, run Python in the browser, or pretend to be a live production monitor.
Requires Python 3.12. The dataset ships with scikit-learn; no dataset download, cloud account, or API key is needed after installing dependencies.
git clone https://github.com/pralav-25/shiftwatch.git
cd shiftwatch
python3.12 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.lock
pip install --no-deps -e .
shiftwatch demoThis writes public/reports/demo.json. Open the dashboard and select Load report to inspect it. For a lightweight library installation without plotting and test tools, use pip install -e . instead of the locked environment.
The executed notebook walks through
training-only exploration, baseline and cross-validation results, holdout error
analysis, and synthetic drift scenarios. Its tables and figures are computed from
the same run_experiment function used by the dashboard. Read it directly on GitHub,
or rerun every cell from the repository root in the Python 3.12 environment above:
pip install -r requirements-notebook.lock
python -m nbconvert --to notebook --execute \
--ExecutePreprocessor.timeout=180 \
--output wine-monitoring-analysis --output-dir /tmp \
notebooks/wine-monitoring-analysis.ipynbThis saves a fresh executed copy under /tmp and preserves the committed notebook.
On Windows, replace /tmp with an existing temporary directory. Notebook tooling
has a separate dependency lock that includes the experiment's numerical lock.
GitHub Actions executes the notebook to catch stale imports and analysis errors.
Use two CSVs containing the same named numerical feature columns. Column order can differ. Remove identifiers and target labels, or select the feature columns explicitly. Each file needs at least five rows; missing numeric values are allowed, infinity and nonnumeric columns are rejected.
Give every column a unique, nonempty header. Empty header cells are rejected before
the CSV parser can replace them with invented Unnamed: ... feature names.
Literal nonempty names are preserved, including a name such as Unnamed: 0.
Rows with extra values are rejected instead of silently treating leading values as an index or dropping columns. Check for extra delimiters if this happens; a rejected comparison leaves any existing output report intact.
shiftwatch compare examples/reference.csv examples/shifted.csv --output report.json
# A data-quality gate: exit 2 if any feature alerts; the report is still saved.
shiftwatch compare reference.csv current.csv \
--psi-threshold 0.2 --alpha 0.05 --fail-on-alert --output report.jsonUse --columns temperature pressure to select features without editing the source
files. Names must be unique and present in both inputs; the report uses the given
order. Unselected identifier, label, and metadata columns are ignored, while CSV
header and row-width checks still apply to the whole file. Without this option,
all columns are compared and both inputs must have the same feature names.
Use --delimiter ';' to compare semicolon-separated exports. Both files use the
same separator; a literal tab or pipe also works. The separator must be one ASCII
character other than a newline, NUL, or double quote. Quoted field names and all
header/row-width checks use that separator. Comma remains the default.
Use --missing-value=-999 --missing-value MISSING for dataset-specific missing
markers. Repeat the option for each marker; they apply to data cells in both
files, while literal header names are preserved. Pandas' default blank/NA markers
remain enabled. Numeric sentinels remain observed numbers unless explicitly
declared. The resulting missing counts feed both missingness alerts and
--fail-on-insufficient-data. The input files are not modified.
Use --missing-threshold 0.1 to alert on a missing-rate change of at least ten
percentage points, or --bins 5 to choose the number of reference quantile bins
for PSI. Defaults are 0.05 and 10; accepted ranges are (0, 1] and integers
from 2 to 50. Both settings are saved in the report's config.
Report files are written atomically: a failed write preserves the previous report. The CLI refuses to overwrite either input CSV, including through a symlink or hard link.
Both compare and demo accept --output - to emit only report JSON on standard
output, suitable for a pipe or redirection. Validation errors go to standard error
and emit no report. Alert and insufficient-data exit flags keep the same meaning.
shiftwatch compare examples/reference.csv examples/shifted.csv --output - | python -m json.toolUse --output report.json for overwrite protection and atomic writes. Shell
redirection such as > report.json is handled by your shell and has neither safeguard.
Load the resulting JSON into the dashboard. Comparisons without a trained model show drift diagnostics without fabricated model-performance metrics. The dashboard accepts up to 200 features, 20 scenarios, and 5 MB per report; the Python library can analyze wider tables.
import pandas as pd
from shiftwatch import compare_frames
result = compare_frames(
pd.read_csv("examples/reference.csv"),
pd.read_csv("examples/shifted.csv"),
psi_threshold=0.2,
)
print(result["alert_count"])Library thresholds must be finite real scalars; booleans, strings, arrays, and
nonfinite values raise ValueError. NumPy scalar settings and integral bin counts
are normalized to ordinary Python numbers so the returned report remains JSON-serializable.
A successful comparison returns 0 unless --fail-on-alert is enabled and
the report contains an alert. With that flag, an alert returns 2 after the
report has been written, so downstream tooling can still inspect the result.
Without the flag, an alert does not make the command fail.
Use --fail-on-insufficient-data when the comparison must have at least five
observed values for every feature in both files. It also returns 2 after
saving diagnostics if any feature is untestable, even when neither distribution
nor missingness alerts fire. The two flags can be combined; neither is enabled
by default. This checks per-feature observations, in addition to the minimum
five rows required in each CSV.
Argument and input-validation errors also use exit status 2. Do not treat that status alone as proof of drift: check the command's error output and whether the current run successfully wrote a report. Use a separate output path for each run to avoid mistaking a preserved report from an earlier run for a fresh result.
Seed 42; 124 training rows; 54 held-out test rows. Training CV selects logistic regression, with mean macro F1 of 0.9834 across five folds. These scores describe a small, relatively easy dataset; they are not evidence of general-purpose model quality.
| Scenario | Accuracy | Features flagged | Change applied |
|---|---|---|---|
| Untouched holdout | 98.1% | 0 / 13 | None |
| Moderate shift | 83.3% | 4 / 13 | +0.75 training standard deviations in four features |
| Severe shift | 53.7% | 4 / 13 | +2 training standard deviations in four features |
| Missing data | 98.1% | 3 / 13 | Approximately 30% missing in three features |
All perturbed scenarios reuse the same test rows and original labels. They are stress tests, not new independent real-world samples. Missingness triggers alerts even when this model's accuracy happens to remain unchanged—one reason to distinguish data quality from prediction quality.
See the full JSON for fold scores, confidence intervals, confusion matrices, feature statistics, exact split indices, and runtime versions.
Requires Node.js 22.18+ and pnpm 11.19.0.
corepack enable
pnpm install --frozen-lockfile
pnpm devOpen the localhost URL printed by the development server. The dashboard uses React, TypeScript, Vinext, accessible Base UI/Shadcn primitives, and SVG charts. pnpm build generates a static dashboard in dist/client.
The Python pipeline and UI communicate through the versioned shiftwatch/v1 JSON contract. This keeps the numerical analysis portable and the public demo inexpensive to host.
flowchart LR
A[UCI Wine or numerical CSVs] --> B[Python validation]
B --> C[Training-only CV and model selection]
B --> D[Reference vs current drift analysis]
C --> E[Held-out evaluation and stress tests]
D --> F[Versioned JSON report]
E --> F
F --> G[React report viewer]
F --> H[CI alert gate]
pytest -q
ruff check src tests
ruff format --check src tests
pnpm test
pnpm typecheck
pnpm lint
pnpm build
# Reproduce and compare the checked-in numerical results:
shiftwatch demo --output /tmp/reproduced.json
python scripts/check_reproduction.py public/reports/demo.json /tmp/reproduced.json
# Regenerate the README figure:
python scripts/plot_results.pyThe suite covers constant distributions, overflow bins, known multiple-testing corrections, missing/all-missing columns, invalid schemas, CLI exit behavior, report validation, filter logic, disjoint train/test splits, deterministic results, and a spy that verifies no preprocessor fits on held-out rows.
GitHub Actions runs Python tests, report reproduction, frontend tests, type checking, linting, and a production build. A separate workflow publishes the dashboard to GitHub Pages after checks pass. Browser interaction and visual regression tests are not included; optional WebMCP support is feature-detected and requires a compatible browser for runtime validation.
src/shiftwatch/ Python drift engine, experiment, and CLI
app/ Interactive report dashboard
lib/report.ts Report contract, validation, and alert/filter logic
tests/ Numerical, leakage, reproducibility, and CLI tests
tests-web/ Report-validation and dashboard-logic tests
examples/ Reference data and a documented synthetic shift
public/reports/ Reproducible experiment output
notebooks/ Executed training exploration and model/drift analysis
docs/ Methodology, model card, and interview preparation
scripts/ Reproduction checks and results figure
.github/workflows/ CI and static dashboard publishing
This project favors inspectable statistics over an opaque “AI health score.” A distribution alert requires PSI ≥ 0.20 and BH-adjusted KS p ≤ 0.05 by default. Missingness alerts use an absolute change of at least 5 percentage points. PSI cutoffs are configurable heuristics, not universal statistical significance thresholds.
The monitor is univariate and numerical. It can miss changes in feature relationships and cannot establish concept drift without label-dependent analysis. Small samples, tied values, dependence between features, repeated monitoring, and exploratory threshold changes affect statistical interpretation. There is no automatic retraining, scheduled ingestion, model registry, or alert delivery service. Read the methodology before adapting the rules to real operations.
Project code: MIT. The bundled demo derives from Wine, by Stefan Aeberhard and M. Forina, UCI Machine Learning Repository, DOI: 10.24432/C5PC7J, licensed under CC BY 4.0. The experiment uses scikit-learn's copy, converts the class labels to zero-based indices, partitions the rows, and creates explicitly labeled synthetic perturbations. These transformations are not endorsed by the dataset authors.
Primary references: UCI dataset, scikit-learn leakage guidance, SciPy KS documentation.
