Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ShiftWatch

A reproducible ML monitoring workbench: detect changes in incoming data and measure their effect on a trained classifier.

Checks Python 3.12 License: MIT

Open the interactive dashboard → · Executed analysis notebook · Methodology · Model card · Interview guide

Reproducible accuracy and drift results

A model's test score is only one part of its lifecycle. ShiftWatch asks what happens when new data no longer resembles the data used for training. It combines a reusable Python drift library, a command-line workflow, a real classification experiment, and an interactive TypeScript report viewer.

The bundled experiment uses UCI Wine. Shifted scenarios are explicitly synthetic; this is an educational monitoring workbench, with no claim of production deployment or business impact.

What it does

  • Trains and selects models: compares a dummy baseline, logistic regression, and a random forest with five-fold stratified cross-validation. Imputation and scaling are fitted inside each fold.
  • Evaluates a separate test set: reports accuracy, macro F1, log loss, a confusion matrix, and bootstrap accuracy intervals.
  • Detects numerical data drift: computes Population Stability Index, two-sample KS tests, Benjamini–Hochberg adjusted p-values, Wasserstein distance, and missingness changes.
  • Works with your data: compares two numerical CSVs and exports a validated JSON report. An optional nonzero exit code supports CI data gates.
  • Explains the result: the dashboard offers scenario switching, feature search, distribution inspection, exploratory thresholds, model comparisons, report import, and export.
  • Reproduces the evidence: committed results, dataset fingerprints, split indices, fixed seeds, dependency locks, and automated tests make the experiment inspectable.

The dashboard reads files locally. It does not send uploaded reports to a server, run Python in the browser, or pretend to be a live production monitor.

Run the experiment

Requires Python 3.12. The dataset ships with scikit-learn; no dataset download, cloud account, or API key is needed after installing dependencies.

git clone https://github.com/pralav-25/shiftwatch.git
cd shiftwatch
python3.12 -m venv .venv
source .venv/bin/activate       # Windows: .venv\Scripts\activate
pip install -r requirements.lock
pip install --no-deps -e .
shiftwatch demo

This writes public/reports/demo.json. Open the dashboard and select Load report to inspect it. For a lightweight library installation without plotting and test tools, use pip install -e . instead of the locked environment.

Analysis notebook

The executed notebook walks through training-only exploration, baseline and cross-validation results, holdout error analysis, and synthetic drift scenarios. Its tables and figures are computed from the same run_experiment function used by the dashboard. Read it directly on GitHub, or rerun every cell from the repository root in the Python 3.12 environment above:

pip install -r requirements-notebook.lock
python -m nbconvert --to notebook --execute \
  --ExecutePreprocessor.timeout=180 \
  --output wine-monitoring-analysis --output-dir /tmp \
  notebooks/wine-monitoring-analysis.ipynb

This saves a fresh executed copy under /tmp and preserves the committed notebook. On Windows, replace /tmp with an existing temporary directory. Notebook tooling has a separate dependency lock that includes the experiment's numerical lock. GitHub Actions executes the notebook to catch stale imports and analysis errors.

Compare your own data

Use two CSVs containing the same named numerical feature columns. Column order can differ. Remove identifiers and target labels, or select the feature columns explicitly. Each file needs at least five rows; missing numeric values are allowed, infinity and nonnumeric columns are rejected.

Give every column a unique, nonempty header. Empty header cells are rejected before the CSV parser can replace them with invented Unnamed: ... feature names. Literal nonempty names are preserved, including a name such as Unnamed: 0. Rows with extra values are rejected instead of silently treating leading values as an index or dropping columns. Check for extra delimiters if this happens; a rejected comparison leaves any existing output report intact.

shiftwatch compare examples/reference.csv examples/shifted.csv --output report.json

# A data-quality gate: exit 2 if any feature alerts; the report is still saved.
shiftwatch compare reference.csv current.csv \
  --psi-threshold 0.2 --alpha 0.05 --fail-on-alert --output report.json

Use --columns temperature pressure to select features without editing the source files. Names must be unique and present in both inputs; the report uses the given order. Unselected identifier, label, and metadata columns are ignored, while CSV header and row-width checks still apply to the whole file. Without this option, all columns are compared and both inputs must have the same feature names.

Use --delimiter ';' to compare semicolon-separated exports. Both files use the same separator; a literal tab or pipe also works. The separator must be one ASCII character other than a newline, NUL, or double quote. Quoted field names and all header/row-width checks use that separator. Comma remains the default.

Use --missing-value=-999 --missing-value MISSING for dataset-specific missing markers. Repeat the option for each marker; they apply to data cells in both files, while literal header names are preserved. Pandas' default blank/NA markers remain enabled. Numeric sentinels remain observed numbers unless explicitly declared. The resulting missing counts feed both missingness alerts and --fail-on-insufficient-data. The input files are not modified.

Use --missing-threshold 0.1 to alert on a missing-rate change of at least ten percentage points, or --bins 5 to choose the number of reference quantile bins for PSI. Defaults are 0.05 and 10; accepted ranges are (0, 1] and integers from 2 to 50. Both settings are saved in the report's config.

Report files are written atomically: a failed write preserves the previous report. The CLI refuses to overwrite either input CSV, including through a symlink or hard link.

Both compare and demo accept --output - to emit only report JSON on standard output, suitable for a pipe or redirection. Validation errors go to standard error and emit no report. Alert and insufficient-data exit flags keep the same meaning.

shiftwatch compare examples/reference.csv examples/shifted.csv --output - | python -m json.tool

Use --output report.json for overwrite protection and atomic writes. Shell redirection such as > report.json is handled by your shell and has neither safeguard.

Load the resulting JSON into the dashboard. Comparisons without a trained model show drift diagnostics without fabricated model-performance metrics. The dashboard accepts up to 200 features, 20 scenarios, and 5 MB per report; the Python library can analyze wider tables.

import pandas as pd
from shiftwatch import compare_frames

result = compare_frames(
    pd.read_csv("examples/reference.csv"),
    pd.read_csv("examples/shifted.csv"),
    psi_threshold=0.2,
)
print(result["alert_count"])

Library thresholds must be finite real scalars; booleans, strings, arrays, and nonfinite values raise ValueError. NumPy scalar settings and integral bin counts are normalized to ordinary Python numbers so the returned report remains JSON-serializable.

Exit status in automation

A successful comparison returns 0 unless --fail-on-alert is enabled and the report contains an alert. With that flag, an alert returns 2 after the report has been written, so downstream tooling can still inspect the result. Without the flag, an alert does not make the command fail.

Use --fail-on-insufficient-data when the comparison must have at least five observed values for every feature in both files. It also returns 2 after saving diagnostics if any feature is untestable, even when neither distribution nor missingness alerts fire. The two flags can be combined; neither is enabled by default. This checks per-feature observations, in addition to the minimum five rows required in each CSV.

Argument and input-validation errors also use exit status 2. Do not treat that status alone as proof of drift: check the command's error output and whether the current run successfully wrote a report. Use a separate output path for each run to avoid mistaking a preserved report from an earlier run for a fresh result.

Results you can reproduce

Seed 42; 124 training rows; 54 held-out test rows. Training CV selects logistic regression, with mean macro F1 of 0.9834 across five folds. These scores describe a small, relatively easy dataset; they are not evidence of general-purpose model quality.

Scenario Accuracy Features flagged Change applied
Untouched holdout 98.1% 0 / 13 None
Moderate shift 83.3% 4 / 13 +0.75 training standard deviations in four features
Severe shift 53.7% 4 / 13 +2 training standard deviations in four features
Missing data 98.1% 3 / 13 Approximately 30% missing in three features

All perturbed scenarios reuse the same test rows and original labels. They are stress tests, not new independent real-world samples. Missingness triggers alerts even when this model's accuracy happens to remain unchanged—one reason to distinguish data quality from prediction quality.

See the full JSON for fold scores, confidence intervals, confusion matrices, feature statistics, exact split indices, and runtime versions.

Run the dashboard locally

Requires Node.js 22.18+ and pnpm 11.19.0.

corepack enable
pnpm install --frozen-lockfile
pnpm dev

Open the localhost URL printed by the development server. The dashboard uses React, TypeScript, Vinext, accessible Base UI/Shadcn primitives, and SVG charts. pnpm build generates a static dashboard in dist/client.

The Python pipeline and UI communicate through the versioned shiftwatch/v1 JSON contract. This keeps the numerical analysis portable and the public demo inexpensive to host.

flowchart LR
  A[UCI Wine or numerical CSVs] --> B[Python validation]
  B --> C[Training-only CV and model selection]
  B --> D[Reference vs current drift analysis]
  C --> E[Held-out evaluation and stress tests]
  D --> F[Versioned JSON report]
  E --> F
  F --> G[React report viewer]
  F --> H[CI alert gate]
Loading

Validation

pytest -q
ruff check src tests
ruff format --check src tests
pnpm test
pnpm typecheck
pnpm lint
pnpm build

# Reproduce and compare the checked-in numerical results:
shiftwatch demo --output /tmp/reproduced.json
python scripts/check_reproduction.py public/reports/demo.json /tmp/reproduced.json

# Regenerate the README figure:
python scripts/plot_results.py

The suite covers constant distributions, overflow bins, known multiple-testing corrections, missing/all-missing columns, invalid schemas, CLI exit behavior, report validation, filter logic, disjoint train/test splits, deterministic results, and a spy that verifies no preprocessor fits on held-out rows.

GitHub Actions runs Python tests, report reproduction, frontend tests, type checking, linting, and a production build. A separate workflow publishes the dashboard to GitHub Pages after checks pass. Browser interaction and visual regression tests are not included; optional WebMCP support is feature-detected and requires a compatible browser for runtime validation.

Project map

src/shiftwatch/       Python drift engine, experiment, and CLI
app/                 Interactive report dashboard
lib/report.ts        Report contract, validation, and alert/filter logic
tests/               Numerical, leakage, reproducibility, and CLI tests
tests-web/           Report-validation and dashboard-logic tests
examples/            Reference data and a documented synthetic shift
public/reports/      Reproducible experiment output
notebooks/           Executed training exploration and model/drift analysis
docs/                Methodology, model card, and interview preparation
scripts/             Reproduction checks and results figure
.github/workflows/   CI and static dashboard publishing

Design decisions and limitations

This project favors inspectable statistics over an opaque “AI health score.” A distribution alert requires PSI ≥ 0.20 and BH-adjusted KS p ≤ 0.05 by default. Missingness alerts use an absolute change of at least 5 percentage points. PSI cutoffs are configurable heuristics, not universal statistical significance thresholds.

The monitor is univariate and numerical. It can miss changes in feature relationships and cannot establish concept drift without label-dependent analysis. Small samples, tied values, dependence between features, repeated monitoring, and exploratory threshold changes affect statistical interpretation. There is no automatic retraining, scheduled ingestion, model registry, or alert delivery service. Read the methodology before adapting the rules to real operations.

Data and license

Project code: MIT. The bundled demo derives from Wine, by Stefan Aeberhard and M. Forina, UCI Machine Learning Repository, DOI: 10.24432/C5PC7J, licensed under CC BY 4.0. The experiment uses scikit-learn's copy, converts the class labels to zero-based indices, partitions the rows, and creates explicitly labeled synthetic perturbations. These transformations are not endorsed by the dataset authors.

Primary references: UCI dataset, scikit-learn leakage guidance, SciPy KS documentation.

About

Reproducible Python ML evaluation and dataset drift monitoring with PSI, KS tests, cross-validation, and an interactive React dashboard.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages