engstat puts the statistics engineers actually use - control charts, capability, designed
experiments, calibration curves, gauge R&R, Weibull life analysis, measurement-uncertainty
budgets, tolerance intervals, equivalence tests and sample-size planning - behind a small, readable
API built on NumPy and SciPy. Every result comes with its uncertainty, and every method is checked
against an authoritative reference.
from engstat import spc, regression, reliability, uncertainty
cap = spc.capability(diameters, lsl=11.95, usl=12.05)
print(cap) # Cp, Cpk with 95 % confidence intervals, Pp, Ppk, expected ppm
cal = regression.Calibration(standards, responses)
cal.concentration([0.842, 0.851, 0.838], replicates=3) # (value, lower, upper)
cal.lod(), cal.loq()
life = reliability.weibull_fit(hours, failed) # suspensions handled correctly
life.b_life(10) # B10 life
budget = uncertainty.propagate(lambda m, d, h: 4*m/(3.14159*d**2*h),
values, standard_uncertainties, dof={"m": 5})
print(budget) # sensitivity coefficients, % contributions, nu_eff, k, U- Answers engineering questions, not just statistical ones. Tolerance intervals for design allowables, TOST for "is the new method equivalent?", inverse prediction for calibration, censored Weibull for tests stopped before every unit fails.
- Validated, not just tested. 171 checks in
docs/VALIDATION.md, rerun in CI: NIST StRD certified results (Longley linear regression to ~1e-11; Misra1a nonlinear to all certified digits of the residual SD; real measurement data - silicon resistivity, the atomic weight of silver, ozone-monitor and ultrasonic calibrations - read from the bundled files), control-chart constants computed by integration and compared with published tables, Montgomery's factorial example, statsmodels cross-checks, exact results and large simulations for interval coverage and estimator bias. - Honest about uncertainty. Confidence intervals for Cpk, Weibull parameters, calibration results and regression predictions; warnings when assumptions are shaky (small expected counts, ridge systems in response surfaces).
- Lightweight. Only NumPy and SciPy are required; plotting and apps are optional extras.
pip install "engstat[viz] @ git+https://github.com/ktwyw/engstat" # library + plotting
# or, for development:
git clone https://github.com/ktwyw/engstat && cd engstat && pip install -e ".[dev]"| Module | Methods |
|---|---|
describe |
summaries (robust MAD/IQR), CIs for mean/variance/SD/proportion (Wilson, Clopper-Pearson), BCa bootstrap, exact normal tolerance intervals |
hypothesis |
t tests (Welch default), paired, F and Levene, chi-square, Shapiro-Wilk and Anderson-Darling, Mann-Whitney, Wilcoxon, TOST equivalence, power and sample size, Grubbs outliers |
regression |
OLS/WLS/polynomial with full inference, confidence and prediction bands, leverage, studentized residuals, Cook's D, VIF, Durbin-Watson; calibration (inverse prediction, LOD, LOQ); nonlinear least squares with parameter CIs; Deming regression |
anova |
one-way with Tukey HSD, two-way with interaction |
doe |
2^k and 2^(k-p) designs, effects with Lenth's method, central-composite and Box-Behnken designs, response surfaces with canonical analysis |
spc |
X-bar/R, X-bar/S, I-MR (Phase I and II), p, np, c, u, EWMA, CUSUM, Western Electric rules, ARL, capability with CIs, non-normal capability (percentile method, Box-Cox) |
msa |
crossed gauge R&R (ANOVA method, AIAG): variance components, %study variation, %tolerance, ndc; attribute agreement (within, between, vs standard; Cohen's and Fleiss' kappa) |
sampling |
acceptance sampling: OC curves (binomial and hypergeometric), single plans designed from AQL/LTPD and risks, AOQ/AOQL/ATI, variables plans (k-method) with exact OC |
reliability |
Weibull MLE with censoring and rank regression, B-lives, Kaplan-Meier with Greenwood limits, MTBF confidence limits |
uncertainty |
GUM propagation with correlation and Welch-Satterthwaite; Monte Carlo (JCGM 101) with shortest or symmetric coverage intervals |
distributions |
maximum-likelihood fitting and AIC ranking, probability-plot data |
datasets |
bundled data files - real NIST reference measurements and course data - with sources and licences |
viz |
control charts, probability and Weibull plots, effects Pareto and half-normal plots, regression bands, residual diagnostics, capability histograms |
Each executed notebook starts from a practical question, shows by simulation why a method works,
applies it to realistic engineering cases, writes the key algorithm out in plain Python ("Inside the
algorithm", checked against the library) and ends with exercises, including an "Implement it
yourself" task. Worked, verified solutions to all 57 exercises are in solutions/.
Every notebook opens in Google Colab and installs engstat automatically.
| # | Notebook | Engineering cases |
|---|---|---|
| 00 | Python, NumPy and pandas for statistics | five probes on one silicon wafer (NIST) |
| 01 | Describing data and checking it | robust statistics, outliers and normality |
| 02 | Distributions and the central limit theorem | fitting and comparing life and strength distributions |
| 03 | Confidence, prediction and tolerance intervals | an A-basis strength allowable |
| 04 | The bootstrap | particle-size percentiles |
| 05 | Hypothesis tests and p-values | do two instruments agree? atomic weight of silver (NIST) |
| 06 | Power, sample size and equivalence | planning a test; method comparison with TOST and Deming regression |
| 07 | Linear regression and diagnostics | Longley collinearity; a heat-transfer correlation |
| 08 | Calibration and nonlinear models | HPLC with LOD/LOQ; Arrhenius; ozone-monitor and ultrasonic calibration (NIST) |
| 09 | Analysis of variance | catalysts; machining; silicon probes (NIST) |
| 10 | Design of experiments | filtration-rate factorial; screening; reactor response surface |
| 11 | Statistical process control and capability | shaft diameters; EWMA/CUSUM; attribute charts; bore capability |
| 12 | Measurement system analysis | gauge R&R |
| 13 | Reliability and life data | bearing life with suspensions; Kaplan-Meier; demonstration tests |
| 14 | Measurement uncertainty | density budget (GUM); orifice meter (Monte Carlo) |
| 15 | From a messy data file to a decision | cleaning supplier certificates, then a specification decision |
| 16 | Acceptance sampling | incoming inspection of bolts: OC curves, AQL/LTPD plans, AOQL, variables plans |
| 17 | Capability for non-normal data | surface roughness: percentile method, Box-Cox, tail uncertainty |
| 18 | Attribute agreement | visual weld inspection: repeatability, reproducibility, kappa, error types |
engstat.datasets ships real, public-domain NIST reference measurements with certified results -
silicon resistivity from five probes (SiRstv.dat), the atomic weight of silver on two instruments
(AtmWtAg.dat), an ozone-monitor calibration (Norris.dat), an ultrasonic calibration
(Chwirut2.dat) and the Longley data - plus a deliberately messy tensile-test file for practising
data cleaning. Read them exactly as you would read your own files:
import numpy as np
from engstat import datasets
data = np.loadtxt(datasets.path("SiRstv.dat"), skiprows=60)
print(datasets.info("AtmWtAg.dat"))- Interactive toolkit - sample size, capability, control charts, Weibull,
calibration and Monte Carlo uncertainty in the browser:
streamlit run showcase/app.py. - Methods and references - the formula and source behind every function.
| Reference | What is checked | Agreement |
|---|---|---|
| NIST StRD Longley (higher difficulty) | 7 coefficients, SEs, residual SD, R², ANOVA | rel. error ≤ 1e-10 |
| NIST StRD Misra1a (nonlinear) | parameters from both NIST start points, SEs, residual SD | ≤ 1e-7 (SD 1e-9) |
| NIST StRD real data: SiRstv, AtmWtAg (ANOVA), Norris (regression), Chwirut2 (nonlinear) | sums of squares, F, R², coefficients, residual SD - read from the bundled files | ≤ 1e-9 (Chwirut2 ≤ 1e-7) |
| Montgomery, SQC tables | d2, d3, c4, A2, D3, D4 for n = 2-10 | all tabulated digits |
| Brute force, simulation, statsmodels | sampling plans (minimal n), AOQL, variables-plan OC; Box-Cox λ; Cohen's and Fleiss' kappa (incl. published example), Clopper-Pearson | exact / ≤ 1e-12 |
| Montgomery, DAE Ex. 6.2 | all 15 factorial effects; active set by Lenth | exact |
| Tolerance-factor tables + simulation | one-sided k; two-sided confidence achieved | ±0.0005; ±0.003 |
R power.t.test, statsmodels |
sample sizes, power, proportion CIs, diagnostics, two-way ANOVA | exact / ≤ 1e-10 |
| Exact results and simulation | GUM, Welch-Satterthwaite, Monte Carlo, gauge R&R bias, Cp CI coverage | within tolerance |
Run it yourself: python docs/validate.py.
If engstat helps your work, please cite it using CITATION.cff (GitHub shows a
"Cite this repository" button).
Yanwei Wang - personal open-source project. GitHub @ktwyw · ORCID 0000-0002-8488-9833 · wangyanwei@gmail.com
Also by the author: fluidmech - a validated fluid-mechanics library and course.
MIT - see LICENSE. Contributions welcome: see CONTRIBUTING.md.






