Skip to content

Repository files navigation

♟️ ChessIQ

AI-Powered Chess Analytics & Improvement Platform

Python PyTorch FastAPI MLflow pytest License


🎯 Overview

ChessIQ analyses 4,709 real Chess.com games and asks whether the outcome can be predicted before the first move.

The answer is no — AUC 0.5001 against a 47.5% base rate. Not "close to chance": chance, to four decimal places. Accuracy is 53.5%, which sounds better only because guessing "loss" every time scores 52.5%.

And accuracy flatters it further. Recall is 0.185: there are 292 wins in the test set and the model flags 54 of them. On F1 the trivial "always guess win" baseline wins, 0.644 to 0.274.

Earlier versions of this project reported 72.49% and 78.21%. Both were produced by data leaks. Locating them, quantifying them and rebuilding the pipeline around genuinely pre-game information is what this project is actually about.

Live demo: chessiq.streamlit.app  ·  API docs: chessiq-api.onrender.com/docs

Both run on free tiers and sleep after inactivity — the first request may take ~50 seconds to wake.

📊 Full analysis — docs/results.md  ·  🏗️ Architecture & diagrams — docs/architecture.md


🔍 The Three Leaks

❌ Leak 1 — Post-game engine metrics (found in V1)

The first model used accuracy and avg_centipawn_loss as features. Both are computed by running Stockfish over the finished game. Removing them dropped reported accuracy from 78.21% to 72.49%.

Correctly identified and fixed. But two more were underneath it.

❌ Leak 2 — Post-game ratings

Chess.com writes the rating after the game into the WhiteElo / BlackElo PGN headers. So the "rating difference" feature already contains the result.

Measured on this archive:

rating change ENTERING game n, by result of game n:
    after a WIN   +7.94
    after a LOSS  -8.62      ← the result is already in the rating

rating change LEAVING game n, by result of game n:
    after a WIN   +1.11
    after a LOSS  +0.27      ← nothing

sign(elo_diff) matches the outcome on 71% of games.

This was visible in V1's own README all along: "88.4% win rate vs opponents 50–100 points weaker." No rating system produces an 88% win rate off a 75-point edge. That was the leak reporting itself.

Fix: the true pre-game rating is the rating recorded on the previous game in the same rating pool (bullet, blitz, rapid and daily are rated separately). The opponent's is recovered from the symmetric points exchange. Cost: 23 games of 4,635.

❌ Leak 3 — The target variable was wrong

V1 mapped PGN result 1-0 to "player won" regardless of colour. The player had White in 2,320 games and Black in 2,317, so the label was correct on only 52.6% of games.

V1 reported corrected
Win rate 49.1% (2,276 wins) 50.4% (2,362 wins)

📊 Combined impact

V1 with engine metrics:        78.21%  ❌
V1 without engine metrics:     72.49%  ❌  (still leaking ratings + wrong label)
V2 fully corrected:            53.50%  ✅  (AUC 0.5001 — chance)

📈 Dataset

Total Games:          4,709 parsed → 4,684 after rating reconstruction
Time Span:            5 years (June 2021 – July 2026)
Total Plies:          332,080
Win Rate:             50.4% (2,362 W / 2,081 L / 241 D)
Engine Accuracy:      91.93% average (Stockfish depth 15)

Ratings are pool-specific, so there is no single trajectory:

pool games first → last range
bullet 2,426 406 → 975 peak 1,222
blitz 1,104 874 → 1,059 peak 1,259
rapid 1,152 773 → 1,413 peak 1,449

🤖 Model Results

All on a held-out chronological test set (615 games, base rate 47.5%).

model acc AUC log loss ECE
majority class 0.4748 0.5000 0.6936 0.0290
logistic (Elo diff + colour) 0.4829 0.4769 0.6878 0.0483
logistic (all pre-game) 0.5350 0.5001 0.6859 0.0111
random forest 0.5301 0.5401 0.7062 0.0787
gradient boosting 0.5187 0.5173 0.7047 0.0689

What accuracy was hiding

model acc precision recall F1 AUC
final model 0.5350 0.5294 0.1849 0.2741 0.5001
always predict "win" 0.4748 0.4748 1.0000 0.6439 0.5000
always predict "not-win" 0.5252 0.0000 0.0000 0.0000 0.5000
                 predicted
               not-win    win
  actual  not-win    275     48
          win        238     54

Move the decision threshold from 0.50 to 0.45 and accuracy falls to 0.4634, well below the base rate. The 1.0-point edge is a property of the threshold, not the model.

And it has no discriminative power where it matters. In the 566 of 615 test games with a rating gap inside ±50 — 92% of them — AUC is 0.4785, below chance.

Does the move sequence help? (PyTorch LSTM, 5 seeds each)

model acc AUC
tabular only 0.5285 ±0.0209 0.5173
sequence only (first 30 plies) 0.4712 ±0.0102 0.4661
hybrid (both) 0.4992 ±0.0271 0.4985

Worse than no. Adding the LSTM branch costs 1.9 points of AUC. The two models differ by exactly one component — whether that branch is connected — so the comparison isolates move information precisely.

Do hyperparameters help? (21 configs, 63 runs, 8 axes)

Also no. Test AUC across the whole sweep spans 0.4962–0.5242 — a range of 0.028, against a typical seed standard deviation of 0.019. Every hyperparameter available moves the result less than re-running the same config with a different random seed.

Worse, validation selection doesn't transfer: across 21 configurations, val and test AUC correlate at only r = +0.375 (p = 0.094), and the config that won on validation ranks 8th of 21 on test. Model selection here selects noise — which is why the /v2/experiments/leaderboard endpoint returns a caveat field saying so.


🎯 Why It's This Hard

pre-game rating gap games win rate
−100 or worse 46 0.304
−100 to −50 128 0.492
−50 to +50 4,197 0.496
+50 to +100 133 0.534
+100 or better 180 0.756

Chess.com matchmaking pairs you with near-equal opponents. 89% of games fall inside a 50-point gap, which means the strongest available pre-game feature has already been neutralised by the matchmaker before you sit down.

An accuracy of 53% isn't a broken model. It's a correct measurement of a task with almost no signal in it.


🎯 Corrected Insights

Recomputed with the right label and pre-game ratings (openings with ≥40 games):

Metric V1 claim Corrected
Best opening D31: 60.26% D31: 62.0% (n=79)
Worst opening A04: 39.19% A41: 37.5% (n=56)
vs weaker (+50–100) 88.4% 53.4% (n=133)
Strongest predictor rating difference (59.63%) rating difference alone is worse than useless (AUC 0.4769)

D31 survives as a genuine strength. Most of the rest did not.


🏗️ Technology Stack

ML & Research — PyTorch (LSTM), scikit-learn, NumPy, python-chess Backend — FastAPI, SQLAlchemy, PostgreSQL Chess Analysis — Stockfish (depth 15) Frontend — Streamlit, Plotly Testing — pytest (93 automated tests)

No GPU required. The LSTM is 240k parameters; one epoch takes 1.1 seconds on a single CPU core, and the full five-seed three-family ablation runs in about six minutes.


📁 Project Structure

ChessIQ/
├── training/                    # V2 research code
│   ├── data/
│   │   ├── pgn.py               # parsing + pre-game rating reconstruction
│   │   ├── vocab.py             # SAN vocabulary (train split only)
│   │   ├── build_dataset.py     # prefix truncation, chronological splits
│   │   └── datasets.py          # PyTorch Dataset / DataLoader
│   ├── features.py              # canonical feature vector (shared with API)
│   ├── models.py                # ChessIQNet — switchable branches
│   ├── train.py                 # training loop, multi-seed CLI
│   ├── baselines.py             # classical comparisons
│   ├── finalize.py              # calibration + artifact export
│   └── diagnostics/             # leakage checks
├── backend/
│   ├── serve_v2.py              # standalone prediction API (no database)
│   ├── ml/predictor.py          # artifact loader
│   ├── routes/prediction_v2.py  # V2 endpoints
│   └── ...                      # V1 service, preserved
├── frontend/app.py              # Streamlit dashboard
├── datasets/                    # manifests committed, .npz gitignored
├── experiments/                 # runs.jsonl, checkpoints, results
├── model/chessiq_v2.joblib      # serving artifact
├── docs/results.md              # full write-up
└── tests/                       # 93 tests

Testing

The repository contains 93 automated tests covering:

• Dataset generation • Leakage prevention • Feature engineering • Model architecture • Training pipeline • Calibration • Production artifact • FastAPI inference • Hybrid sequence analyzer

Run all tests:

pytest


🚀 Quick Start

git clone https://github.com/keyurc2332/ChessIQ.git
cd ChessIQ

python -m venv .venv
.\.venv\Scripts\Activate.ps1          # Windows
source .venv/bin/activate             # Mac/Linux
pip install -r requirements-training.txt

Reproduce every number above (~10 minutes, one CPU core):

pytest                                                    # 93 tests
python -m training.data.build_dataset --pgn backend/all_games.pgn \
    --player keyur_2332 --max-plies 30 --out datasets/
python -m training.baselines
python -m training.train --family all --seeds 5 --max-epochs 50
python -m training.finalize

Run it — two terminals:

# terminal 1 — prediction API (no database needed)
pip install -r backend/requirements-serve.txt
cd backend && uvicorn serve_v2:app --reload       # http://127.0.0.1:8000/docs

# terminal 2 — dashboard
pip install -r frontend/requirements.txt
cd frontend && python -m streamlit run app.py     # http://localhost:8501

Or with Docker:

docker compose up api                              # http://localhost:8000/docs
docker compose run --rm training                   # rebuild the model from raw PGN

🔬 Methodology

Prefix truncation. Games are cut to their first 30 plies, and only games longer than 30 plies are kept — a model fed a complete game sees the checkmate.

Chronological splits, snapped to day boundaries. Oldest games train, newest test. A random split would let the model learn from 2026 games to predict 2022 ones. Day-snapping stops one session — up to 104 games — straddling two splits.

Vocabulary fitted on train only. A move appearing solely in test should collapse to <unk>, exactly as at inference. 953 tokens, 0.41% OOV on test.

One feature module. training/features.py is imported by both the dataset builder and the API, with a test asserting identical output. Training/serving skew produces confident wrong answers and no error anywhere.

Five seeds on every neural result. With signal this weak, one run's accuracy is mostly noise.


📸 Screenshots

The predictor tells you when it doesn't know. Two near-equal players fall in the range where the model scored AUC 0.4785 — below chance — so it says so instead of rendering a confident number:

Predictor with no signal

With a real rating gap, the gauge crosses the base-rate line:

Predictor with signal

The whole project in one chart. 480 of 615 test predictions (78%) land between 0.45 and 0.52 — the model almost never commits, which is the correct response to data this uninformative:

Prediction distribution


📊 Dashboard Pages

  1. Dashboard — performance KPIs, results distribution, accuracy histogram
  2. Opening Analysis — win rate by ECO code with sample sizes
  3. Opponent Strength — win rate by pre-game rating gap
  4. Time Control — performance across bullet, blitz, rapid
  5. 5-Year Progress — rating progression per pool
  6. Win Predictor — calibrated probability from the V2 API
  7. AI Tips — recommendations, only where sample size supports them
  8. ML Insights — model comparison, ablation, calibration

Every figure on the dashboard is read from the V2 API or from the JSON the training pipeline writes. Nothing is hardcoded and nothing is simulated — and when the API is unreachable the affected panel is disabled rather than falling back to an invented estimate.


📄 Resume Impact

One-liner:

ChessIQ: discovered and corrected a data leak in Chess.com PGN exports that had
inflated reported chess-outcome prediction accuracy from 53% to 78%; the honest
model scores AUC 0.5001 — exactly chance.

Full bullet:

ChessIQ V2: Deep Learning & MLOps
• Discovered post-game rating leakage in Chess.com PGN exports; sign(elo_diff)
  matched game outcome on 71% of games, inflating reported accuracy from
  53.5% to 78.2%. Reconstructed true pre-game ratings per rating pool
• Found and fixed a target-variable bug that labelled "White won" as "player
  won", making the label correct on only 52.6% of games
• Built a leakage-controlled PyTorch pipeline (prefix truncation, chronological
  day-aligned splits, train-only vocabulary) and ran a 5-seed ablation showing
  move sequences *reduce* performance (hybrid AUC 0.4985 vs tabular 0.5173)
• Demonstrated that matchmaking places 89% of games within a 50-point rating
  gap, making the task near-unpredictable by construction
• Shipped a calibrated FastAPI service (ECE 0.011) with 93 automated tests
  covering the data pipeline, models, production inference and deployment artifacts.
• Python, PyTorch, scikit-learn, FastAPI, pytest, Stockfish, Streamlit

⚠️ Limitations

Single player, 4,709 games — nothing generalises without testing. Opponent pre-game ratings assume a symmetric points exchange, which is approximate under provisional ratings and differing K-factors. Draws count as non-wins. The negative result on move sequences may reflect data volume rather than absence of signal.

See docs/results.md §6.


🚀 Future Work

• Evaluate on the public Lichess database • Transformer-based sequence encoder • Online model monitoring • Explainability using SHAP • Continuous retraining pipeline • Multi-player generalization study

🎓 What This Project Demonstrates

The models are unremarkable — a logistic regression that barely beats a coin flip. The process is the substance:

  • A leak that every standard metric passed — accuracy, AUC, log loss, Brier, ECE 0.011, five-seed replication, 93 automated tests — caught only by comparing output against domain knowledge
  • A target-variable bug inherited from V1
  • An ablation where two models differ by exactly one component
  • Chronological, day-aligned splits that exposed shift a random split would hide
  • A negative result reported as a negative result

The lesson: statistical validation cannot detect a leak that is consistent across every split. Only domain knowledge can.


🔗 Links

📝 License

MIT License — free to use for learning and research.

🙏 Acknowledgments

Chess.com (game data API) · Stockfish (engine) · PyTorch · Streamlit · scikit-learn


Built with ❤️ for Data Science & Chess | Last Updated: July 2026

About

Chess outcome prediction, and the three data leaks that made it look possible. Honest result: AUC 0.5001 - chance.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages