ChessIQ analyses 4,709 real Chess.com games and asks whether the outcome can be predicted before the first move.
The answer is no — AUC 0.5001 against a 47.5% base rate. Not "close to chance": chance, to four decimal places. Accuracy is 53.5%, which sounds better only because guessing "loss" every time scores 52.5%.
And accuracy flatters it further. Recall is 0.185: there are 292 wins in the test set and the model flags 54 of them. On F1 the trivial "always guess win" baseline wins, 0.644 to 0.274.
Earlier versions of this project reported 72.49% and 78.21%. Both were produced by data leaks. Locating them, quantifying them and rebuilding the pipeline around genuinely pre-game information is what this project is actually about.
Live demo: chessiq.streamlit.app · API docs: chessiq-api.onrender.com/docs
Both run on free tiers and sleep after inactivity — the first request may take ~50 seconds to wake.
📊 Full analysis — docs/results.md · 🏗️ Architecture & diagrams — docs/architecture.md
The first model used accuracy and avg_centipawn_loss as features. Both are
computed by running Stockfish over the finished game. Removing them dropped
reported accuracy from 78.21% to 72.49%.
Correctly identified and fixed. But two more were underneath it.
Chess.com writes the rating after the game into the WhiteElo / BlackElo
PGN headers. So the "rating difference" feature already contains the result.
Measured on this archive:
rating change ENTERING game n, by result of game n:
after a WIN +7.94
after a LOSS -8.62 ← the result is already in the rating
rating change LEAVING game n, by result of game n:
after a WIN +1.11
after a LOSS +0.27 ← nothing
sign(elo_diff) matches the outcome on 71% of games.
This was visible in V1's own README all along: "88.4% win rate vs opponents 50–100 points weaker." No rating system produces an 88% win rate off a 75-point edge. That was the leak reporting itself.
Fix: the true pre-game rating is the rating recorded on the previous game in the same rating pool (bullet, blitz, rapid and daily are rated separately). The opponent's is recovered from the symmetric points exchange. Cost: 23 games of 4,635.
V1 mapped PGN result 1-0 to "player won" regardless of colour. The player had
White in 2,320 games and Black in 2,317, so the label was correct on only
52.6% of games.
| V1 reported | corrected | |
|---|---|---|
| Win rate | 49.1% (2,276 wins) | 50.4% (2,362 wins) |
V1 with engine metrics: 78.21% ❌
V1 without engine metrics: 72.49% ❌ (still leaking ratings + wrong label)
V2 fully corrected: 53.50% ✅ (AUC 0.5001 — chance)
Total Games: 4,709 parsed → 4,684 after rating reconstruction
Time Span: 5 years (June 2021 – July 2026)
Total Plies: 332,080
Win Rate: 50.4% (2,362 W / 2,081 L / 241 D)
Engine Accuracy: 91.93% average (Stockfish depth 15)
Ratings are pool-specific, so there is no single trajectory:
| pool | games | first → last | range |
|---|---|---|---|
| bullet | 2,426 | 406 → 975 | peak 1,222 |
| blitz | 1,104 | 874 → 1,059 | peak 1,259 |
| rapid | 1,152 | 773 → 1,413 | peak 1,449 |
All on a held-out chronological test set (615 games, base rate 47.5%).
| model | acc | AUC | log loss | ECE |
|---|---|---|---|---|
| majority class | 0.4748 | 0.5000 | 0.6936 | 0.0290 |
| logistic (Elo diff + colour) | 0.4829 | 0.4769 | 0.6878 | 0.0483 |
| logistic (all pre-game) | 0.5350 | 0.5001 | 0.6859 | 0.0111 |
| random forest | 0.5301 | 0.5401 | 0.7062 | 0.0787 |
| gradient boosting | 0.5187 | 0.5173 | 0.7047 | 0.0689 |
| model | acc | precision | recall | F1 | AUC |
|---|---|---|---|---|---|
| final model | 0.5350 | 0.5294 | 0.1849 | 0.2741 | 0.5001 |
| always predict "win" | 0.4748 | 0.4748 | 1.0000 | 0.6439 | 0.5000 |
| always predict "not-win" | 0.5252 | 0.0000 | 0.0000 | 0.0000 | 0.5000 |
predicted
not-win win
actual not-win 275 48
win 238 54
Move the decision threshold from 0.50 to 0.45 and accuracy falls to 0.4634, well below the base rate. The 1.0-point edge is a property of the threshold, not the model.
And it has no discriminative power where it matters. In the 566 of 615 test games with a rating gap inside ±50 — 92% of them — AUC is 0.4785, below chance.
| model | acc | AUC |
|---|---|---|
| tabular only | 0.5285 ±0.0209 | 0.5173 |
| sequence only (first 30 plies) | 0.4712 ±0.0102 | 0.4661 |
| hybrid (both) | 0.4992 ±0.0271 | 0.4985 |
Worse than no. Adding the LSTM branch costs 1.9 points of AUC. The two models differ by exactly one component — whether that branch is connected — so the comparison isolates move information precisely.
Also no. Test AUC across the whole sweep spans 0.4962–0.5242 — a range of 0.028, against a typical seed standard deviation of 0.019. Every hyperparameter available moves the result less than re-running the same config with a different random seed.
Worse, validation selection doesn't transfer: across 21 configurations, val and
test AUC correlate at only r = +0.375 (p = 0.094), and the config that won
on validation ranks 8th of 21 on test. Model selection here selects noise —
which is why the /v2/experiments/leaderboard endpoint returns a caveat field
saying so.
| pre-game rating gap | games | win rate |
|---|---|---|
| −100 or worse | 46 | 0.304 |
| −100 to −50 | 128 | 0.492 |
| −50 to +50 | 4,197 | 0.496 |
| +50 to +100 | 133 | 0.534 |
| +100 or better | 180 | 0.756 |
Chess.com matchmaking pairs you with near-equal opponents. 89% of games fall inside a 50-point gap, which means the strongest available pre-game feature has already been neutralised by the matchmaker before you sit down.
An accuracy of 53% isn't a broken model. It's a correct measurement of a task with almost no signal in it.
Recomputed with the right label and pre-game ratings (openings with ≥40 games):
| Metric | V1 claim | Corrected |
|---|---|---|
| Best opening | D31: 60.26% | D31: 62.0% (n=79) |
| Worst opening | A04: 39.19% | A41: 37.5% (n=56) |
| vs weaker (+50–100) | 88.4% | 53.4% (n=133) |
| Strongest predictor | rating difference (59.63%) | rating difference alone is worse than useless (AUC 0.4769) |
D31 survives as a genuine strength. Most of the rest did not.
ML & Research — PyTorch (LSTM), scikit-learn, NumPy, python-chess Backend — FastAPI, SQLAlchemy, PostgreSQL Chess Analysis — Stockfish (depth 15) Frontend — Streamlit, Plotly Testing — pytest (93 automated tests)
No GPU required. The LSTM is 240k parameters; one epoch takes 1.1 seconds on a single CPU core, and the full five-seed three-family ablation runs in about six minutes.
ChessIQ/
├── training/ # V2 research code
│ ├── data/
│ │ ├── pgn.py # parsing + pre-game rating reconstruction
│ │ ├── vocab.py # SAN vocabulary (train split only)
│ │ ├── build_dataset.py # prefix truncation, chronological splits
│ │ └── datasets.py # PyTorch Dataset / DataLoader
│ ├── features.py # canonical feature vector (shared with API)
│ ├── models.py # ChessIQNet — switchable branches
│ ├── train.py # training loop, multi-seed CLI
│ ├── baselines.py # classical comparisons
│ ├── finalize.py # calibration + artifact export
│ └── diagnostics/ # leakage checks
├── backend/
│ ├── serve_v2.py # standalone prediction API (no database)
│ ├── ml/predictor.py # artifact loader
│ ├── routes/prediction_v2.py # V2 endpoints
│ └── ... # V1 service, preserved
├── frontend/app.py # Streamlit dashboard
├── datasets/ # manifests committed, .npz gitignored
├── experiments/ # runs.jsonl, checkpoints, results
├── model/chessiq_v2.joblib # serving artifact
├── docs/results.md # full write-up
└── tests/ # 93 tests
The repository contains 93 automated tests covering:
• Dataset generation • Leakage prevention • Feature engineering • Model architecture • Training pipeline • Calibration • Production artifact • FastAPI inference • Hybrid sequence analyzer
Run all tests:
pytest
git clone https://github.com/keyurc2332/ChessIQ.git
cd ChessIQ
python -m venv .venv
.\.venv\Scripts\Activate.ps1 # Windows
source .venv/bin/activate # Mac/Linux
pip install -r requirements-training.txtReproduce every number above (~10 minutes, one CPU core):
pytest # 93 tests
python -m training.data.build_dataset --pgn backend/all_games.pgn \
--player keyur_2332 --max-plies 30 --out datasets/
python -m training.baselines
python -m training.train --family all --seeds 5 --max-epochs 50
python -m training.finalizeRun it — two terminals:
# terminal 1 — prediction API (no database needed)
pip install -r backend/requirements-serve.txt
cd backend && uvicorn serve_v2:app --reload # http://127.0.0.1:8000/docs
# terminal 2 — dashboard
pip install -r frontend/requirements.txt
cd frontend && python -m streamlit run app.py # http://localhost:8501Or with Docker:
docker compose up api # http://localhost:8000/docs
docker compose run --rm training # rebuild the model from raw PGNPrefix truncation. Games are cut to their first 30 plies, and only games longer than 30 plies are kept — a model fed a complete game sees the checkmate.
Chronological splits, snapped to day boundaries. Oldest games train, newest test. A random split would let the model learn from 2026 games to predict 2022 ones. Day-snapping stops one session — up to 104 games — straddling two splits.
Vocabulary fitted on train only. A move appearing solely in test should
collapse to <unk>, exactly as at inference. 953 tokens, 0.41% OOV on test.
One feature module. training/features.py is imported by both the dataset
builder and the API, with a test asserting identical output. Training/serving
skew produces confident wrong answers and no error anywhere.
Five seeds on every neural result. With signal this weak, one run's accuracy is mostly noise.
The predictor tells you when it doesn't know. Two near-equal players fall in the range where the model scored AUC 0.4785 — below chance — so it says so instead of rendering a confident number:
With a real rating gap, the gauge crosses the base-rate line:
The whole project in one chart. 480 of 615 test predictions (78%) land between 0.45 and 0.52 — the model almost never commits, which is the correct response to data this uninformative:
- Dashboard — performance KPIs, results distribution, accuracy histogram
- Opening Analysis — win rate by ECO code with sample sizes
- Opponent Strength — win rate by pre-game rating gap
- Time Control — performance across bullet, blitz, rapid
- 5-Year Progress — rating progression per pool
- Win Predictor — calibrated probability from the V2 API
- AI Tips — recommendations, only where sample size supports them
- ML Insights — model comparison, ablation, calibration
Every figure on the dashboard is read from the V2 API or from the JSON the training pipeline writes. Nothing is hardcoded and nothing is simulated — and when the API is unreachable the affected panel is disabled rather than falling back to an invented estimate.
One-liner:
ChessIQ: discovered and corrected a data leak in Chess.com PGN exports that had
inflated reported chess-outcome prediction accuracy from 53% to 78%; the honest
model scores AUC 0.5001 — exactly chance.
Full bullet:
ChessIQ V2: Deep Learning & MLOps
• Discovered post-game rating leakage in Chess.com PGN exports; sign(elo_diff)
matched game outcome on 71% of games, inflating reported accuracy from
53.5% to 78.2%. Reconstructed true pre-game ratings per rating pool
• Found and fixed a target-variable bug that labelled "White won" as "player
won", making the label correct on only 52.6% of games
• Built a leakage-controlled PyTorch pipeline (prefix truncation, chronological
day-aligned splits, train-only vocabulary) and ran a 5-seed ablation showing
move sequences *reduce* performance (hybrid AUC 0.4985 vs tabular 0.5173)
• Demonstrated that matchmaking places 89% of games within a 50-point rating
gap, making the task near-unpredictable by construction
• Shipped a calibrated FastAPI service (ECE 0.011) with 93 automated tests
covering the data pipeline, models, production inference and deployment artifacts.
• Python, PyTorch, scikit-learn, FastAPI, pytest, Stockfish, Streamlit
Single player, 4,709 games — nothing generalises without testing. Opponent pre-game ratings assume a symmetric points exchange, which is approximate under provisional ratings and differing K-factors. Draws count as non-wins. The negative result on move sequences may reflect data volume rather than absence of signal.
See docs/results.md §6.
• Evaluate on the public Lichess database • Transformer-based sequence encoder • Online model monitoring • Explainability using SHAP • Continuous retraining pipeline • Multi-player generalization study
The models are unremarkable — a logistic regression that barely beats a coin flip. The process is the substance:
- A leak that every standard metric passed — accuracy, AUC, log loss, Brier, ECE 0.011, five-seed replication, 93 automated tests — caught only by comparing output against domain knowledge
- A target-variable bug inherited from V1
- An ablation where two models differ by exactly one component
- Chronological, day-aligned splits that exposed shift a random split would hide
- A negative result reported as a negative result
The lesson: statistical validation cannot detect a leak that is consistent across every split. Only domain knowledge can.
MIT License — free to use for learning and research.
Chess.com (game data API) · Stockfish (engine) · PyTorch · Streamlit · scikit-learn
Built with ❤️ for Data Science & Chess | Last Updated: July 2026


