Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 24 additions & 2 deletions .github/workflows/backtest.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,9 @@
# Manual: walk-forward replay of the scanner over the last N sessions on the
# runner's open internet. The Markdown report lands in the job summary.
# gh workflow run backtest.yml [-f days=63] [-f horizon=40] [-f min_score=60]
# gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60
# (the second form replays two years with each scan seeing the nightly's 500-bar
# window and reports the two years side by side: the out-of-sample check).
# Also runs on pushes that touch its own files (first registration / smoke test).
name: backtest

Expand Down Expand Up @@ -31,6 +34,18 @@ on:
description: "space-separated KEY=VALUE constant overrides, e.g. CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45"
required: false
default: ""
period:
description: "yfinance history period (2y = what the nightly downloads; 5y for a two-year replay)"
required: false
default: "2y"
bars:
description: "bars of history each scan sees (empty = all; 500 mimics the nightly 2y window in a long replay)"
required: false
default: ""
split:
description: "YYYY-MM-DD: also report the sessions before and from this date as separate windows"
required: false
default: ""
push:
paths:
- tools/backtest.py
Expand All @@ -42,7 +57,9 @@ permissions:
jobs:
backtest:
runs-on: ubuntu-latest
timeout-minutes: 90
# A runner replays about 5 s per session per profile over the full index; --grid runs two
# profiles, so 63 sessions take about 10 minutes and a two-year (500-session) replay about 90.
timeout-minutes: 180
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
Expand All @@ -59,9 +76,14 @@ jobs:
PROFILE: ${{ github.event.inputs.profile || 'spec' }}
ABLATE: ${{ github.event.inputs.ablate || 'false' }}
OVERRIDES: ${{ github.event.inputs.overrides || '' }}
PERIOD: ${{ github.event.inputs.period || '2y' }}
BARS: ${{ github.event.inputs.bars || '' }}
SPLIT: ${{ github.event.inputs.split || '' }}
run: |
if [ "$ABLATE" = "true" ]; then MODE=--ablate; else MODE=--grid; fi
cmd=(python tools/backtest.py --days "$DAYS" --horizon "$HORIZON" --min-score "$MIN_SCORE"
--profile "$PROFILE" "$MODE" --json backtest.json)
--profile "$PROFILE" --period "$PERIOD" "$MODE" --json backtest.json)
if [ -n "$BARS" ]; then cmd+=(--bars "$BARS"); fi
if [ -n "$SPLIT" ]; then cmd+=(--split "$SPLIT"); fi
for kv in $OVERRIDES; do cmd+=(--set "$kv"); done
"${cmd[@]}" | tee -a "$GITHUB_STEP_SUMMARY"
22 changes: 22 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,28 @@ All notable changes to this project are documented here. Format follows

## [Unreleased]

### Added (backtest statistics and out-of-sample split, #91)
- Every backtest summary (overall, per pattern, per score bucket, other
profile) carries median R, standard deviation, total R, the deepest drawdown
of the cumulative R curve (1 R per trade in scan order), a 95 % bootstrap
interval of the mean R and the drawdown exceeded in 5 % of resamples. The
bootstrap resamples scan months, not trades: signals cluster in time (33 in
June 2025, 2 in March 2026), so a per-trade interval would be too narrow.
- `--split YYYY-MM-DD` reports the sessions before and from a date as
separate windows next to the pooled one, so a rule chosen on one window is
judged on the other; `--bars N` limits what each scan sees so a `--period
5y --days 500 --bars 500` replay reproduces two years of nightly runs.
Workflow inputs `period`, `bars`, `split`. Background: every calibration
to date was judged on the year the values were chosen on; a first two-year
replay found tuned at +0.30 R out of sample against +0.28 in-sample, legacy
at +0.34 against +0.09.

### Fixed
- Backtest: a next open at or below the stop was scored as a stop with an
undefined R, counted in the hit rate but dropped from mean R (INVH
2024-10-01, UAL 2025-11-17). It is now `below_stop`: not traded, counted
with the gaps (#94).

### Changed (calibration, 2026-09-06 replay of the review rules)
- Tuned profile: `MIN_REWARD_RISK = 1.0`. Year-long replay (250 sessions,
horizon 60, tuned): no minimum 177 signals at +0.28 R, 1.0 -> 158 at
Expand Down
4 changes: 3 additions & 1 deletion docs/wiki/03-Configuration-and-Tuning.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,4 +173,6 @@ For sensitivity, run the negative-control mutations (`test_patterns.py -k violat
python -m pytest test_patterns.py -k "violations or controls" -q
```

For real-data effect, run the backtest (see [Testing and Contributing](04-Testing-and-Contributing.md)): it replays the active profile and the other one on the same year of prices, and re-scores the signals under stop and target variants. Any constant can be overridden for one replay without a code change: `gh workflow run backtest.yml -f overrides="CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45"` (locally `--set KEY=VALUE`). Change one constant, re-run, and only then change the default.
For real-data effect, run the backtest (see [Testing and Contributing](04-Testing-and-Contributing.md)): it replays the active profile and the other one on the same prices, and re-scores the signals under stop and target variants. Any constant can be overridden for one replay without a code change: `gh workflow run backtest.yml -f overrides="CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45"` (locally `--set KEY=VALUE`). Change one constant, re-run, and only then change the default.

Every calibration above was judged on a single year, the one the candidate values were chosen on. The backtest now reports a 95 % interval of the mean R (resampling scan months, because signals cluster in time) and, with `-f split=YYYY-MM-DD`, the sessions before and from a date as separate windows, so a value chosen on one window can be confirmed on the other before it becomes the default: `gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60`. A first such replay (2026-09-07, local) found the tuned profile at +0.30 R on 167 signals in the year before 2025-09-09, which no decision had used, against +0.28 R on 159 in the tuning year; legacy fell from +0.34 to +0.09 between the same two windows.
16 changes: 15 additions & 1 deletion docs/wiki/04-Testing-and-Contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,20 @@ Work on a feature branch and open a PR to `main`; the `tests` workflow must pass

## Measuring signal outcomes

`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open (opens above the row's Max buy are "gapped", not traded), and classifies the outcome within a horizon. Output: overall, per-pattern and per-score-bucket hit rates and mean R, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE, every signal, and (with `--grid`) the same signals re-scored under stop-distance, stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile. Caveats: today's constituents only (survivorship bias), and the last `horizon` sessions are still open. Dispatch with `gh workflow run backtest.yml -f days=63 -f horizon=40`; about 2 minutes for the full index.
`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open, and classifies the outcome within a horizon. Two kinds of open are not traded and counted separately: one above the row's Max buy (`gap`) and one at or below the stop (`below_stop`, where nobody buys and R would be undefined). Output, overall, per pattern and per score bucket: hit rate, mean R, median R, standard deviation, total R, the deepest drawdown of the cumulative R curve (1 R per trade, in scan order), a 95 % bootstrap interval of the mean R and the drawdown exceeded in only 5 % of resamples, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE; then every signal; and (with `--grid`) the same signals re-scored under stop-distance, stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile.

The bootstrap resamples scan **months**, not trades: signals arrive in clusters (33 in one month and 2 in another over 2024-26), so treating them as independent draws would make the interval far too narrow. An interval is only printed with at least two months and five trades.

**Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets.

A runner replays about 5 seconds per session per profile over the full index, and the workflow always runs `--grid`, which adds the other profile's pass:

```bash
gh workflow run backtest.yml -f days=63 -f horizon=40 # about 10 minutes
```

```bash
gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60 # two years, about 90 minutes
```

`tools/evaluate_signals.py` reads every version of `output/signals.json` from git history (the daily scan commits one per run), keeps the first appearance of each `CONFIRMED` signal keyed on `(ticker, pattern, stop)`, fetches the bars that followed, and classifies each as `target` (High reached the target before Low touched the stop), `stop`, `open` (neither within the horizon, marked to the last close) or `no_data`. R multiples are `(exit − entry) / (entry − stop)`. Run it on GitHub (`gh workflow run evaluate-signals.yml -f horizon=60`) because market-data hosts may be blocked locally.
108 changes: 97 additions & 11 deletions test_backtest.py
Original file line number Diff line number Diff line change
@@ -1,11 +1,12 @@
#!/usr/bin/env python3
"""Tests for tools/backtest.py: no look-ahead, first-seen signals, fills and breakdowns (offline)."""
"""Tests for tools/backtest.py: no look-ahead, first-seen signals, fills, statistics and breakdowns (offline)."""

from __future__ import annotations

import os
import sys

import numpy as np
import pandas as pd
import pytest

Expand Down Expand Up @@ -40,6 +41,34 @@ def test_walk_forward_marks_gaps_and_no_data(mini_universe):
assert [r["outcome"] for r in rows] == ["no_data"]


def test_walk_forward_does_not_trade_an_open_at_or_below_the_stop(mini_universe):
"""An open already through the stop is no trade (R would be undefined), counted next to the gaps."""
cup = mini_universe["CUP"].copy()
(today,) = scan.detect_cup_and_handle(cup, "CUP") # seen at bar -2, filled at bar -1's open
cup.loc[cup.index[-1], ["Open", "Low"]] = [today.stop - 0.01, today.stop - 0.5]
(r,) = bt.walk_forward({"CUP": cup}, days=5, horizon=10)
assert r["outcome"] == "below_stop" and r["r"] is None and r["fill"] == round(today.stop - 0.01, 2)
stats = bt.breakdown([r])["overall"]
assert (stats["below_stop"], stats["gap"], stats["n"], stats["mean_r"]) == (1, 0, 0, None)
md = bt.render([r], bt.report_sections([r], 5, 10), 5, 10)
assert "| all | 1 | 1 | 0 | 0 | 0 | - | - |" in md # "Not traded" counts it


def test_walk_forward_bars_limits_what_each_scan_sees(mini_universe, monkeypatch):
seen = []
real = scan.scan_symbol

def spy(sym, df, detectors=None):
seen.append(len(df))
return real(sym, df, detectors)

monkeypatch.setattr(scan, "scan_symbol", spy)
bt.walk_forward(mini_universe, days=3, horizon=5, bars=120)
assert seen and set(seen) == {120} # every scan saw exactly the last 120 bars
seen.clear()
assert bt.walk_forward(mini_universe, days=3, horizon=5, bars=50) == [] and seen == [] # < 60 bars: skipped


def _bars(*rows):
"""rows: (open, high, low, close) per day."""
idx = pd.bdate_range("2026-01-05", periods=len(rows))
Expand Down Expand Up @@ -72,15 +101,64 @@ def test_classify_variant_close_vs_intraday_stops():
"outcome": "stop", "bars": 1, "exit": 93.5, "r": -1.3}


# --------------------------------------------------------------------------- #
# Statistics: drawdown, month-block bootstrap, windows
# --------------------------------------------------------------------------- #
def _trade(day, r, ticker="T"):
return {"scan_day": day, "ticker": ticker, "r": r}


def test_max_drawdown():
assert bt.max_drawdown([]) is None
assert bt.max_drawdown([1.0, 1.0]) == 0.0
assert bt.max_drawdown([-1.0, 2.0]) == -1.0 # the curve starts at 0: a losing first trade counts
assert bt.max_drawdown([2.0, -1.0, -1.0, 0.5]) == -2.0 # peak 2 -> trough 0
assert bt.max_drawdown([0.5, -1.0, -1.0, 3.0, -0.5]) == -2.0 # 0.5 -> -1.5 is a fall of 2; -0.5 later is less


def test_block_bootstrap_resamples_months():
flat = [_trade(f"2026-{m:02d}-10", 1.0) for m in (1, 1, 2, 2, 3, 3)]
assert bt.block_bootstrap(flat, n_boot=50) == {"ci_low": 1.0, "ci_high": 1.0, "dd_p95": 0.0, "blocks": 3}
# Fewer than two months, or fewer than five trades: an interval would mean nothing.
assert bt.block_bootstrap(flat[:2], n_boot=10)["ci_low"] is None
assert bt.block_bootstrap(flat[:4], n_boot=10) == {"ci_low": None, "ci_high": None, "dd_p95": None, "blocks": 2}
# A mixed sample: the interval brackets the mean, the worst-case drawdown is at least the observed one,
# rows without r are ignored, and the result is reproducible.
rs = (2.0, -1.0, -1.0, 2.0, -1.0, 2.0, -1.0, -1.0)
mixed = [_trade(f"2026-{m:02d}-10", r) for m, r in zip((1, 1, 2, 2, 3, 3, 4, 4), rs)] + [_trade("2026-05-10", None)]
b = bt.block_bootstrap(mixed, n_boot=500, seed=1)
assert b["blocks"] == 4 and b["ci_low"] < sum(rs) / len(rs) < b["ci_high"]
assert b["dd_p95"] <= bt.max_drawdown(rs) <= 0
assert bt.block_bootstrap(mixed, n_boot=200) == bt.block_bootstrap(mixed, n_boot=200)


def test_split_windows_and_report_sections(mini_universe):
rows = bt.walk_forward(mini_universe, days=5, horizon=10)
assert bt.split_windows(rows, None) == [("all sessions", rows)]
split = max(r["scan_day"] for r in rows)
wins = bt.split_windows(rows, split)
assert [w[0] for w in wins] == ["all sessions", f"before {split}", f"from {split}"]
assert len(wins[1][1]) + len(wins[2][1]) == len(rows) and wins[2][1]
assert all(r["scan_day"] < split for r in wins[1][1]) and all(r["scan_day"] >= split for r in wins[2][1])
sections = bt.report_sections(rows, 5, 10, split=split, data=mini_universe, do_grid=True)
assert [s["label"] for s in sections] == [w[0] for w in wins]
assert sections[0]["stats"]["overall"]["signals"] == len(rows) and len(sections[0]["grid"]) == len(bt.GRID)
assert sections[2]["stats"]["overall"]["signals"] == len(wins[2][1])
md = bt.render(rows, sections, 5, 10)
assert "## All sessions" in md and f"## Before {split}" in md and f"## From {split}" in md
assert md.count("### Stop / target variants") == 3 and "judged on the other" in md
assert "## Signals" in md


def test_grid_rescores_the_same_signals(mini_universe):
rows = bt.walk_forward(mini_universe, days=5, horizon=10)
g = bt.grid(rows, mini_universe, 10)
assert len(g) == len(bt.GRID) == 18
assert {(x["stop_extra_atr"], x["stop_basis"], x["target_mode"]) for x in g} == set(bt.GRID)
traded = sum(1 for r in rows if r["outcome"] in ("target", "stop", "open"))
assert all(x["n"] == traded for x in g)
md = bt.render(rows, bt.breakdown(rows), 5, 10, g)
assert "## Stop / target variants" in md and "| 0.75 | close | breakout |" in md
md = bt.render(rows, bt.report_sections(rows, 5, 10, data=mini_universe, do_grid=True), 5, 10)
assert "### Stop / target variants" in md and "| 0.75 | close | breakout |" in md


def test_breakout_level_target_applies_to_cups_only(mini_universe):
Expand Down Expand Up @@ -126,8 +204,11 @@ def test_profile_pass_replays_the_other_rule_set_and_restores_the_active_one(min
legacy_targets = {r["ticker"]: r["target"] for r in other["rows"] if r["pattern"] == "Cup & Handle"}
spec_targets = {r["ticker"]: r["target"] for r in rows if r["pattern"] == "Cup & Handle"}
assert spec_targets and legacy_targets and spec_targets != legacy_targets # left-rim vs right-rim measure
md = bt.render(rows, bt.breakdown(rows), 5, 10, None, other)
assert "## Rule profile comparison: spec (above) vs legacy (below)" in md and "| legacy: all |" in md
sections = bt.report_sections(rows, 5, 10, other_rows=other["rows"], other_name="legacy")
assert sections[0]["other"]["stats"]["overall"]["signals"] == other["stats"]["overall"]["signals"]
md = bt.render(rows, sections, 5, 10)
assert "### Rule profile comparison: spec (above) vs legacy (below), all sessions" in md
assert "| legacy: all |" in md


def test_breakdown_and_render():
Expand All @@ -140,13 +221,18 @@ def row(pattern, score, outcome, r):
rows = [row("Cup & Handle", 65, "target", 2.0), row("Cup & Handle", 72, "stop", -1.0),
row("Bullish Wolfe Wave", 85, "open", 0.4), row("Cup & Handle", 91, "gap", None)]
stats = bt.breakdown(rows)
assert stats["overall"]["signals"] == 4 and stats["overall"]["gap"] == 1
assert stats["overall"]["target"] == 1 and stats["overall"]["stop"] == 1 and stats["overall"]["open"] == 1
assert stats["overall"]["hit_rate"] == 0.5 and stats["overall"]["mean_r"] == round((2.0 - 1.0 + 0.4) / 3, 3)
o = stats["overall"]
assert (o["signals"], o["gap"], o["below_stop"]) == (4, 1, 0)
assert (o["target"], o["stop"], o["open"]) == (1, 1, 1)
assert o["hit_rate"] == 0.5 and o["mean_r"] == round((2.0 - 1.0 + 0.4) / 3, 3)
assert o["median_r"] == 0.4 and o["total_r"] == 1.4
assert o["std_r"] == round(float(np.std([2.0, -1.0, 0.4], ddof=1)), 3)
assert o["max_dd"] == -1.0 # 2 -> 1 in scan order
assert o["ci_low"] is None and o["dd_p95"] is None and o["blocks"] == 1 # one month: no interval
assert set(stats["by_score"]) == {"60-69", "70-79", "80-89", "90-100"}
assert stats["by_pattern"]["Bullish Wolfe Wave"]["hit_rate"] is None # nothing resolved
assert stats["overall"]["success5"] == 0.5 and stats["overall"]["mfe"] == 0.08
md = bt.render(rows, stats, 63, 40)
assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | 50% | +8.0% | -2.0% |" in md
assert o["success5"] == 0.5 and o["mfe"] == 0.08
md = bt.render(rows, bt.report_sections(rows, 63, 40), 63, 40)
assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | +0.40 | +1.40 | - | -1.00 | - | 50% | +8.0% | -2.0% |" in md
assert "| 2026-01-05 | T | Cup & Handle | 91 | 100.0 | 100.5 | 95.0 | 110.0 | gap | 0 | - |" in md
assert isinstance(pd.DataFrame(rows), pd.DataFrame)
Loading
Loading