From 1fdddd75957d2d5b12d4150025f3995a08c578b2 Mon Sep 17 00:00:00 2001 From: Yaniv Bernhard Date: Mon, 7 Sep 2026 22:10:48 +0300 Subject: [PATCH 1/2] feat: backtest confidence intervals, drawdown, --split and --bars; a fill through the stop is not traded Every backtest summary now carries median R, standard deviation, total R, the deepest drawdown of the cumulative R curve, a 95 % bootstrap interval of the mean R and the drawdown exceeded in 5 % of resamples. The bootstrap resamples scan months rather than trades because signals cluster in time. --split YYYY-MM-DD reports the sessions before and from a date as separate windows so a rule chosen on one window is judged on the other; --bars N limits what each scan sees so a long download replays the nightly's 2y window. Workflow inputs period, bars and split. A next open at or below the stop was scored as a stop with an undefined R, counted in the hit rate but dropped from mean R; it is now below_stop, not traded, counted with the gaps. Closes #91, closes #94. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/backtest.yml | 22 +- CHANGELOG.md | 22 ++ docs/wiki/03-Configuration-and-Tuning.md | 4 +- docs/wiki/04-Testing-and-Contributing.md | 14 +- test_backtest.py | 108 +++++++- tools/backtest.py | 302 ++++++++++++++++++----- 6 files changed, 394 insertions(+), 78 deletions(-) diff --git a/.github/workflows/backtest.yml b/.github/workflows/backtest.yml index 6c684ab..314b2a9 100644 --- a/.github/workflows/backtest.yml +++ b/.github/workflows/backtest.yml @@ -1,6 +1,9 @@ # Manual: walk-forward replay of the scanner over the last N sessions on the # runner's open internet. The Markdown report lands in the job summary. # gh workflow run backtest.yml [-f days=63] [-f horizon=40] [-f min_score=60] +# gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60 +# (the second form replays two years with each scan seeing the nightly's 500-bar +# window and reports the two years side by side: the out-of-sample check). # Also runs on pushes that touch its own files (first registration / smoke test). name: backtest @@ -31,6 +34,18 @@ on: description: "space-separated KEY=VALUE constant overrides, e.g. CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45" required: false default: "" + period: + description: "yfinance history period (2y = what the nightly downloads; 5y for a two-year replay)" + required: false + default: "2y" + bars: + description: "bars of history each scan sees (empty = all; 500 mimics the nightly 2y window in a long replay)" + required: false + default: "" + split: + description: "YYYY-MM-DD: also report the sessions before and from this date as separate windows" + required: false + default: "" push: paths: - tools/backtest.py @@ -59,9 +74,14 @@ jobs: PROFILE: ${{ github.event.inputs.profile || 'spec' }} ABLATE: ${{ github.event.inputs.ablate || 'false' }} OVERRIDES: ${{ github.event.inputs.overrides || '' }} + PERIOD: ${{ github.event.inputs.period || '2y' }} + BARS: ${{ github.event.inputs.bars || '' }} + SPLIT: ${{ github.event.inputs.split || '' }} run: | if [ "$ABLATE" = "true" ]; then MODE=--ablate; else MODE=--grid; fi cmd=(python tools/backtest.py --days "$DAYS" --horizon "$HORIZON" --min-score "$MIN_SCORE" - --profile "$PROFILE" "$MODE" --json backtest.json) + --profile "$PROFILE" --period "$PERIOD" "$MODE" --json backtest.json) + if [ -n "$BARS" ]; then cmd+=(--bars "$BARS"); fi + if [ -n "$SPLIT" ]; then cmd+=(--split "$SPLIT"); fi for kv in $OVERRIDES; do cmd+=(--set "$kv"); done "${cmd[@]}" | tee -a "$GITHUB_STEP_SUMMARY" diff --git a/CHANGELOG.md b/CHANGELOG.md index e628277..42cdc69 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,28 @@ All notable changes to this project are documented here. Format follows ## [Unreleased] +### Added (backtest statistics and out-of-sample split, #91) +- Every backtest summary (overall, per pattern, per score bucket, other + profile) carries median R, standard deviation, total R, the deepest drawdown + of the cumulative R curve (1 R per trade in scan order), a 95 % bootstrap + interval of the mean R and the drawdown exceeded in 5 % of resamples. The + bootstrap resamples scan months, not trades: signals cluster in time (33 in + June 2025, 2 in March 2026), so a per-trade interval would be too narrow. +- `--split YYYY-MM-DD` reports the sessions before and from a date as + separate windows next to the pooled one, so a rule chosen on one window is + judged on the other; `--bars N` limits what each scan sees so a `--period + 5y --days 500 --bars 500` replay reproduces two years of nightly runs. + Workflow inputs `period`, `bars`, `split`. Background: every calibration + to date was judged on the year the values were chosen on; a first two-year + replay found tuned at +0.30 R out of sample against +0.28 in-sample, legacy + at +0.34 against +0.09. + +### Fixed +- Backtest: a next open at or below the stop was scored as a stop with an + undefined R, counted in the hit rate but dropped from mean R (INVH + 2024-10-01, UAL 2025-11-17). It is now `below_stop`: not traded, counted + with the gaps (#94). + ### Changed (calibration, 2026-09-06 replay of the review rules) - Tuned profile: `MIN_REWARD_RISK = 1.0`. Year-long replay (250 sessions, horizon 60, tuned): no minimum 177 signals at +0.28 R, 1.0 -> 158 at diff --git a/docs/wiki/03-Configuration-and-Tuning.md b/docs/wiki/03-Configuration-and-Tuning.md index 9f7d480..e5f88df 100644 --- a/docs/wiki/03-Configuration-and-Tuning.md +++ b/docs/wiki/03-Configuration-and-Tuning.md @@ -173,4 +173,6 @@ For sensitivity, run the negative-control mutations (`test_patterns.py -k violat python -m pytest test_patterns.py -k "violations or controls" -q ``` -For real-data effect, run the backtest (see [Testing and Contributing](04-Testing-and-Contributing.md)): it replays the active profile and the other one on the same year of prices, and re-scores the signals under stop and target variants. Any constant can be overridden for one replay without a code change: `gh workflow run backtest.yml -f overrides="CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45"` (locally `--set KEY=VALUE`). Change one constant, re-run, and only then change the default. +For real-data effect, run the backtest (see [Testing and Contributing](04-Testing-and-Contributing.md)): it replays the active profile and the other one on the same prices, and re-scores the signals under stop and target variants. Any constant can be overridden for one replay without a code change: `gh workflow run backtest.yml -f overrides="CUP_TRIGGER=rim_b WW_TIME_SYM_TOL=0.45"` (locally `--set KEY=VALUE`). Change one constant, re-run, and only then change the default. + +Every calibration above was judged on a single year, the one the candidate values were chosen on. The backtest now reports a 95 % interval of the mean R (resampling scan months, because signals cluster in time) and, with `-f split=YYYY-MM-DD`, the sessions before and from a date as separate windows, so a value chosen on one window can be confirmed on the other before it becomes the default: `gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60`. A first such replay (2026-09-07, local) found the tuned profile at +0.30 R on 167 signals in the year before 2025-09-09, which no decision had used, against +0.28 R on 159 in the tuning year; legacy fell from +0.34 to +0.09 between the same two windows. diff --git a/docs/wiki/04-Testing-and-Contributing.md b/docs/wiki/04-Testing-and-Contributing.md index 9631272..b23988a 100644 --- a/docs/wiki/04-Testing-and-Contributing.md +++ b/docs/wiki/04-Testing-and-Contributing.md @@ -78,6 +78,18 @@ Work on a feature branch and open a PR to `main`; the `tests` workflow must pass ## Measuring signal outcomes -`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open (opens above the row's Max buy are "gapped", not traded), and classifies the outcome within a horizon. Output: overall, per-pattern and per-score-bucket hit rates and mean R, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE, every signal, and (with `--grid`) the same signals re-scored under stop-distance, stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile. Caveats: today's constituents only (survivorship bias), and the last `horizon` sessions are still open. Dispatch with `gh workflow run backtest.yml -f days=63 -f horizon=40`; about 2 minutes for the full index. +`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open, and classifies the outcome within a horizon. Two kinds of open are not traded and counted separately: one above the row's Max buy (`gap`) and one at or below the stop (`below_stop`, where nobody buys and R would be undefined). Output, overall, per pattern and per score bucket: hit rate, mean R, median R, standard deviation, total R, the deepest drawdown of the cumulative R curve (1 R per trade, in scan order), a 95 % bootstrap interval of the mean R and the drawdown exceeded in only 5 % of resamples, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE; then every signal; and (with `--grid`) the same signals re-scored under stop-distance, stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile. + +The bootstrap resamples scan **months**, not trades: signals arrive in clusters (33 in one month and 2 in another over 2024-26), so treating them as independent draws would make the interval far too narrow. An interval is only printed with at least two months and five trades. + +**Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets. + +```bash +gh workflow run backtest.yml -f days=63 -f horizon=40 # about 2 minutes +``` + +```bash +gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60 # two years, about 35 minutes with --grid +``` `tools/evaluate_signals.py` reads every version of `output/signals.json` from git history (the daily scan commits one per run), keeps the first appearance of each `CONFIRMED` signal keyed on `(ticker, pattern, stop)`, fetches the bars that followed, and classifies each as `target` (High reached the target before Low touched the stop), `stop`, `open` (neither within the horizon, marked to the last close) or `no_data`. R multiples are `(exit − entry) / (entry − stop)`. Run it on GitHub (`gh workflow run evaluate-signals.yml -f horizon=60`) because market-data hosts may be blocked locally. diff --git a/test_backtest.py b/test_backtest.py index 7c863d5..4ab6932 100644 --- a/test_backtest.py +++ b/test_backtest.py @@ -1,11 +1,12 @@ #!/usr/bin/env python3 -"""Tests for tools/backtest.py: no look-ahead, first-seen signals, fills and breakdowns (offline).""" +"""Tests for tools/backtest.py: no look-ahead, first-seen signals, fills, statistics and breakdowns (offline).""" from __future__ import annotations import os import sys +import numpy as np import pandas as pd import pytest @@ -40,6 +41,34 @@ def test_walk_forward_marks_gaps_and_no_data(mini_universe): assert [r["outcome"] for r in rows] == ["no_data"] +def test_walk_forward_does_not_trade_an_open_at_or_below_the_stop(mini_universe): + """An open already through the stop is no trade (R would be undefined), counted next to the gaps.""" + cup = mini_universe["CUP"].copy() + (today,) = scan.detect_cup_and_handle(cup, "CUP") # seen at bar -2, filled at bar -1's open + cup.loc[cup.index[-1], ["Open", "Low"]] = [today.stop - 0.01, today.stop - 0.5] + (r,) = bt.walk_forward({"CUP": cup}, days=5, horizon=10) + assert r["outcome"] == "below_stop" and r["r"] is None and r["fill"] == round(today.stop - 0.01, 2) + stats = bt.breakdown([r])["overall"] + assert (stats["below_stop"], stats["gap"], stats["n"], stats["mean_r"]) == (1, 0, 0, None) + md = bt.render([r], bt.report_sections([r], 5, 10), 5, 10) + assert "| all | 1 | 1 | 0 | 0 | 0 | - | - |" in md # "Not traded" counts it + + +def test_walk_forward_bars_limits_what_each_scan_sees(mini_universe, monkeypatch): + seen = [] + real = scan.scan_symbol + + def spy(sym, df, detectors=None): + seen.append(len(df)) + return real(sym, df, detectors) + + monkeypatch.setattr(scan, "scan_symbol", spy) + bt.walk_forward(mini_universe, days=3, horizon=5, bars=120) + assert seen and set(seen) == {120} # every scan saw exactly the last 120 bars + seen.clear() + assert bt.walk_forward(mini_universe, days=3, horizon=5, bars=50) == [] and seen == [] # < 60 bars: skipped + + def _bars(*rows): """rows: (open, high, low, close) per day.""" idx = pd.bdate_range("2026-01-05", periods=len(rows)) @@ -72,6 +101,55 @@ def test_classify_variant_close_vs_intraday_stops(): "outcome": "stop", "bars": 1, "exit": 93.5, "r": -1.3} +# --------------------------------------------------------------------------- # +# Statistics: drawdown, month-block bootstrap, windows +# --------------------------------------------------------------------------- # +def _trade(day, r, ticker="T"): + return {"scan_day": day, "ticker": ticker, "r": r} + + +def test_max_drawdown(): + assert bt.max_drawdown([]) is None + assert bt.max_drawdown([1.0, 1.0]) == 0.0 + assert bt.max_drawdown([-1.0, 2.0]) == -1.0 # the curve starts at 0: a losing first trade counts + assert bt.max_drawdown([2.0, -1.0, -1.0, 0.5]) == -2.0 # peak 2 -> trough 0 + assert bt.max_drawdown([0.5, -1.0, -1.0, 3.0, -0.5]) == -2.0 # 0.5 -> -1.5 is a fall of 2; -0.5 later is less + + +def test_block_bootstrap_resamples_months(): + flat = [_trade(f"2026-{m:02d}-10", 1.0) for m in (1, 1, 2, 2, 3, 3)] + assert bt.block_bootstrap(flat, n_boot=50) == {"ci_low": 1.0, "ci_high": 1.0, "dd_p95": 0.0, "blocks": 3} + # Fewer than two months, or fewer than five trades: an interval would mean nothing. + assert bt.block_bootstrap(flat[:2], n_boot=10)["ci_low"] is None + assert bt.block_bootstrap(flat[:4], n_boot=10) == {"ci_low": None, "ci_high": None, "dd_p95": None, "blocks": 2} + # A mixed sample: the interval brackets the mean, the worst-case drawdown is at least the observed one, + # rows without r are ignored, and the result is reproducible. + rs = (2.0, -1.0, -1.0, 2.0, -1.0, 2.0, -1.0, -1.0) + mixed = [_trade(f"2026-{m:02d}-10", r) for m, r in zip((1, 1, 2, 2, 3, 3, 4, 4), rs)] + [_trade("2026-05-10", None)] + b = bt.block_bootstrap(mixed, n_boot=500, seed=1) + assert b["blocks"] == 4 and b["ci_low"] < sum(rs) / len(rs) < b["ci_high"] + assert b["dd_p95"] <= bt.max_drawdown(rs) <= 0 + assert bt.block_bootstrap(mixed, n_boot=200) == bt.block_bootstrap(mixed, n_boot=200) + + +def test_split_windows_and_report_sections(mini_universe): + rows = bt.walk_forward(mini_universe, days=5, horizon=10) + assert bt.split_windows(rows, None) == [("all sessions", rows)] + split = max(r["scan_day"] for r in rows) + wins = bt.split_windows(rows, split) + assert [w[0] for w in wins] == ["all sessions", f"before {split}", f"from {split}"] + assert len(wins[1][1]) + len(wins[2][1]) == len(rows) and wins[2][1] + assert all(r["scan_day"] < split for r in wins[1][1]) and all(r["scan_day"] >= split for r in wins[2][1]) + sections = bt.report_sections(rows, 5, 10, split=split, data=mini_universe, do_grid=True) + assert [s["label"] for s in sections] == [w[0] for w in wins] + assert sections[0]["stats"]["overall"]["signals"] == len(rows) and len(sections[0]["grid"]) == len(bt.GRID) + assert sections[2]["stats"]["overall"]["signals"] == len(wins[2][1]) + md = bt.render(rows, sections, 5, 10) + assert "## All sessions" in md and f"## Before {split}" in md and f"## From {split}" in md + assert md.count("### Stop / target variants") == 3 and "judged on the other" in md + assert "## Signals" in md + + def test_grid_rescores_the_same_signals(mini_universe): rows = bt.walk_forward(mini_universe, days=5, horizon=10) g = bt.grid(rows, mini_universe, 10) @@ -79,8 +157,8 @@ def test_grid_rescores_the_same_signals(mini_universe): assert {(x["stop_extra_atr"], x["stop_basis"], x["target_mode"]) for x in g} == set(bt.GRID) traded = sum(1 for r in rows if r["outcome"] in ("target", "stop", "open")) assert all(x["n"] == traded for x in g) - md = bt.render(rows, bt.breakdown(rows), 5, 10, g) - assert "## Stop / target variants" in md and "| 0.75 | close | breakout |" in md + md = bt.render(rows, bt.report_sections(rows, 5, 10, data=mini_universe, do_grid=True), 5, 10) + assert "### Stop / target variants" in md and "| 0.75 | close | breakout |" in md def test_breakout_level_target_applies_to_cups_only(mini_universe): @@ -126,8 +204,11 @@ def test_profile_pass_replays_the_other_rule_set_and_restores_the_active_one(min legacy_targets = {r["ticker"]: r["target"] for r in other["rows"] if r["pattern"] == "Cup & Handle"} spec_targets = {r["ticker"]: r["target"] for r in rows if r["pattern"] == "Cup & Handle"} assert spec_targets and legacy_targets and spec_targets != legacy_targets # left-rim vs right-rim measure - md = bt.render(rows, bt.breakdown(rows), 5, 10, None, other) - assert "## Rule profile comparison: spec (above) vs legacy (below)" in md and "| legacy: all |" in md + sections = bt.report_sections(rows, 5, 10, other_rows=other["rows"], other_name="legacy") + assert sections[0]["other"]["stats"]["overall"]["signals"] == other["stats"]["overall"]["signals"] + md = bt.render(rows, sections, 5, 10) + assert "### Rule profile comparison: spec (above) vs legacy (below), all sessions" in md + assert "| legacy: all |" in md def test_breakdown_and_render(): @@ -140,13 +221,18 @@ def row(pattern, score, outcome, r): rows = [row("Cup & Handle", 65, "target", 2.0), row("Cup & Handle", 72, "stop", -1.0), row("Bullish Wolfe Wave", 85, "open", 0.4), row("Cup & Handle", 91, "gap", None)] stats = bt.breakdown(rows) - assert stats["overall"]["signals"] == 4 and stats["overall"]["gap"] == 1 - assert stats["overall"]["target"] == 1 and stats["overall"]["stop"] == 1 and stats["overall"]["open"] == 1 - assert stats["overall"]["hit_rate"] == 0.5 and stats["overall"]["mean_r"] == round((2.0 - 1.0 + 0.4) / 3, 3) + o = stats["overall"] + assert (o["signals"], o["gap"], o["below_stop"]) == (4, 1, 0) + assert (o["target"], o["stop"], o["open"]) == (1, 1, 1) + assert o["hit_rate"] == 0.5 and o["mean_r"] == round((2.0 - 1.0 + 0.4) / 3, 3) + assert o["median_r"] == 0.4 and o["total_r"] == 1.4 + assert o["std_r"] == round(float(np.std([2.0, -1.0, 0.4], ddof=1)), 3) + assert o["max_dd"] == -1.0 # 2 -> 1 in scan order + assert o["ci_low"] is None and o["dd_p95"] is None and o["blocks"] == 1 # one month: no interval assert set(stats["by_score"]) == {"60-69", "70-79", "80-89", "90-100"} assert stats["by_pattern"]["Bullish Wolfe Wave"]["hit_rate"] is None # nothing resolved - assert stats["overall"]["success5"] == 0.5 and stats["overall"]["mfe"] == 0.08 - md = bt.render(rows, stats, 63, 40) - assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | 50% | +8.0% | -2.0% |" in md + assert o["success5"] == 0.5 and o["mfe"] == 0.08 + md = bt.render(rows, bt.report_sections(rows, 63, 40), 63, 40) + assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | +0.40 | +1.40 | - | -1.00 | - | 50% | +8.0% | -2.0% |" in md assert "| 2026-01-05 | T | Cup & Handle | 91 | 100.0 | 100.5 | 95.0 | 110.0 | gap | 0 | - |" in md assert isinstance(pd.DataFrame(rows), pd.DataFrame) diff --git a/tools/backtest.py b/tools/backtest.py index fd49b43..8823a9d 100644 --- a/tools/backtest.py +++ b/tools/backtest.py @@ -12,6 +12,7 @@ runner, or locally where Yahoo is reachable): python tools/backtest.py [--days 63] [--horizon 40] [--tickers A,B] [--json out.json] + [--period 5y --days 500 --bars 500 --split 2025-09-09] Method: @@ -21,15 +22,31 @@ * The fill is the **next session's open** (the e-mail arrives before the US open). An open above the row's ``max_buy`` (trigger + ``MAX_RUNAWAY``, or where the risk at the fill reaches ``MAX_BUY_RISK_MULT`` x the planned risk) - is a ``gap``: no trade, counted separately. + is a ``gap``; an open at or below the stop is ``below_stop`` (nobody buys an + open that is already through the stop, and R would be undefined). Neither + is traded; both are counted separately. * Outcomes use ``evaluate_signals.classify``: ``target`` / ``stop`` / ``open`` within ``horizon`` bars after the fill bar, gaps filled at the open, R multiple = (exit - fill) / (fill - stop). * No look-ahead: each scan only sees bars <= D; asserted per signal. + ``--bars N`` further limits what each scan sees to its last N bars, so a + long download replays what the nightly job sees (its 2y download is about + 500 bars). * Per signal it also records the maximum favourable / adverse excursion over the horizon and the chart-book style ``success5``: did the high reach +5 % above the fill before any *close* below the stop (no target, no intraday stop), which is what published "pattern success rates" measure. +* Every summary carries, next to hit rate and mean R: median R, standard + deviation, total R, the deepest drawdown of the cumulative R curve (1 R per + trade, in scan order), a 95 % bootstrap interval of the mean R and the + drawdown exceeded in only 5 % of resamples. The bootstrap resamples scan + *months*, not trades: signals cluster in time (33 in one month, 2 in + another, over 2024-26), so resampling trades as if independent would + understate the uncertainty. +* ``--split YYYY-MM-DD`` reports the sessions before and from that date as + separate windows next to the pooled one, so a rule chosen on one window is + judged on the other -- the out-of-sample check a calibration decision + should pass (docs/wiki/03). * ``--grid`` re-scores the same signals under stop / target variants: extra ATR below the reported stop (0 / 0.25 / 0.75, i.e. ~0.25 / 0.5 / 1.0 ATR under the structural low), intraday vs close-based stops, and three target @@ -61,6 +78,7 @@ import time from typing import Any, Dict, List, Optional, Sequence, Tuple +import numpy as np import pandas as pd sys.path.insert(0, os.path.join(os.path.dirname(os.path.abspath(__file__)), "..")) @@ -74,6 +92,8 @@ TARGET_MODES = ("full", "half", "breakout") GRID = [(extra, basis, mode) for extra in (0.0, 0.25, 0.75) for basis in ("intraday", "close") for mode in TARGET_MODES] +TRADED = ("target", "stop", "open") # outcomes with a position; "gap" / "below_stop" / "no_data" have none +N_BOOT = 2000 # month-block bootstrap resamples per summary def variant_target(row: dict, mode: str) -> Optional[float]: @@ -145,10 +165,81 @@ def classify_variant(fill: float, stop: float, target: Optional[float], bars: pd return {"outcome": "open", "bars": len(w), "exit": last, "r": (last - fill) / risk if risk > 0 else None} +# --------------------------------------------------------------------------- # +# Statistics: drawdown, month-block bootstrap, windows +# --------------------------------------------------------------------------- # +def max_drawdown(rs: Sequence[float]) -> Optional[float]: + """Deepest fall of the cumulative R curve from its running peak, 1 R risked per trade, in the given order. + + The curve starts at 0, so a losing first trade already counts as a drawdown. + + :returns: A number <= 0 rounded to 3 dp, or ``None`` without trades. + """ + if len(rs) == 0: + return None + cum = np.cumsum(np.asarray(rs, dtype=float)) + peak = np.maximum(np.maximum.accumulate(cum), 0.0) + return round(float((cum - peak).min()), 3) + + +def block_bootstrap(traded: Sequence[dict], n_boot: int = N_BOOT, seed: int = 0) -> Dict[str, Any]: + """Bootstrap the mean R and the drawdown by resampling scan *months* with replacement. + + Why months: signals arrive in clusters (a correction ends and twenty bases + break out in the same fortnight), so their outcomes are not independent + draws and a per-trade bootstrap is too narrow. Each resample picks as many + months as the window has, concatenates their trades in the picked order, + and records the mean R and the max drawdown of that sequence. + + :param traded: Rows with ``scan_day`` and ``r`` (rows without ``r`` are + ignored), in scan order. + :param n_boot: Resamples; ``seed`` makes the result reproducible. + :returns: ``ci_low`` / ``ci_high`` (95 % percentile interval of the mean R), + ``dd_p95`` (the drawdown exceeded in only 5 % of resamples, <= 0) and + ``blocks`` (months with trades); the first three are ``None`` with fewer + than 2 months or 5 trades, where an interval would mean nothing. + + Complexity: O(n_boot * trades). + """ + months: Dict[str, List[float]] = {} + for r in traded: + if r.get("r") is not None: + months.setdefault(str(r["scan_day"])[:7], []).append(float(r["r"])) + blocks = [np.asarray(months[m]) for m in sorted(months)] + if len(blocks) < 2 or sum(len(b) for b in blocks) < 5: + return {"ci_low": None, "ci_high": None, "dd_p95": None, "blocks": len(blocks)} + rng = np.random.default_rng(seed) + means = np.empty(n_boot) + dds = np.empty(n_boot) + for i in range(n_boot): + pick = rng.integers(0, len(blocks), len(blocks)) + seq = np.concatenate([blocks[j] for j in pick]) + means[i] = seq.mean() + dds[i] = max_drawdown(seq) + lo, hi = np.percentile(means, [2.5, 97.5]) + return {"ci_low": round(float(lo), 3), "ci_high": round(float(hi), 3), + "dd_p95": round(float(np.percentile(dds, 5)), 3), "blocks": len(blocks)} + + +def split_windows(rows: Sequence[dict], split: Optional[str]) -> List[Tuple[str, List[dict]]]: + """``[(label, rows)]``: the pooled rows and, with ``split``, the rows before and from that ISO date. + + ``scan_day`` is an ISO string, so the comparison is lexical. + """ + out = [("all sessions", list(rows))] + if split: + out.append((f"before {split}", [r for r in rows if r["scan_day"] < split])) + out.append((f"from {split}", [r for r in rows if r["scan_day"] >= split])) + return out + + +# --------------------------------------------------------------------------- # +# Replay +# --------------------------------------------------------------------------- # def grid(rows: Sequence[dict], data: Dict[str, pd.DataFrame], horizon: int) -> List[dict]: """Re-score the traded signals under every stop / target variant in ``GRID``.""" out = [] - traded = [r for r in rows if r["outcome"] in ("target", "stop", "open")] + traded = [r for r in rows if r["outcome"] in TRADED] for extra, basis, mode in GRID: scored = [] for r in traded: @@ -161,20 +252,20 @@ def grid(rows: Sequence[dict], data: Dict[str, pd.DataFrame], horizon: int) -> L return out -def ablation(data: Dict[str, pd.DataFrame], days: int, horizon: int) -> List[dict]: +def ablation(data: Dict[str, pd.DataFrame], days: int, horizon: int, bars: Optional[int] = None) -> List[dict]: """Leave-one-rule-out over the spec profile. Replays the active (spec) profile, then once per key in ``RULE_PROFILES["legacy"]`` with only that key set to its legacy value. :returns: one summary row per replay: ``{"rule", "value", "delta", **breakdown(...)["overall"]}``. """ - base = breakdown(walk_forward(data, days, horizon))["overall"] + base = breakdown(walk_forward(data, days, horizon, bars=bars))["overall"] rows = [{"rule": "spec (all rules)", "value": "-", "delta": 0, **base}] for key, legacy_value in scan.RULE_PROFILES["legacy"].items(): saved = getattr(scan, key) setattr(scan, key, legacy_value) try: - s = breakdown(walk_forward(data, days, horizon))["overall"] + s = breakdown(walk_forward(data, days, horizon, bars=bars))["overall"] finally: setattr(scan, key, saved) rows.append({"rule": key, "value": str(legacy_value), "delta": s["signals"] - base["signals"], **s}) @@ -191,8 +282,7 @@ def render_ablation(rows: Sequence[dict], days: int, horizon: int) -> str: "|---|---|---|---|---|---|---|---|---|---|"] for r in sorted(rows, key=lambda r: -abs(r["delta"])): lines.append(f"| {r['rule']} | {r['value']} | {r['signals']} | {r['delta']:+d} | {r['target']} | {r['stop']} | " - f"{r['open']} | {_pct(r['hit_rate'])} | {'-' if r['mean_r'] is None else f'{r['mean_r']:+.2f}'} | " - f"{_pct(r['success5'])} |") + f"{r['open']} | {_pct(r['hit_rate'])} | {_r(r['mean_r'])} | {_pct(r['success5'])} |") return "\n".join(lines) @@ -213,36 +303,45 @@ def apply_override(item: str) -> Tuple[str, Any]: return key, value -def profile_pass(data: Dict[str, pd.DataFrame], days: int, horizon: int, name: str) -> dict: +def profile_pass(data: Dict[str, pd.DataFrame], days: int, horizon: int, name: str, + bars: Optional[int] = None) -> dict: """Second full walk-forward under rule profile ``name`` (the active profile is restored). :returns: ``{"profile", "stats", "rows"}`` with the same ``breakdown`` slices as the main report. """ with rule_profile(name): - rows = walk_forward(data, days, horizon) + rows = walk_forward(data, days, horizon, bars=bars) return {"profile": name, "stats": breakdown(rows), "rows": rows} def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, - detectors: Optional[Sequence] = None) -> List[dict]: + detectors: Optional[Sequence] = None, bars: Optional[int] = None) -> List[dict]: """Replay the scanner over the last ``days`` sessions and score each first-seen signal. :param data: ``{symbol: OHLCV frame}`` as returned by ``download_history``. :param days: Number of most recent sessions to scan (each as if it were "today"). :param horizon: Bars after the fill before an unresolved signal counts as ``open``. :param detectors: Subset of detectors to run (default: all). + :param bars: Bars of history each scan sees (default: everything up to the + day). 500 mimics the nightly job's 2y download in a longer replay. :returns: One dict per first-seen CONFIRMED signal with the signal fields plus ``fill``, ``outcome``, ``bars``, ``exit``, ``r`` (and for cups the parsed ``cup_bottom`` / ``cup_trigger`` for the breakout-level target variant). + ``outcome`` is ``target`` / ``stop`` / ``open`` for traded rows, ``gap`` + (open above Max buy) or ``below_stop`` (open at or below the stop) for + rows not traded, ``no_data`` without a bar after the scan day. Complexity: O(days * symbols * detector cost); ~2 ms per symbol-day. """ sessions = sorted({d for df in data.values() for d in df.index}) scan_days = sessions[-days:] if days < len(sessions) else sessions seen: Dict[tuple, dict] = {} + not_traded = dict(exit=None, r=None, mfe=None, mae=None, success5=None) for d in scan_days: for sym, df in data.items(): hist = df[df.index <= d] + if bars is not None: + hist = hist.iloc[-bars:] if len(hist) < 60 or hist.index[-1] != d: continue # symbol had no bar on d (halted, listed later) for s in scan.scan_symbol(sym, hist, detectors): @@ -260,14 +359,15 @@ def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, if m: row["cup_bottom"], row["cup_trigger"] = float(m[1]), float(m[2]) if after.empty: - row.update(fill=None, outcome="no_data", bars=0, exit=None, r=None, - mfe=None, mae=None, success5=None) + row.update(fill=None, outcome="no_data", bars=0, **not_traded) else: fill = float(after["Open"].iloc[0]) max_buy = s.max_buy if s.max_buy is not None else s.entry * (1 + scan.MAX_RUNAWAY) if fill > max_buy: - row.update(fill=round(fill, 2), outcome="gap", bars=0, exit=None, r=None, - mfe=None, mae=None, success5=None) + row.update(fill=round(fill, 2), outcome="gap", bars=0, **not_traded) + elif fill <= s.stop: + # The open is already through the stop: no trade, and R would be undefined. + row.update(fill=round(fill, 2), outcome="below_stop", bars=0, **not_traded) else: res = ev.classify(fill, s.stop, s.target, after, horizon) row.update(fill=round(fill, 2), **res, **excursions(fill, s.stop, after, horizon)) @@ -276,12 +376,28 @@ def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, def breakdown(rows: Sequence[dict]) -> dict: - """Summary overall, per pattern and per score bucket (gaps and no_data excluded from rates).""" + """Summary overall, per pattern and per score bucket (rows not traded are excluded from the rates). + + Each summary carries the ``evaluate_signals.summarise`` counts plus + ``signals``, ``gap``, ``below_stop``, ``median_r``, ``std_r``, ``total_r``, + ``max_dd`` (cumulative R curve in scan order), the month-block bootstrap + ``ci_low`` / ``ci_high`` / ``dd_p95`` / ``blocks`` (see + :func:`block_bootstrap`), ``success5``, ``mfe`` and ``mae``. + """ def summ(sub): - traded = [r for r in sub if r["outcome"] in ("target", "stop", "open")] + traded = sorted((r for r in sub if r["outcome"] in TRADED), key=lambda r: (r["scan_day"], r["ticker"])) + rs = [float(r["r"]) for r in traded if r["r"] is not None] decided = [r["success5"] for r in traded if r.get("success5") is not None] - return {**ev.summarise(traded), "gap": sum(1 for r in sub if r["outcome"] == "gap"), - "signals": len(sub), + boot = block_bootstrap(traded) + return {**ev.summarise(traded), "signals": len(sub), + "gap": sum(1 for r in sub if r["outcome"] == "gap"), + "below_stop": sum(1 for r in sub if r["outcome"] == "below_stop"), + "median_r": round(float(np.median(rs)), 3) if rs else None, + "std_r": round(float(np.std(rs, ddof=1)), 3) if len(rs) > 1 else None, + "total_r": round(float(np.sum(rs)), 2) if rs else None, + "max_dd": max_drawdown(rs), + "ci_low": boot["ci_low"], "ci_high": boot["ci_high"], + "dd_p95": boot["dd_p95"], "blocks": boot["blocks"], "success5": round(sum(decided) / len(decided), 3) if decided else None, "mfe": round(sum(r["mfe"] for r in traded) / len(traded), 4) if traded else None, "mae": round(sum(r["mae"] for r in traded) / len(traded), 4) if traded else None} @@ -295,48 +411,95 @@ def summ(sub): return out +def report_sections(rows: Sequence[dict], days: int, horizon: int, split: Optional[str] = None, + data: Optional[Dict[str, pd.DataFrame]] = None, do_grid: bool = False, + other_rows: Optional[Sequence[dict]] = None, other_name: Optional[str] = None) -> List[dict]: + """One section per window (see :func:`split_windows`): its rows, ``breakdown`` + stats, the stop / target ``grid`` (with ``do_grid`` and ``data``) and the other + profile's stats over the same window (with ``other_rows``).""" + windows = split_windows(rows, split) + others = split_windows(other_rows, split) if other_rows is not None else [(None, None)] * len(windows) + sections = [] + for (label, rws), (_, o_rws) in zip(windows, others): + sec: Dict[str, Any] = {"label": label, "rows": rws, "stats": breakdown(rws), "grid": None, "other": None} + if do_grid and data is not None: + sec["grid"] = grid(rws, data, horizon) + if o_rws is not None: + sec["other"] = {"profile": other_name, "stats": breakdown(o_rws)} + sections.append(sec) + return sections + + +# --------------------------------------------------------------------------- # +# Rendering +# --------------------------------------------------------------------------- # def _pct(x, signed=False): return "-" if x is None else (f"{x:+.1%}" if signed else f"{x:.0%}") -def render(rows: Sequence[dict], stats: dict, days: int, horizon: int, - grid_rows: Optional[Sequence[dict]] = None, other: Optional[dict] = None) -> str: - """Markdown report (readable as a GitHub step summary).""" - def line(name, s): - return (f"| {name} | {s['signals']} | {s['gap']} | {s['target']} | {s['stop']} | {s['open']} | " - f"{_pct(s['hit_rate'])} | {'-' if s['mean_r'] is None else f'{s['mean_r']:+.2f}'} | " - f"{_pct(s['success5'])} | {_pct(s['mfe'], True)} | {_pct(s['mae'], True)} |") - hdr = ("| Slice | Signals | Gapped | Target | Stop | Open | Hit rate | Mean R | +5% first | MFE | MAE |\n" - "|---|---|---|---|---|---|---|---|---|---|---|") +def _r(x): + return "-" if x is None else f"{x:+.2f}" + + +def _ci(s: dict) -> str: + return "-" if s.get("ci_low") is None else f"[{s['ci_low']:+.2f}, {s['ci_high']:+.2f}]" + + +SUMMARY_HEADER = ("| Slice | Signals | Not traded | Target | Stop | Open | Hit rate | Mean R | Median R | Total R | " + "95% CI | Max DD | DD p95 | +5% first | MFE | MAE |\n" + "|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|") + + +def _line(name: str, s: dict) -> str: + return (f"| {name} | {s['signals']} | {s['gap'] + s['below_stop']} | {s['target']} | {s['stop']} | {s['open']} | " + f"{_pct(s['hit_rate'])} | {_r(s['mean_r'])} | {_r(s['median_r'])} | {_r(s['total_r'])} | {_ci(s)} | " + f"{_r(s['max_dd'])} | {_r(s['dd_p95'])} | {_pct(s['success5'])} | {_pct(s['mfe'], True)} | " + f"{_pct(s['mae'], True)} |") + + +def _summary_table(stats: dict, prefix: str = "") -> List[str]: + lines = [SUMMARY_HEADER, _line(f"{prefix}all", stats["overall"])] + lines += [_line(f"{prefix}{p}", s) for p, s in stats["by_pattern"].items()] + lines += [_line(f"{prefix}score {b}", s) for b, s in stats["by_score"].items()] + return lines + + +def render(rows: Sequence[dict], sections: Sequence[dict], days: int, horizon: int) -> str: + """Markdown report (readable as a GitHub step summary): one block per window, then every signal.""" lines = [f"# Walk-forward backtest: last {days} sessions, horizon {horizon} bars " f"(rule profile: {scan.ACTIVE_PROFILE})", "", "Hit rate = target / (target + stop). Mean R over traded signals (open ones marked to the last close). " - "Fill = next session's open; opens above the row's Max buy are gapped (not traded). " + "Fill = next session's open; not traded = the open was above the row's Max buy (gap) or at or below " + "the stop. 95% CI = bootstrap interval of the mean R that resamples scan months, not trades (signals " + "cluster in time). Max DD = deepest fall of the cumulative R curve, 1 R per trade in scan order; " + "DD p95 = the drawdown exceeded in 5 % of the month resamples. " "+5% first = share of signals whose high reached +5 % above the fill before any close below the stop " "(the chart-book success definition). MFE / MAE = mean best / worst excursion from the fill " - "within the horizon.", - "", hdr, line("all", stats["overall"])] - lines += [line(p, s) for p, s in stats["by_pattern"].items()] - lines += [line(f"score {b}", s) for b, s in stats["by_score"].items()] - if grid_rows: - lines += ["", "## Stop / target variants (same signals)", "", - "Extra ATR = distance added below the reported stop (which already sits 0.25 ATR under the " - "structural low). Target: full = reported measured move; half = halfway to it; breakout = " - "cups measured from the bottom to the handle breakout level (Investopedia), others unchanged.", "", - "| Extra ATR | Stop basis | Target | Target | Stop | Open | Hit rate | Mean R |", - "|---|---|---|---|---|---|---|---|"] - for g in grid_rows: - lines.append(f"| {g['stop_extra_atr']} | {g['stop_basis']} | {g['target_mode']} | " - f"{g['target']} | {g['stop']} | {g['open']} | {_pct(g['hit_rate'])} | " - f"{'-' if g['mean_r'] is None else f'{g['mean_r']:+.2f}'} |") - if other: - o = other["stats"] - lines += ["", f"## Rule profile comparison: {scan.ACTIVE_PROFILE} (above) vs {other['profile']} (below)", "", - "A second full walk-forward on the same data under the other rule profile " - "(see scan.RULE_PROFILES).", "", hdr, line(f"{other['profile']}: all", o["overall"])] - lines += [line(f"{other['profile']}: {p}", s) for p, s in o["by_pattern"].items()] - lines += [line(f"{other['profile']}: score {b}", s) for b, s in o["by_score"].items()] - lines += ["", "| Day | Ticker | Pattern | Score | Entry | Fill | Stop | Target | Outcome | Bars | R |", + "within the horizon."] + if len(sections) > 1: + lines += ["", "The windows below report the sessions before and from the split date separately: a rule " + "chosen on one window is judged on the other."] + for sec in sections: + lines += ["", f"## {sec['label'].capitalize()}", ""] + _summary_table(sec["stats"]) + if sec["grid"]: + lines += ["", f"### Stop / target variants (same signals, {sec['label']})", "", + "Extra ATR = distance added below the reported stop (which already sits 0.25 ATR under the " + "structural low). Target: full = reported measured move; half = halfway to it; breakout = " + "cups measured from the bottom to the handle breakout level (Investopedia), others " + "unchanged.", "", + "| Extra ATR | Stop basis | Target | Target | Stop | Open | Hit rate | Mean R |", + "|---|---|---|---|---|---|---|---|"] + for g in sec["grid"]: + lines.append(f"| {g['stop_extra_atr']} | {g['stop_basis']} | {g['target_mode']} | " + f"{g['target']} | {g['stop']} | {g['open']} | {_pct(g['hit_rate'])} | {_r(g['mean_r'])} |") + if sec["other"]: + o = sec["other"] + lines += ["", f"### Rule profile comparison: {scan.ACTIVE_PROFILE} (above) vs {o['profile']} (below), " + f"{sec['label']}", "", + "A second full walk-forward on the same data under the other rule profile " + "(see scan.RULE_PROFILES).", ""] + _summary_table(o["stats"], prefix=f"{o['profile']}: ") + lines += ["", "## Signals", "", + "| Day | Ticker | Pattern | Score | Entry | Fill | Stop | Target | Outcome | Bars | R |", "|---|---|---|---|---|---|---|---|---|---|---|"] for r in sorted(rows, key=lambda r: (r["scan_day"], r["ticker"])): lines.append(f"| {r['scan_day']} | {r['ticker']} | {r['pattern']} | {r['score']} | {r['entry']} | " @@ -352,7 +515,12 @@ def main(argv: Optional[Sequence[str]] = None) -> int: ap.add_argument("--horizon", type=int, default=40, help="bars after the fill before a signal is 'open'") ap.add_argument("--tickers", help="comma separated symbols (default: full S&P 500)") ap.add_argument("--csv", help="local constituents CSV") - ap.add_argument("--period", default="2y") + ap.add_argument("--period", default="2y", help="yfinance history period (default 2y, what the nightly downloads)") + ap.add_argument("--bars", type=int, default=None, + help="bars of history each scan sees (default: all); 500 mimics the nightly 2y window " + "in a long replay") + ap.add_argument("--split", metavar="YYYY-MM-DD", + help="also report the sessions before and from this date as separate windows (out-of-sample check)") ap.add_argument("--min-score", type=int, default=None, help="override MIN_SCORE for the replay") ap.add_argument("--json", help="also write rows and stats here") ap.add_argument("--grid", action="store_true", @@ -379,32 +547,38 @@ def main(argv: Optional[Sequence[str]] = None) -> int: if not data: print("no price data") return 2 - log.info("downloaded %d symbols in %.0fs; replaying %d sessions", len(data), time.time() - t0, args.days) + log.info("downloaded %d symbols in %.0fs; replaying %d sessions%s", len(data), time.time() - t0, args.days, + f", each scan sees {args.bars} bars" if args.bars else "") if args.ablate: scan.apply_profile("spec") - table = ablation(data, args.days, args.horizon) + table = ablation(data, args.days, args.horizon, args.bars) log.info("ablation done in %.0fs: %d replays", time.time() - t0, len(table)) print(render_ablation(table, args.days, args.horizon)) if args.json: with open(args.json, "w", encoding="utf-8") as fh: - json.dump({"days": args.days, "horizon": args.horizon, "ablation": table}, fh, indent=2) + json.dump({"days": args.days, "horizon": args.horizon, "bars": args.bars, "ablation": table}, + fh, indent=2) return 0 - rows = walk_forward(data, args.days, args.horizon) - stats = breakdown(rows) + rows = walk_forward(data, args.days, args.horizon, bars=args.bars) log.info("replay done in %.0fs: %d first-seen confirmed signals", time.time() - t0, len(rows)) - grid_rows = other = None + other_rows = other_name = None if args.grid: - grid_rows = grid(rows, data, args.horizon) other_name = "legacy" if scan.ACTIVE_PROFILE == "spec" else "spec" - other = profile_pass(data, args.days, args.horizon, other_name) + other_rows = profile_pass(data, args.days, args.horizon, other_name, args.bars)["rows"] log.info("%s-profile pass done in %.0fs: %d first-seen confirmed signals", - other_name, time.time() - t0, len(other["rows"])) - print(render(rows, stats, args.days, args.horizon, grid_rows, other)) + other_name, time.time() - t0, len(other_rows)) + sections = report_sections(rows, args.days, args.horizon, args.split, data, args.grid, other_rows, other_name) + print(render(rows, sections, args.days, args.horizon)) if args.json: + pooled = sections[0] with open(args.json, "w", encoding="utf-8") as fh: json.dump({"days": args.days, "horizon": args.horizon, "profile": scan.ACTIVE_PROFILE, - "min_score": scan.MIN_SCORE, "stats": stats, "grid": grid_rows, - "other_profile": other, "rows": rows}, fh, indent=2) + "min_score": scan.MIN_SCORE, "bars": args.bars, "split": args.split, + "stats": pooled["stats"], "grid": pooled["grid"], + "other_profile": ({"profile": other_name, "stats": pooled["other"]["stats"], "rows": other_rows} + if other_rows is not None else None), + "windows": [{k: s[k] for k in ("label", "stats", "grid", "other")} for s in sections[1:]], + "rows": rows}, fh, indent=2) return 0 From e1ff498f9e02f28c832f6bdcf0221c5c74464eb4 Mon Sep 17 00:00:00 2001 From: Yaniv Bernhard Date: Mon, 7 Sep 2026 22:23:38 +0300 Subject: [PATCH 2/2] ops: backtest job timeout 180 minutes; state the real runner timings in the wiki The first CI run of this branch took 11 minutes for 63 sessions with --grid, in line with every earlier push run (6 to 12 minutes), not the "about 2 minutes" the wiki claimed. At about 5 s per session per profile a two-year --grid replay needs roughly 90 minutes, the previous job timeout. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/backtest.yml | 4 +++- docs/wiki/04-Testing-and-Contributing.md | 6 ++++-- 2 files changed, 7 insertions(+), 3 deletions(-) diff --git a/.github/workflows/backtest.yml b/.github/workflows/backtest.yml index 314b2a9..72502d6 100644 --- a/.github/workflows/backtest.yml +++ b/.github/workflows/backtest.yml @@ -57,7 +57,9 @@ permissions: jobs: backtest: runs-on: ubuntu-latest - timeout-minutes: 90 + # A runner replays about 5 s per session per profile over the full index; --grid runs two + # profiles, so 63 sessions take about 10 minutes and a two-year (500-session) replay about 90. + timeout-minutes: 180 steps: - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 diff --git a/docs/wiki/04-Testing-and-Contributing.md b/docs/wiki/04-Testing-and-Contributing.md index b23988a..69ea51e 100644 --- a/docs/wiki/04-Testing-and-Contributing.md +++ b/docs/wiki/04-Testing-and-Contributing.md @@ -84,12 +84,14 @@ The bootstrap resamples scan **months**, not trades: signals arrive in clusters **Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets. +A runner replays about 5 seconds per session per profile over the full index, and the workflow always runs `--grid`, which adds the other profile's pass: + ```bash -gh workflow run backtest.yml -f days=63 -f horizon=40 # about 2 minutes +gh workflow run backtest.yml -f days=63 -f horizon=40 # about 10 minutes ``` ```bash -gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60 # two years, about 35 minutes with --grid +gh workflow run backtest.yml -f period=5y -f days=500 -f bars=500 -f split=2025-09-09 -f horizon=60 # two years, about 90 minutes ``` `tools/evaluate_signals.py` reads every version of `output/signals.json` from git history (the daily scan commits one per run), keeps the first appearance of each `CONFIRMED` signal keyed on `(ticker, pattern, stop)`, fetches the bars that followed, and classifies each as `target` (High reached the target before Low touched the stop), `stop`, `open` (neither within the horizon, marked to the last close) or `no_data`. R multiples are `(exit − entry) / (entry − stop)`. Run it on GitHub (`gh workflow run evaluate-signals.yml -f horizon=60`) because market-data hosts may be blocked locally.