Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/backtest.yml
Original file line number Diff line number Diff line change
Expand Up @@ -107,3 +107,9 @@ jobs:
if [ "$ASOF" = "true" ]; then cmd+=(--constituents-asof); fi
for kv in $OVERRIDES; do cmd+=(--set "$kv"); done
"${cmd[@]}" | tee -a "$GITHUB_STEP_SUMMARY"
- name: Keep the JSON (rows, later listings, tables) so several runs can be pooled exactly
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: backtest-${{ github.run_id }}
path: backtest.json
if-no-files-found: ignore
16 changes: 16 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,22 @@ All notable changes to this project are documented here. Format follows

## [Unreleased]

### Added (late-entry replay, #115)
- The backtest records every later listing of an already-seen signal (the
same structure still `CONFIRMED` on a following scan day, which the report
shows again), filled at the next open under the same rules, and reports a
**Late entry** table per window: per day on the list (1 = the first
report), listings, traded, hit rate, mean and median R with the month-block
interval, and the paired comparison with the same signals bought on day 1
(pairs, both means, the mean difference and its interval). `walk_forward`
takes `repeats=`; first-seen rows carry `first_day` and `listed_day`; the
JSON gains `late` and `repeats`; each `backtest` run keeps its JSON as a
workflow artifact (`backtest-<run id>`) so several runs pool exactly.
Replayed over the eleven yearly point-in-time runs (tuning page): a row on
its second day is as good as new; from the third day the same signal pays
about 0.1 R less than on day 1, in ten of eleven years; from day 7 no edge
remains, for a Wolfe from day 6. No rule changed.

### Added (report: the F&G column)
- Every report row carries `F&G`, the stock's own fear-and-greed reading at
the last close, 0 to 100, one value per symbol; `fear_greed` in
Expand Down
16 changes: 16 additions & 0 deletions docs/wiki/03-Configuration-and-Tuning.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,22 @@ A confirmed breakout is almost never fearful by construction (one row below 20 i

The one candidate rule this produces is the Bollinger stretch: a breakout bar that closes above its upper band. Leaving those 453 signals out would keep 1179 signals at +0.30 R against 1632 at +0.22, at a cost of 16 R of the ten-year total of 367. Under the protocol it remains a hypothesis: the gate has to be replayed as a rule on its own, so that its effect on the drawdown and per pattern is measured (#113). The reading itself is on every report row as the `F&G` column, one value per symbol at the last close, information only.

### Late entry, tested (2026-09-09)

The owner's question on HAL, listed as a confirmed inverse head and shoulders in every report from 2026-09-04: the rules keep an H&S or Wolfe breakout listed for up to 8 bars and re-quote the entry at each close, and #98 settled that the age at *first* report is no reason to cut that window. But every replay fills a signal once, at the open after its first report, and nothing had measured what a reader gets who buys a row on its second, third or eighth day on the list, which is what the report invites by showing the row again. `tools/backtest.py` now records every later CONFIRMED listing of an already-seen signal (`listed_day` = sessions since the first report) and fills it at the next open under the same rules, and its **Late entry** table pairs each later listing with the same signal's first report (#115; the eleven yearly point-in-time runs 34323013559 to 34323039560, tuned, `bars=500`, `horizon=60`: 1692 first-seen signals, 4915 later listings).

| Day on the list | Listed | Traded | Hit rate | Mean R | 95 % CI | Pairs | Same signals on day 1 | On day N | Difference | 95 % CI |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1692 | 1632 | 35 % | +0.22 | [+0.09, +0.36] | | | | | |
| 2 | 1140 | 1111 | 35 % | +0.28 | [+0.13, +0.45] | 1090 | +0.28 | +0.27 | −0.01 | [−0.07, +0.07] |
| 3 | 995 | 952 | 35 % | +0.22 | [+0.08, +0.36] | 928 | +0.31 | +0.22 | −0.09 | [−0.15, −0.02] |
| 4-5 | 1541 | 1479 | 36 % | +0.19 | [+0.05, +0.32] | 1440 | +0.31 | +0.18 | −0.12 | [−0.18, −0.06] |
| 6-9 | 1230 | 1169 | 36 % | +0.14 | [−0.04, +0.32] | 1135 | +0.26 | +0.14 | −0.12 | [−0.22, −0.01] |

Day by day the later entries run +0.28, +0.22, +0.18, +0.20, +0.27, +0.12, +0.05 and −0.00 R from day 2 to day 9, the paired difference −0.01, −0.09, −0.12, −0.12, −0.05, −0.14, −0.16 and −0.19 (the day-9 interval [−0.31, −0.06]). Per year, days 2-3 were worse than day 1 on the same signals in 5 of 11 years, days 4-9 in 10 of 11 (2024 the exception, +0.06). The price paid is not the reason: the quoted entry drifts only 0.3 % (day 2) to 1.2 % (day 9) above the first report's. What is gone is the move of the first sessions, which is why the signals still listed on day N had done better on day 1 (+0.26 to +0.32 R) than the whole population (+0.22): a row that survives its first days is the better population, and buying it late gives part of that back. Per pattern: an inverse H&S on its second day is as good as new (+0.01 on 748 pairs), costs 0.14 R on days 4-5 (interval [−0.21, −0.08]) and is still +0.21 R unconditionally there; a Wolfe on days 6-9 lost money, −0.45 R on 44 with a 13 % hit rate (paired −0.38, [−0.73, −0.03]); cups, near zero at any age, lose 0.07 to 0.13 R from day 2 on. Three cups re-detected weeks later with the same stop (a new handle over the same low, nine listings) are left out of the buckets.

What this settles: a row on its second day is as good as a fresh one; from the third day the same signal pays about 0.1 R less than it did on day 1, in ten of eleven years, and stays positive on average through day 6 only because survivors are a better population; from day 7 the replay shows no edge (+0.12, +0.05, −0.00, intervals spanning zero), for a Wolfe from day 6. No constant changes: the age window is a first-report question (#98). The day on the list belongs in the report and on the site next to the age, with the plain entry line reserved for days 1-2, a late note from day 3 and no entry line from day 7 (Wolfe from day 6); whether the scanner should also stop listing a confirmed row after its sixth session is a rule decision under the protocol, with this table as its evidence, and stays open on #115.

### Market context (2026-09-08)

The report header and every backtest row carry the SPY regime (close and SMA50 against the SMA200), the VIX and the breadth of the universe, and the backtest summaries add a per-regime slice with VIX and breadth among the feature buckets. Over the ten years on the point-in-time index (the section above), signals scanned in a SPY bear regime ran +0.59 R on 226 against +0.16 on 1190 in bull regimes, positive in five of the six years with bear sessions, and signals scanned at a VIX above 25 ran +0.50 R in eight of eight years; a VIX below 15, negative in the two-year window, ran +0.21 R pooled. Nothing gates on the context; #101 tracks how the report should present it.
Expand Down
4 changes: 3 additions & 1 deletion docs/wiki/04-Testing-and-Contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ Work on a feature branch and open a PR to `main`; the `tests` workflow must pass
| `daily-scan` | 01:17 UTC daily, manual | tests, full scan, commit `output/` |
| `debug-last-bar` | manual, pushes touching its files | read-only per-symbol bar diagnostics |
| `evaluate-signals` | manual | replay every committed `CONFIRMED` signal against later prices; Markdown table in the job summary |
| `backtest` | manual, pushes touching its files | walk-forward replay of the scanner over the last N sessions; overall / per-pattern / per-score-bucket hit rates and R multiples in the job summary |
| `backtest` | manual, pushes touching its files | walk-forward replay of the scanner over the last N sessions; overall / per-pattern / per-score-bucket hit rates and R multiples in the job summary; the run's `backtest.json` is kept as a workflow artifact (`backtest-<run id>`) |
| `sync-wiki` | pushes to `main` touching `docs/wiki/`, manual | mirror `docs/wiki/` into the GitHub wiki |

## Measuring signal outcomes
Expand All @@ -85,6 +85,8 @@ The bootstrap resamples scan **months**, not trades: signals arrive in clusters

Two more tables follow each summary. **Excursions by horizon** gives, overall and per pattern, the median MFE and MAE within the first 5, 10, 20, 40 and 60 bars after the fill, in percent and in ATR, and the share of signals that had reached the target, hit the stop or done neither by then (medians, because a few runaway winners dominate the means). **Outcome by feature** buckets the traded signals by features measured at the scan day from the history the scan saw: breakout volume ratio and its z-score against the prior 20 bars, close vs SMA200, SMA50 vs SMA200, the SMA200's change over 40 bars, the distance from the SMA200 in ATR, the stop distance in ATR, the pattern's depth in ATR, risk %, reward:risk, the bars from the pattern's last anchor to the breakout, the breakout age when first reported, the breakout bar's close within its range and whether it cleared the prior bar's high, the VIX and the breadth, for cups the handle's volume against the cup's and the handle's volume slope, and a per-ticker fear-and-greed reading, the equal-weight 0-100 composite of RSI 14, the MACD histogram's percentile within the trailing year and Bollinger %B that the TradingView community indicators of that name share, at the scan day and at the pattern's last low, with its components. Bucket edges are fixed (`FEATURE_BUCKETS` in `tools/backtest.py`) so two replays, or the two windows of a split, compare bucket by bucket; each row shows N, hit rate, mean and median R and the month-block interval. Nothing is added to `signals.json` or the report: the features exist to be tested, and a feature earns a rule only if its buckets separate outcomes by more than their intervals on both windows of a split.

A third table, **Late entry**, follows: every later `CONFIRMED` listing of an already-seen signal (the same structure still confirmed on a following scan day, which the nightly report shows again) is filled at the next open under the same rules, and the table gives, per day on the list (1 = the first report), the listings, the traded ones, hit rate, mean and median R with the month-block interval, and the paired comparison with the same signals bought on day 1: the pairs, both means, the mean difference and its month-block interval (months of the first report). The later listings are `repeats` in `backtest.json`; first-seen rows carry `first_day` and `listed_day` 1. This is the test behind #115: every replay fills a signal once, at its first report, while the report keeps showing a confirmed row for up to 8 bars.

**Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets.

**Point-in-time universe.** By default a replay scans today's constituents, so symbols that left the index during the window are missing and symbols that joined are present before they joined: survivorship bias, which flatters bullish patterns (the dataset's 2016 snapshot holds 179 symbols that are not members today). `--constituents-asof` (workflow input `asof=true`) reconstructs the membership as of each scan day from the git history of the constituent dataset the scanner already pins: `tools/universe_history.py` makes a bare, blobless clone of `datasets/s-and-p-500-companies` under `.cache/` (under 1 MB, refreshed by a fetch on later runs), reads `data/constituents.csv` at the last commit on or before each day, downloads the union of members over the window, and scans a symbol only on the days it was a member. The report opens with a coverage table: members per year and how many of them have Yahoo history. Two limits remain. The dataset records a change some days after the index does, with about weekly commits since 2020 but one a year in 2017-2019. And a symbol that no longer trades, or was renamed, has no Yahoo history, so the replay still cannot trade it; the coverage table is the measure of the survivorship bias that is left (a transient Yahoo miss counts the same way, so compare the share across runs before reading much into one). Eleven such runs, one per calendar year from 2016 to 2026, are the multi-year evidence on the [tuning page](03-Configuration-and-Tuning.md): the tuned profile at +0.23 R per trade, positive in ten of eleven years, with drawdowns of up to 41 R inside a year. `--end YYYY-MM-DD` (input `end`) replays the `--days` sessions on or before a date, so a past year is one dispatch, and `grid=false` skips the variants and the other profile's pass when only the primary profile matters. The download must reach the window: `period=10y` for anything from 2016 on.
Expand Down
74 changes: 74 additions & 0 deletions test_backtest.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@

from __future__ import annotations

import json
import os
import re
import sys
Expand Down Expand Up @@ -435,3 +436,76 @@ def row(pattern, score, outcome, r):
assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | +0.40 | +1.40 | - | -1.00 | - | 50% | +8.0% | -2.0% |" in md
assert "| 2026-01-05 | T | Cup & Handle | 91 | 100.0 | 100.5 | 95.0 | 110.0 | gap | 0 | - |" in md
assert isinstance(pd.DataFrame(rows), pd.DataFrame)


# --------------------------------------------------------------------------- #
# Late entry: later listings of an already-seen signal
# --------------------------------------------------------------------------- #
def _extend(df: pd.DataFrame, closes) -> pd.DataFrame:
"""Append sessions closing at ``closes``: each opens at the prior close, with a small range and average volume."""
idx = pd.bdate_range(df.index[-1] + pd.Timedelta(days=1), periods=len(closes))
prev, rows = float(df["Close"].iloc[-1]), []
for c in closes:
rows.append({"Open": prev, "High": max(prev, c) * 1.005, "Low": min(prev, c) * 0.995, "Close": c,
"Volume": float(df["Volume"].iloc[-30:].mean())})
prev = c
return pd.concat([df, pd.DataFrame(rows, index=idx)])


def test_walk_forward_records_later_listings_as_late_entries(mini_universe):
cup = _extend(mini_universe["CUP"], [102.5, 103.0, 103.5]) # breakout at bar -5, listed for three more days
repeats = []
rows = bt.walk_forward({"CUP": cup}, days=6, horizon=10, repeats=repeats)
(first,) = rows
assert first["listed_day"] == 1 and first["first_day"] == first["scan_day"] == str(cup.index[-5].date())
assert [r["listed_day"] for r in repeats] == [2, 3, 4] # a cup is dropped after age 3
for k, r in enumerate(repeats, start=2):
assert r["scan_day"] == str(cup.index[-6 + k].date()) and r["first_day"] == first["scan_day"]
assert r["stop"] == first["stop"] and r["bars_since_break"] == k - 1
assert r["fill"] == round(float(cup["Open"].iloc[-5 + k]), 2) # the next session's open
assert r["outcome"] == "open" and r["bars"] == 5 - k
assert bt.walk_forward({"CUP": cup}, days=6, horizon=10) == rows # opt-in: the first-seen rows do not change
table = bt.late_entry_table(rows, repeats)
assert [t["listed_day"] for t in table] == [1, 2, 3, 4] and [t["listed"] for t in table] == [1, 1, 1, 1]
assert table[0]["pairs"] is None and table[1]["pairs"] == 1
assert table[1]["pair_day1_mean_r"] == round(first["r"], 3) and table[1]["pair_mean_r"] == round(repeats[0]["r"], 3)
assert table[1]["diff_mean_r"] == round(repeats[0]["r"] - first["r"], 3)


def test_late_entry_table_pairs_each_later_listing_with_its_first_report():
def sig(ticker, day, r, outcome="open", first_day=None, listed_day=1):
return {"ticker": ticker, "pattern": "Cup & Handle", "stop": 90.0, "scan_day": day,
"first_day": first_day or day, "listed_day": listed_day, "outcome": outcome, "r": r,
"success5": None, "mfe": 0.01 if outcome != "gap" else None, "mae": -0.01 if outcome != "gap" else None}
rows = [sig("A", "2026-01-05", 1.0), sig("B", "2026-02-03", -1.0), sig("C", "2026-03-02", 0.5)]
repeats = [sig("A", "2026-01-06", 0.5, first_day="2026-01-05", listed_day=2),
sig("B", "2026-02-04", None, outcome="gap", first_day="2026-02-03", listed_day=2),
sig("C", "2026-03-03", -1.0, first_day="2026-03-02", listed_day=2),
sig("A", "2026-01-07", 2.0, first_day="2026-01-05", listed_day=3),
sig("Z", "2026-01-07", 9.0, first_day="2026-01-05", listed_day=3)] # no first report in rows: ignored
d1, d2, d3 = bt.late_entry_table(rows, repeats)
assert (d1["listed_day"], d1["listed"], d1["n"], d1["mean_r"], d1["pairs"]) == (1, 3, 3, 0.167, None)
assert (d2["listed_day"], d2["listed"], d2["n"], d2["gap"]) == (2, 3, 2, 1) # the gap is listed, not traded
assert (d2["pairs"], d2["pair_day1_mean_r"], d2["pair_mean_r"], d2["diff_mean_r"]) == (2, 0.75, -0.25, -1.0)
assert d2["diff_ci_low"] is None # two pairs: no interval
assert (d3["listed_day"], d3["listed"], d3["pairs"], d3["diff_mean_r"]) == (3, 1, 1, 1.0)


def test_render_late_entry_section_and_json(mini_universe, monkeypatch, tmp_path, capsys):
cup = _extend(mini_universe["CUP"], [102.5, 103.0, 103.5])
repeats = []
rows = bt.walk_forward({"CUP": cup}, days=6, horizon=10, repeats=repeats)
sections = bt.report_sections(rows, 6, 10, repeats=repeats)
md = bt.render(rows, sections, 6, 10)
assert "### Late entry: buying a signal on its Nth day on the list (all sessions)" in md
assert f"| 2 | 1 | 1 | - | {bt._r(repeats[0]['r'])} |" in md # nothing resolved: no hit rate
assert bt.report_sections(rows, 6, 10)[0]["late"] is None # without repeats: no table
# End to end: main collects the later listings and writes the table and the rows to the JSON.
monkeypatch.setattr(scan, "load_sp500_symbols", lambda csv=None: ["CUP"])
monkeypatch.setattr(scan, "download_history",
lambda symbols, period="2y": {"CUP": cup} if "CUP" in symbols else {})
out = tmp_path / "bt.json"
assert bt.main(["--days", "6", "--horizon", "10", "--json", str(out)]) == 0
doc = json.loads(out.read_text())
assert [t["listed_day"] for t in doc["late"]] == [1, 2, 3, 4] and len(doc["repeats"]) == 3
assert "Late entry" in capsys.readouterr().out
Loading
Loading