Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,25 @@ All notable changes to this project are documented here. Format follows

## [Unreleased]

### Added (backtest features from the second external review, 2026-09-08)
- Replay rows and the outcome-by-feature table gain the pattern's depth in
ATR, the breakout bar's close within its range and whether it cleared the
prior bar's high, the breakout volume z-score against the prior 20 bars,
and for cups the handle's volume against the cup's and the handle's volume
slope. The stop / target grid gains a 1.25-ATR stop variant. Recorded to be
tested; no rule uses them. Tested over the same eleven years (tuning page):
the volume z-score rises monotonically with mean R but a gate at 1.5 would
keep 19 % of the signals; a close above the prior bar's high ran +0.19 R
against +0.42 without it; the close's position within the bar and the
handle's volume decay had no effect; a cup R² of 0.80 made cups worse; the
1.25-ATR stop ran +0.20 R against +0.29. The review's other proposals were
already in place (converging Wolfe lines and the intersection ETA, a convex
quadratic cup fit with the U-versus-V test, close-based confirmation,
ATR-buffered stops), already tested and rejected (a volume gate, a
reward:risk floor of 2.0), or contradicted by the ten-year replay
(inhibiting alerts below the SPY 200-day average or above VIX 25, where the
system's best trades were).

### Changed (docs: ten years on the index as it was)
- Tuning page: the tuned profile replayed one calendar year at a time from
2016 to 2026 with point-in-time membership (runs 34246882614 to
Expand Down
16 changes: 16 additions & 0 deletions docs/wiki/03-Configuration-and-Tuning.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,22 @@ What the ten years overturned from the two-year window, each a lesson in what on

**Open questions:** #99 and #101 (the deep-down-trend and bear-regime contexts, information first, no rule), #108 (the drawdowns), #109 (cups), and #97 (a per-signal probability model, whose first revisit condition the ten years now meet).

### A second review, tested (2026-09-08)

A second external review proposed a converging-wedge check for Wolfe (already the code: the 2-4 line must fall faster than 1-3 and the target is their intersection), a convex quadratic cup fit (already the code, with the U-versus-V test the review lacks), close-based confirmation (already the code), ATR-buffered stops (already the code, at 0.25 ATR), a volume z-score gate, a reward:risk floor of 2.0, a handle volume decay rule for cups, a confirmation candle, and a macro filter inhibiting alerts below the SPY 200-day average or above VIX 25. The testable items were replayed over the same eleven years with the features recorded on every row (runs 34255129970 to 34255999955; the second set with `CUP_MIN_ROUNDNESS=0.8`; the stop variant on the two-year grid, run 34256064869):

| Proposal | Ten-year result | Verdict |
|---|---|---|
| Volume z-score of 1.5 as a gate | monotone: at or below 0 +0.16 R on 795 signals, 0-1.5 +0.25 on 520, 1.5-3 +0.31 on 171, above 3 +0.34 on 144; the gate keeps 315 of 1630 signals at +0.32 and drops 1315 at +0.20 | a feature, not a gate |
| Confirmation candle: close above the prior bar's high | with it +0.19 R on 1405; without it +0.42 R on 225, positive in 9 of 11 years | reversed |
| Strong close within the breakout bar's range | at or below 50 % +0.24, 50-80 % +0.24, above 80 % +0.21 | no effect |
| Handle volume must decay (cups) | handle over cup volume at or below 0.7 +0.06 R, 0.7-1.0 +0.01, above 1.0 +0.06; falling slope +0.05, rising +0.01 | no effect |
| Cup R² of 0.80 instead of 0.70 | cups 127 traded at −0.10 R against 307 at +0.03; the 180 removed ran +0.12 | worse |
| Larger swings (the ATR-pivot idea), read as pattern depth in ATR | at or below 3 ATR +0.26 R on 949 with a 42 % hit rate, above 10 ATR +0.26 on 116 with a 15 % hit rate | shallow patterns hit more often; no gate |
| Stop 1.2 ATR under the pivot | two-year grid on 327 signals: the reported stop (0.25 ATR under the low) +0.29 R at a 41 % hit rate; 1.25 ATR under it +0.20 R at 49 %, and the same order in both windows (+0.30 against +0.21 unseen, +0.27 against +0.19 tuning year); the risk unit grows faster than the stop-outs fall | worse per trade |
| Reward:risk floor of 2.0 | rows at or below 1.5 ran +0.20 R with a 54 % hit rate over ten years (#100) | rejected |
| Inhibit alerts in a SPY bear regime or above VIX 25 | bear regime +0.59 R on 226, VIX above 25 +0.50 R in 8 of 8 years (#101) | contradicted |

### Market context (2026-09-08)

The report header and every backtest row carry the SPY regime (close and SMA50 against the SMA200), the VIX and the breadth of the universe, and the backtest summaries add a per-regime slice with VIX and breadth among the feature buckets. Over the ten years on the point-in-time index (the section above), signals scanned in a SPY bear regime ran +0.59 R on 226 against +0.16 on 1190 in bull regimes, positive in five of the six years with bear sessions, and signals scanned at a VIX above 25 ran +0.50 R in eight of eight years; a VIX below 15, negative in the two-year window, ran +0.21 R pooled. Nothing gates on the context; #101 tracks how the report should present it.
Expand Down
4 changes: 2 additions & 2 deletions docs/wiki/04-Testing-and-Contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,11 +79,11 @@ Work on a feature branch and open a PR to `main`; the `tests` workflow must pass

## Measuring signal outcomes

`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open, and classifies the outcome within a horizon. Two kinds of open are not traded and counted separately: one above the row's Max buy (`gap`) and one at or below the stop (`below_stop`, where nobody buys and R would be undefined). Output, overall, per pattern and per score bucket: hit rate, mean R, median R, standard deviation, total R, the deepest drawdown of the cumulative R curve (1 R per trade, in scan order), a 95 % bootstrap interval of the mean R and the drawdown exceeded in only 5 % of resamples, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE; then every signal; and (with `--grid`) the same signals re-scored under stop-distance, stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile.
`tools/backtest.py` is the fast path: for each of the last N sessions it truncates every symbol's history at that day, runs the scanner exactly as the nightly job would have, takes each `CONFIRMED` signal on the day it first appears, fills at the next session's open, and classifies the outcome within a horizon. Two kinds of open are not traded and counted separately: one above the row's Max buy (`gap`) and one at or below the stop (`below_stop`, where nobody buys and R would be undefined). Output, overall, per pattern and per score bucket: hit rate, mean R, median R, standard deviation, total R, the deepest drawdown of the cumulative R curve (1 R per trade, in scan order), a 95 % bootstrap interval of the mean R and the drawdown exceeded in only 5 % of resamples, the chart-book "+5 % before a close below the stop" success share and mean MFE / MAE; then every signal; and (with `--grid`) the same signals re-scored under stop-distance (0.25 to 1.25 ATR under the structural low), stop-basis and target variants (reported measured move, half of it, and the Investopedia bottom-to-breakout measure for cups), plus a second full walk-forward under the other rule profile (`spec` vs `legacy`) so both rule sets are compared on the same data, per pattern and per score bucket. `--profile` picks the primary profile.

The bootstrap resamples scan **months**, not trades: signals arrive in clusters (33 in one month and 2 in another over 2024-26), so treating them as independent draws would make the interval far too narrow. An interval is only printed with at least two months and five trades.

Two more tables follow each summary. **Excursions by horizon** gives, overall and per pattern, the median MFE and MAE within the first 5, 10, 20, 40 and 60 bars after the fill, in percent and in ATR, and the share of signals that had reached the target, hit the stop or done neither by then (medians, because a few runaway winners dominate the means). **Outcome by feature** buckets the traded signals by features measured at the scan day from the history the scan saw: breakout volume ratio, close vs SMA200, SMA50 vs SMA200, the SMA200's change over 40 bars, the distance from the SMA200 in ATR, the stop distance in ATR, risk %, reward:risk, the bars from the pattern's last anchor to the breakout, and the breakout age when first reported. Bucket edges are fixed (`FEATURE_BUCKETS` in `tools/backtest.py`) so two replays, or the two windows of a split, compare bucket by bucket; each row shows N, hit rate, mean and median R and the month-block interval. Nothing is added to `signals.json` or the report: the features exist to be tested, and a feature earns a rule only if its buckets separate outcomes by more than their intervals on both windows of a split.
Two more tables follow each summary. **Excursions by horizon** gives, overall and per pattern, the median MFE and MAE within the first 5, 10, 20, 40 and 60 bars after the fill, in percent and in ATR, and the share of signals that had reached the target, hit the stop or done neither by then (medians, because a few runaway winners dominate the means). **Outcome by feature** buckets the traded signals by features measured at the scan day from the history the scan saw: breakout volume ratio and its z-score against the prior 20 bars, close vs SMA200, SMA50 vs SMA200, the SMA200's change over 40 bars, the distance from the SMA200 in ATR, the stop distance in ATR, the pattern's depth in ATR, risk %, reward:risk, the bars from the pattern's last anchor to the breakout, the breakout age when first reported, the breakout bar's close within its range and whether it cleared the prior bar's high, the VIX and the breadth, and for cups the handle's volume against the cup's and the handle's volume slope. Bucket edges are fixed (`FEATURE_BUCKETS` in `tools/backtest.py`) so two replays, or the two windows of a split, compare bucket by bucket; each row shows N, hit rate, mean and median R and the month-block interval. Nothing is added to `signals.json` or the report: the features exist to be tested, and a feature earns a rule only if its buckets separate outcomes by more than their intervals on both windows of a split.

**Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets.

Expand Down
40 changes: 38 additions & 2 deletions test_backtest.py
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,22 @@ def test_row_features_match_hand_computation(mini_universe):
assert f["target_atr"] == round((s.target - s.entry) / atr_last, 3) > 0
anchor = cup.index.get_loc(pd.Timestamp(re.search(r"handle low (\S+)", s.notes)[1]))
assert f["wait_bars"] == len(cup) - 1 - s.bars_since_break - anchor >= 1 # the handle low precedes the break
# The breakout bar and the cup / handle volumes, recomputed from the frame and the notes.
b = len(cup) - 1 - s.bars_since_break
high, low, vol = (cup[c].to_numpy() for c in ("High", "Low", "Volume"))
assert f["break_close_pos"] == round((close.iloc[b] - low[b]) / (high[b] - low[b]), 3)
assert 0 <= f["break_close_pos"] <= 1
assert f["break_over_prior_high"] == (1.0 if close.iloc[b] > high[b - 1] else 0.0)
base = vol[b - scan.VOLUME_AVG_LEN:b]
assert f["volume_z"] == round((vol[b] - base.mean()) / base.std(ddof=1), 2) > 1.0 # the fixture's 3x volume bar
a = cup.index.get_loc(pd.Timestamp(re.search(r"left rim (\S+)", s.notes)[1]))
rb = cup.index.get_loc(pd.Timestamp(re.search(r"right rim (\S+)", s.notes)[1]))
assert f["handle_volume_ratio"] == round(vol[rb + 1:b].mean() / vol[a:rb + 1].mean(), 3)
handle_v = vol[rb + 1:b]
assert f["handle_volume_slope"] == round(np.polyfit(np.arange(len(handle_v)), handle_v, 1)[0] / handle_v.mean(), 4)
bottom = float(re.search(r"bottom \S+ @([\d.]+)", s.notes)[1])
rim_b = float(re.search(r"right rim \S+ @([\d.]+)", s.notes)[1])
assert f["depth_atr"] == round((rim_b - bottom) / atr_last, 3) > 3
# Too little history for the averages: those features are None, the ATR-based ones and the wait remain.
short = bt.row_features(cup.iloc[-150:], s, atr_last)
assert short["close_vs_sma200"] is None and short["sma200_slope"] is None and short["dist_sma200_atr"] is None
Expand All @@ -96,6 +112,24 @@ def test_row_features_match_hand_computation(mini_universe):
None, "", "5 2020-01-01 @95.00")
g = bt.row_features(cup, no_target, 0.0)
assert g["target_atr"] is None and g["stop_atr"] is None and g["wait_bars"] is None # anchor date not in hist
assert g["depth_atr"] is None and g["handle_volume_ratio"] is None # no anchors, not a cup
assert g["break_close_pos"] is not None and g["volume_z"] is not None # the last bar still is a bar


def test_row_features_per_pattern(mini_universe):
ihs, ww = mini_universe["IHS"], mini_universe["WW"]
(si,) = scan.detect_inverse_hs(ihs, "IHS")
fi = bt.row_features(ihs, si, float(scan.atr(ihs).iloc[-1]))
anchors = re.search(r"LS \S+ @([\d.]+), head \S+ @([\d.]+), RS \S+ @([\d.]+)", si.notes)
ls, head, rs = (float(x) for x in anchors.groups())
assert fi["depth_atr"] == round((min(ls, rs) - head) / float(scan.atr(ihs).iloc[-1]), 3) > 0
assert fi["handle_volume_ratio"] is None and fi["handle_volume_slope"] is None # cups only
(sw,) = scan.detect_bullish_wolfe(ww, "WW")
fw = bt.row_features(ww, sw, float(scan.atr(ww).iloc[-1]))
p1 = float(re.search(r"^1 \S+ @([\d.]+)", sw.notes)[1])
p5 = float(re.search(r", 5 \S+ @([\d.]+);", sw.notes)[1])
assert fw["depth_atr"] == round((p1 - p5) / float(scan.atr(ww).iloc[-1]), 3) > 0
assert fw["break_over_prior_high"] in (0.0, 1.0) and fw["volume_z"] is not None


def test_walk_forward_rows_carry_replay_features(mini_universe):
Expand Down Expand Up @@ -270,19 +304,21 @@ def test_split_windows_and_report_sections(mini_universe):
assert "## All sessions" in md and f"## Before {split}" in md and f"## From {split}" in md
assert md.count("### Stop / target variants") == 3 and "judged on the other" in md
assert "### Excursions by horizon (all sessions)" in md and "### Outcome by feature (all sessions)" in md
assert "| entry minus stop, in ATR |" in md
assert "| entry minus stop, in ATR |" in md and "| breakout bar close within its range |" in md
assert "| handle volume / cup volume (cups) |" in md and "| pattern depth in ATR |" in md
assert "## Signals" in md


def test_grid_rescores_the_same_signals(mini_universe):
rows = bt.walk_forward(mini_universe, days=5, horizon=10)
g = bt.grid(rows, mini_universe, 10)
assert len(g) == len(bt.GRID) == 18
assert len(g) == len(bt.GRID) == 24
assert {(x["stop_extra_atr"], x["stop_basis"], x["target_mode"]) for x in g} == set(bt.GRID)
traded = sum(1 for r in rows if r["outcome"] in ("target", "stop", "open"))
assert all(x["n"] == traded for x in g)
md = bt.render(rows, bt.report_sections(rows, 5, 10, data=mini_universe, do_grid=True), 5, 10)
assert "### Stop / target variants" in md and "| 0.75 | close | breakout |" in md
assert "| 1.0 | intraday | full |" in md and "1.25 ATR under it" in md


def test_breakout_level_target_applies_to_cups_only(mini_universe):
Expand Down
Loading
Loading