diff --git a/.github/workflows/backtest.yml b/.github/workflows/backtest.yml index b5a7d23..ca6b759 100644 --- a/.github/workflows/backtest.yml +++ b/.github/workflows/backtest.yml @@ -107,3 +107,9 @@ jobs: if [ "$ASOF" = "true" ]; then cmd+=(--constituents-asof); fi for kv in $OVERRIDES; do cmd+=(--set "$kv"); done "${cmd[@]}" | tee -a "$GITHUB_STEP_SUMMARY" + - name: Keep the JSON (rows, later listings, tables) so several runs can be pooled exactly + uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 + with: + name: backtest-${{ github.run_id }} + path: backtest.json + if-no-files-found: ignore diff --git a/CHANGELOG.md b/CHANGELOG.md index ef25b8a..373680f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,22 @@ All notable changes to this project are documented here. Format follows ## [Unreleased] +### Added (late-entry replay, #115) +- The backtest records every later listing of an already-seen signal (the + same structure still `CONFIRMED` on a following scan day, which the report + shows again), filled at the next open under the same rules, and reports a + **Late entry** table per window: per day on the list (1 = the first + report), listings, traded, hit rate, mean and median R with the month-block + interval, and the paired comparison with the same signals bought on day 1 + (pairs, both means, the mean difference and its interval). `walk_forward` + takes `repeats=`; first-seen rows carry `first_day` and `listed_day`; the + JSON gains `late` and `repeats`; each `backtest` run keeps its JSON as a + workflow artifact (`backtest-`) so several runs pool exactly. + Replayed over the eleven yearly point-in-time runs (tuning page): a row on + its second day is as good as new; from the third day the same signal pays + about 0.1 R less than on day 1, in ten of eleven years; from day 7 no edge + remains, for a Wolfe from day 6. No rule changed. + ### Added (report: the F&G column) - Every report row carries `F&G`, the stock's own fear-and-greed reading at the last close, 0 to 100, one value per symbol; `fear_greed` in diff --git a/docs/wiki/03-Configuration-and-Tuning.md b/docs/wiki/03-Configuration-and-Tuning.md index 724743e..100c7eb 100644 --- a/docs/wiki/03-Configuration-and-Tuning.md +++ b/docs/wiki/03-Configuration-and-Tuning.md @@ -128,6 +128,22 @@ A confirmed breakout is almost never fearful by construction (one row below 20 i The one candidate rule this produces is the Bollinger stretch: a breakout bar that closes above its upper band. Leaving those 453 signals out would keep 1179 signals at +0.30 R against 1632 at +0.22, at a cost of 16 R of the ten-year total of 367. Under the protocol it remains a hypothesis: the gate has to be replayed as a rule on its own, so that its effect on the drawdown and per pattern is measured (#113). The reading itself is on every report row as the `F&G` column, one value per symbol at the last close, information only. +### Late entry, tested (2026-09-09) + +The owner's question on HAL, listed as a confirmed inverse head and shoulders in every report from 2026-09-04: the rules keep an H&S or Wolfe breakout listed for up to 8 bars and re-quote the entry at each close, and #98 settled that the age at *first* report is no reason to cut that window. But every replay fills a signal once, at the open after its first report, and nothing had measured what a reader gets who buys a row on its second, third or eighth day on the list, which is what the report invites by showing the row again. `tools/backtest.py` now records every later CONFIRMED listing of an already-seen signal (`listed_day` = sessions since the first report) and fills it at the next open under the same rules, and its **Late entry** table pairs each later listing with the same signal's first report (#115; the eleven yearly point-in-time runs 34323013559 to 34323039560, tuned, `bars=500`, `horizon=60`: 1692 first-seen signals, 4915 later listings). + +| Day on the list | Listed | Traded | Hit rate | Mean R | 95 % CI | Pairs | Same signals on day 1 | On day N | Difference | 95 % CI | +|---|---|---|---|---|---|---|---|---|---|---| +| 1 | 1692 | 1632 | 35 % | +0.22 | [+0.09, +0.36] | | | | | | +| 2 | 1140 | 1111 | 35 % | +0.28 | [+0.13, +0.45] | 1090 | +0.28 | +0.27 | −0.01 | [−0.07, +0.07] | +| 3 | 995 | 952 | 35 % | +0.22 | [+0.08, +0.36] | 928 | +0.31 | +0.22 | −0.09 | [−0.15, −0.02] | +| 4-5 | 1541 | 1479 | 36 % | +0.19 | [+0.05, +0.32] | 1440 | +0.31 | +0.18 | −0.12 | [−0.18, −0.06] | +| 6-9 | 1230 | 1169 | 36 % | +0.14 | [−0.04, +0.32] | 1135 | +0.26 | +0.14 | −0.12 | [−0.22, −0.01] | + +Day by day the later entries run +0.28, +0.22, +0.18, +0.20, +0.27, +0.12, +0.05 and −0.00 R from day 2 to day 9, the paired difference −0.01, −0.09, −0.12, −0.12, −0.05, −0.14, −0.16 and −0.19 (the day-9 interval [−0.31, −0.06]). Per year, days 2-3 were worse than day 1 on the same signals in 5 of 11 years, days 4-9 in 10 of 11 (2024 the exception, +0.06). The price paid is not the reason: the quoted entry drifts only 0.3 % (day 2) to 1.2 % (day 9) above the first report's. What is gone is the move of the first sessions, which is why the signals still listed on day N had done better on day 1 (+0.26 to +0.32 R) than the whole population (+0.22): a row that survives its first days is the better population, and buying it late gives part of that back. Per pattern: an inverse H&S on its second day is as good as new (+0.01 on 748 pairs), costs 0.14 R on days 4-5 (interval [−0.21, −0.08]) and is still +0.21 R unconditionally there; a Wolfe on days 6-9 lost money, −0.45 R on 44 with a 13 % hit rate (paired −0.38, [−0.73, −0.03]); cups, near zero at any age, lose 0.07 to 0.13 R from day 2 on. Three cups re-detected weeks later with the same stop (a new handle over the same low, nine listings) are left out of the buckets. + +What this settles: a row on its second day is as good as a fresh one; from the third day the same signal pays about 0.1 R less than it did on day 1, in ten of eleven years, and stays positive on average through day 6 only because survivors are a better population; from day 7 the replay shows no edge (+0.12, +0.05, −0.00, intervals spanning zero), for a Wolfe from day 6. No constant changes: the age window is a first-report question (#98). The day on the list belongs in the report and on the site next to the age, with the plain entry line reserved for days 1-2, a late note from day 3 and no entry line from day 7 (Wolfe from day 6); whether the scanner should also stop listing a confirmed row after its sixth session is a rule decision under the protocol, with this table as its evidence, and stays open on #115. + ### Market context (2026-09-08) The report header and every backtest row carry the SPY regime (close and SMA50 against the SMA200), the VIX and the breadth of the universe, and the backtest summaries add a per-regime slice with VIX and breadth among the feature buckets. Over the ten years on the point-in-time index (the section above), signals scanned in a SPY bear regime ran +0.59 R on 226 against +0.16 on 1190 in bull regimes, positive in five of the six years with bear sessions, and signals scanned at a VIX above 25 ran +0.50 R in eight of eight years; a VIX below 15, negative in the two-year window, ran +0.21 R pooled. Nothing gates on the context; #101 tracks how the report should present it. diff --git a/docs/wiki/04-Testing-and-Contributing.md b/docs/wiki/04-Testing-and-Contributing.md index b10fed5..3d5d5d2 100644 --- a/docs/wiki/04-Testing-and-Contributing.md +++ b/docs/wiki/04-Testing-and-Contributing.md @@ -74,7 +74,7 @@ Work on a feature branch and open a PR to `main`; the `tests` workflow must pass | `daily-scan` | 01:17 UTC daily, manual | tests, full scan, commit `output/` | | `debug-last-bar` | manual, pushes touching its files | read-only per-symbol bar diagnostics | | `evaluate-signals` | manual | replay every committed `CONFIRMED` signal against later prices; Markdown table in the job summary | -| `backtest` | manual, pushes touching its files | walk-forward replay of the scanner over the last N sessions; overall / per-pattern / per-score-bucket hit rates and R multiples in the job summary | +| `backtest` | manual, pushes touching its files | walk-forward replay of the scanner over the last N sessions; overall / per-pattern / per-score-bucket hit rates and R multiples in the job summary; the run's `backtest.json` is kept as a workflow artifact (`backtest-`) | | `sync-wiki` | pushes to `main` touching `docs/wiki/`, manual | mirror `docs/wiki/` into the GitHub wiki | ## Measuring signal outcomes @@ -85,6 +85,8 @@ The bootstrap resamples scan **months**, not trades: signals arrive in clusters Two more tables follow each summary. **Excursions by horizon** gives, overall and per pattern, the median MFE and MAE within the first 5, 10, 20, 40 and 60 bars after the fill, in percent and in ATR, and the share of signals that had reached the target, hit the stop or done neither by then (medians, because a few runaway winners dominate the means). **Outcome by feature** buckets the traded signals by features measured at the scan day from the history the scan saw: breakout volume ratio and its z-score against the prior 20 bars, close vs SMA200, SMA50 vs SMA200, the SMA200's change over 40 bars, the distance from the SMA200 in ATR, the stop distance in ATR, the pattern's depth in ATR, risk %, reward:risk, the bars from the pattern's last anchor to the breakout, the breakout age when first reported, the breakout bar's close within its range and whether it cleared the prior bar's high, the VIX and the breadth, for cups the handle's volume against the cup's and the handle's volume slope, and a per-ticker fear-and-greed reading, the equal-weight 0-100 composite of RSI 14, the MACD histogram's percentile within the trailing year and Bollinger %B that the TradingView community indicators of that name share, at the scan day and at the pattern's last low, with its components. Bucket edges are fixed (`FEATURE_BUCKETS` in `tools/backtest.py`) so two replays, or the two windows of a split, compare bucket by bucket; each row shows N, hit rate, mean and median R and the month-block interval. Nothing is added to `signals.json` or the report: the features exist to be tested, and a feature earns a rule only if its buckets separate outcomes by more than their intervals on both windows of a split. +A third table, **Late entry**, follows: every later `CONFIRMED` listing of an already-seen signal (the same structure still confirmed on a following scan day, which the nightly report shows again) is filled at the next open under the same rules, and the table gives, per day on the list (1 = the first report), the listings, the traded ones, hit rate, mean and median R with the month-block interval, and the paired comparison with the same signals bought on day 1: the pairs, both means, the mean difference and its month-block interval (months of the first report). The later listings are `repeats` in `backtest.json`; first-seen rows carry `first_day` and `listed_day` 1. This is the test behind #115: every replay fills a signal once, at its first report, while the report keeps showing a confirmed row for up to 8 bars. + **Out-of-sample check.** `--split YYYY-MM-DD` prints every table three times: all sessions, the sessions before the date, and the sessions from it. A rule chosen on one window is judged on the other; how the windows are used is the protocol in [Configuration and Tuning](03-Configuration-and-Tuning.md). `--bars 500` makes each scan see only its last 500 bars, which is what the nightly's `2y` download gives it, so a `--period 5y --days 500` replay reproduces two years of nightly runs rather than scans with ever-longer histories. Caveats: today's constituents only (survivorship bias), the last `horizon` sessions are still open, and both replayed years to 2026-09 were mostly bull markets. **Point-in-time universe.** By default a replay scans today's constituents, so symbols that left the index during the window are missing and symbols that joined are present before they joined: survivorship bias, which flatters bullish patterns (the dataset's 2016 snapshot holds 179 symbols that are not members today). `--constituents-asof` (workflow input `asof=true`) reconstructs the membership as of each scan day from the git history of the constituent dataset the scanner already pins: `tools/universe_history.py` makes a bare, blobless clone of `datasets/s-and-p-500-companies` under `.cache/` (under 1 MB, refreshed by a fetch on later runs), reads `data/constituents.csv` at the last commit on or before each day, downloads the union of members over the window, and scans a symbol only on the days it was a member. The report opens with a coverage table: members per year and how many of them have Yahoo history. Two limits remain. The dataset records a change some days after the index does, with about weekly commits since 2020 but one a year in 2017-2019. And a symbol that no longer trades, or was renamed, has no Yahoo history, so the replay still cannot trade it; the coverage table is the measure of the survivorship bias that is left (a transient Yahoo miss counts the same way, so compare the share across runs before reading much into one). Eleven such runs, one per calendar year from 2016 to 2026, are the multi-year evidence on the [tuning page](03-Configuration-and-Tuning.md): the tuned profile at +0.23 R per trade, positive in ten of eleven years, with drawdowns of up to 41 R inside a year. `--end YYYY-MM-DD` (input `end`) replays the `--days` sessions on or before a date, so a past year is one dispatch, and `grid=false` skips the variants and the other profile's pass when only the primary profile matters. The download must reach the window: `period=10y` for anything from 2016 on. diff --git a/test_backtest.py b/test_backtest.py index ee4926b..c9ffb9c 100644 --- a/test_backtest.py +++ b/test_backtest.py @@ -3,6 +3,7 @@ from __future__ import annotations +import json import os import re import sys @@ -435,3 +436,76 @@ def row(pattern, score, outcome, r): assert "| all | 4 | 1 | 1 | 1 | 1 | 50% | +0.47 | +0.40 | +1.40 | - | -1.00 | - | 50% | +8.0% | -2.0% |" in md assert "| 2026-01-05 | T | Cup & Handle | 91 | 100.0 | 100.5 | 95.0 | 110.0 | gap | 0 | - |" in md assert isinstance(pd.DataFrame(rows), pd.DataFrame) + + +# --------------------------------------------------------------------------- # +# Late entry: later listings of an already-seen signal +# --------------------------------------------------------------------------- # +def _extend(df: pd.DataFrame, closes) -> pd.DataFrame: + """Append sessions closing at ``closes``: each opens at the prior close, with a small range and average volume.""" + idx = pd.bdate_range(df.index[-1] + pd.Timedelta(days=1), periods=len(closes)) + prev, rows = float(df["Close"].iloc[-1]), [] + for c in closes: + rows.append({"Open": prev, "High": max(prev, c) * 1.005, "Low": min(prev, c) * 0.995, "Close": c, + "Volume": float(df["Volume"].iloc[-30:].mean())}) + prev = c + return pd.concat([df, pd.DataFrame(rows, index=idx)]) + + +def test_walk_forward_records_later_listings_as_late_entries(mini_universe): + cup = _extend(mini_universe["CUP"], [102.5, 103.0, 103.5]) # breakout at bar -5, listed for three more days + repeats = [] + rows = bt.walk_forward({"CUP": cup}, days=6, horizon=10, repeats=repeats) + (first,) = rows + assert first["listed_day"] == 1 and first["first_day"] == first["scan_day"] == str(cup.index[-5].date()) + assert [r["listed_day"] for r in repeats] == [2, 3, 4] # a cup is dropped after age 3 + for k, r in enumerate(repeats, start=2): + assert r["scan_day"] == str(cup.index[-6 + k].date()) and r["first_day"] == first["scan_day"] + assert r["stop"] == first["stop"] and r["bars_since_break"] == k - 1 + assert r["fill"] == round(float(cup["Open"].iloc[-5 + k]), 2) # the next session's open + assert r["outcome"] == "open" and r["bars"] == 5 - k + assert bt.walk_forward({"CUP": cup}, days=6, horizon=10) == rows # opt-in: the first-seen rows do not change + table = bt.late_entry_table(rows, repeats) + assert [t["listed_day"] for t in table] == [1, 2, 3, 4] and [t["listed"] for t in table] == [1, 1, 1, 1] + assert table[0]["pairs"] is None and table[1]["pairs"] == 1 + assert table[1]["pair_day1_mean_r"] == round(first["r"], 3) and table[1]["pair_mean_r"] == round(repeats[0]["r"], 3) + assert table[1]["diff_mean_r"] == round(repeats[0]["r"] - first["r"], 3) + + +def test_late_entry_table_pairs_each_later_listing_with_its_first_report(): + def sig(ticker, day, r, outcome="open", first_day=None, listed_day=1): + return {"ticker": ticker, "pattern": "Cup & Handle", "stop": 90.0, "scan_day": day, + "first_day": first_day or day, "listed_day": listed_day, "outcome": outcome, "r": r, + "success5": None, "mfe": 0.01 if outcome != "gap" else None, "mae": -0.01 if outcome != "gap" else None} + rows = [sig("A", "2026-01-05", 1.0), sig("B", "2026-02-03", -1.0), sig("C", "2026-03-02", 0.5)] + repeats = [sig("A", "2026-01-06", 0.5, first_day="2026-01-05", listed_day=2), + sig("B", "2026-02-04", None, outcome="gap", first_day="2026-02-03", listed_day=2), + sig("C", "2026-03-03", -1.0, first_day="2026-03-02", listed_day=2), + sig("A", "2026-01-07", 2.0, first_day="2026-01-05", listed_day=3), + sig("Z", "2026-01-07", 9.0, first_day="2026-01-05", listed_day=3)] # no first report in rows: ignored + d1, d2, d3 = bt.late_entry_table(rows, repeats) + assert (d1["listed_day"], d1["listed"], d1["n"], d1["mean_r"], d1["pairs"]) == (1, 3, 3, 0.167, None) + assert (d2["listed_day"], d2["listed"], d2["n"], d2["gap"]) == (2, 3, 2, 1) # the gap is listed, not traded + assert (d2["pairs"], d2["pair_day1_mean_r"], d2["pair_mean_r"], d2["diff_mean_r"]) == (2, 0.75, -0.25, -1.0) + assert d2["diff_ci_low"] is None # two pairs: no interval + assert (d3["listed_day"], d3["listed"], d3["pairs"], d3["diff_mean_r"]) == (3, 1, 1, 1.0) + + +def test_render_late_entry_section_and_json(mini_universe, monkeypatch, tmp_path, capsys): + cup = _extend(mini_universe["CUP"], [102.5, 103.0, 103.5]) + repeats = [] + rows = bt.walk_forward({"CUP": cup}, days=6, horizon=10, repeats=repeats) + sections = bt.report_sections(rows, 6, 10, repeats=repeats) + md = bt.render(rows, sections, 6, 10) + assert "### Late entry: buying a signal on its Nth day on the list (all sessions)" in md + assert f"| 2 | 1 | 1 | - | {bt._r(repeats[0]['r'])} |" in md # nothing resolved: no hit rate + assert bt.report_sections(rows, 6, 10)[0]["late"] is None # without repeats: no table + # End to end: main collects the later listings and writes the table and the rows to the JSON. + monkeypatch.setattr(scan, "load_sp500_symbols", lambda csv=None: ["CUP"]) + monkeypatch.setattr(scan, "download_history", + lambda symbols, period="2y": {"CUP": cup} if "CUP" in symbols else {}) + out = tmp_path / "bt.json" + assert bt.main(["--days", "6", "--horizon", "10", "--json", str(out)]) == 0 + doc = json.loads(out.read_text()) + assert [t["listed_day"] for t in doc["late"]] == [1, 2, 3, 4] and len(doc["repeats"]) == 3 + assert "Late entry" in capsys.readouterr().out diff --git a/tools/backtest.py b/tools/backtest.py index 60334ba..3001f5c 100644 --- a/tools/backtest.py +++ b/tools/backtest.py @@ -56,6 +56,12 @@ ``scan.market_series``: the SPY regime (close and SMA50 against the SMA200), the VIX and the breadth of the universe. The summaries add a per-regime slice; VIX and breadth join the feature buckets. No rule reads any of it. +* Every **later listing** of an already-seen signal (the same structure still + CONFIRMED on a following scan day, which the nightly report shows again) is + recorded too, with ``listed_day`` = sessions since the first report (1 = the + first report) and filled at the next open under the same rules, so the + late-entry table can say what a reader gets who buys a row on its Nth day + on the list against the same signals bought on day 1. * Every summary carries, next to hit rate and mean R: median R, standard deviation, total R, the deepest drawdown of the cumulative R curve (1 R per trade, in scan order), a 95 % bootstrap interval of the mean R and the @@ -528,11 +534,32 @@ def scan_sessions(data: Dict[str, pd.DataFrame], days: int, end: Optional[str] = return sessions[-days:] if days < len(sessions) else sessions +def _fill_outcome(s: scan.Signal, after: pd.DataFrame, horizon: int) -> dict: + """Fill a signal at the next open and score it: ``fill``, ``outcome``, ``bars``, ``exit``, ``r``, excursions. + + An open above the row's Max buy is a ``gap``, one at or below the stop is + ``below_stop``; neither is traded. Shared by the first report of a signal + and its later listings so both are judged by the same rules. + """ + not_traded = dict(exit=None, r=None, mfe=None, mae=None, success5=None) + if after.empty: + return dict(fill=None, outcome="no_data", bars=0, **not_traded) + fill = float(after["Open"].iloc[0]) + max_buy = s.max_buy if s.max_buy is not None else s.entry * (1 + scan.MAX_RUNAWAY) + if fill > max_buy: + return dict(fill=round(fill, 2), outcome="gap", bars=0, **not_traded) + if fill <= s.stop: + # The open is already through the stop: no trade, and R would be undefined. + return dict(fill=round(fill, 2), outcome="below_stop", bars=0, **not_traded) + res = ev.classify(fill, s.stop, s.target, after, horizon) + return dict(fill=round(fill, 2), **res, **excursions(fill, s.stop, after, horizon)) + + def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, detectors: Optional[Sequence] = None, bars: Optional[int] = None, market: Optional[Mapping[str, pd.Series]] = None, members: Optional[Callable[[Any], Container[str]]] = None, - end: Optional[str] = None) -> List[dict]: + end: Optional[str] = None, repeats: Optional[List[dict]] = None) -> List[dict]: """Replay the scanner over the last ``days`` sessions and score each first-seen signal. :param data: ``{symbol: OHLCV frame}`` as returned by ``download_history``. @@ -547,6 +574,11 @@ def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, is scanned only on days it is in that set (``universe_history.Membership.members``). :param end: Last scan day (ISO); the window is the ``days`` sessions on or before it, so a past year can be replayed on its own. + :param repeats: When given, every later CONFIRMED listing of an already-seen + signal is appended to it: the day's signal fields plus ``first_day``, + ``listed_day`` (sessions since the first report, 1 = that report) and the + fill / outcome of buying it at the next open (:func:`_fill_outcome`). + First-seen rows carry ``first_day`` = ``scan_day`` and ``listed_day`` 1. :returns: One dict per first-seen CONFIRMED signal with the signal fields plus ``fill``, ``outcome``, ``bars``, ``exit``, ``r``, the :func:`row_features`, the market context (and for cups the parsed ``cup_bottom`` / ``cup_trigger`` @@ -559,8 +591,8 @@ def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, """ scan_days = scan_sessions(data, days, end) seen: Dict[tuple, dict] = {} - not_traded = dict(exit=None, r=None, mfe=None, mae=None, success5=None) - for d in scan_days: + first_no: Dict[tuple, int] = {} # key -> position of its first report among the scan days + for no, d in enumerate(scan_days): ctx = _market_at(market, d) today = members(d) if members is not None else None for sym, df in data.items(): @@ -576,29 +608,23 @@ def walk_forward(data: Dict[str, pd.DataFrame], days: int, horizon: int, continue assert s.last_date == str(d.date()), "look-ahead: signal dated after the scan day" key = (sym, s.pattern, round(s.stop, 2)) + after = df[df.index > d] if key in seen: + if repeats is not None: # the report shows the row again: a late entry + repeats.append({**scan.asdict(s), "scan_day": str(d.date()), + "first_day": seen[key]["scan_day"], "listed_day": no - first_no[key] + 1, + **_fill_outcome(s, after, horizon)}) continue - after = df[df.index > d] + first_no[key] = no atr_last = round(float(scan.atr(hist).iloc[-1]), 4) - row = {**scan.asdict(s), "scan_day": str(d.date()), "atr": atr_last, - "cup_bottom": None, "cup_trigger": None, **row_features(hist, s, atr_last), **ctx} + row = {**scan.asdict(s), "scan_day": str(d.date()), "first_day": str(d.date()), "listed_day": 1, + "atr": atr_last, "cup_bottom": None, "cup_trigger": None, + **row_features(hist, s, atr_last), **ctx} if s.pattern == "Cup & Handle": m = re.search(r"bottom \S+ @([\d.]+).*trigger ([\d.]+)", s.notes) if m: row["cup_bottom"], row["cup_trigger"] = float(m[1]), float(m[2]) - if after.empty: - row.update(fill=None, outcome="no_data", bars=0, **not_traded) - else: - fill = float(after["Open"].iloc[0]) - max_buy = s.max_buy if s.max_buy is not None else s.entry * (1 + scan.MAX_RUNAWAY) - if fill > max_buy: - row.update(fill=round(fill, 2), outcome="gap", bars=0, **not_traded) - elif fill <= s.stop: - # The open is already through the stop: no trade, and R would be undefined. - row.update(fill=round(fill, 2), outcome="below_stop", bars=0, **not_traded) - else: - res = ev.classify(fill, s.stop, s.target, after, horizon) - row.update(fill=round(fill, 2), **res, **excursions(fill, s.stop, after, horizon)) + row.update(_fill_outcome(s, after, horizon)) seen[key] = row return list(seen.values()) @@ -718,19 +744,67 @@ def feature_buckets(rows: Sequence[dict]) -> List[dict]: return out +def _mean(xs: Sequence[float]) -> Optional[float]: + return round(sum(xs) / len(xs), 3) if xs else None + + +def late_entry_table(rows: Sequence[dict], repeats: Sequence[dict]) -> List[dict]: + """Outcome of buying a signal on its Nth day on the list, against the same signals bought on day 1. + + Day 1 is the first report (``rows``); a later listing (``repeats``, see + :func:`walk_forward`) belongs to the table when its first report is in + ``rows``, so a split window keeps a signal and its repeats together. Each + day's row summarises the listings filled that day and pairs every traded + one with its own first report: ``pairs``, ``pair_day1_mean_r``, + ``pair_mean_r``, ``diff_mean_r`` (day N minus day 1 over the pairs) and the + month-block interval of that difference (``diff_ci_low`` / ``diff_ci_high``, + months of the first report). + + :returns: One dict per day on the list with listings, in day order. + """ + def key(r: dict) -> tuple: + return (r["ticker"], r["pattern"], round(float(r["stop"]), 2), r["first_day"]) + + first = {key(r): r for r in rows} + later = [r for r in repeats if key(r) in first] + out = [] + for day in sorted({1} | {int(r["listed_day"]) for r in later}): + sub = list(rows) if day == 1 else [r for r in later if int(r["listed_day"]) == day] + s = summarise_rows(sub) + entry: Dict[str, Any] = {"listed_day": day, "listed": len(sub), "pairs": None, "pair_day1_mean_r": None, + "pair_mean_r": None, "diff_mean_r": None, "diff_ci_low": None, "diff_ci_high": None, + **{k: s[k] for k in ("n", "gap", "below_stop", "target", "stop", "open", "hit_rate", + "mean_r", "median_r", "ci_low", "ci_high")}} + if day > 1: + pairs = [(first[key(r)], r) for r in sub + if r["outcome"] in TRADED and r["r"] is not None + and first[key(r)]["outcome"] in TRADED and first[key(r)]["r"] is not None] + diffs = [{"scan_day": a["scan_day"], "r": float(b["r"]) - float(a["r"])} for a, b in pairs] + boot = block_bootstrap(diffs) + entry.update(pairs=len(pairs), pair_day1_mean_r=_mean([float(a["r"]) for a, _ in pairs]), + pair_mean_r=_mean([float(b["r"]) for _, b in pairs]), + diff_mean_r=_mean([x["r"] for x in diffs]), + diff_ci_low=boot["ci_low"], diff_ci_high=boot["ci_high"]) + out.append(entry) + return out + + def report_sections(rows: Sequence[dict], days: int, horizon: int, split: Optional[str] = None, data: Optional[Dict[str, pd.DataFrame]] = None, do_grid: bool = False, - other_rows: Optional[Sequence[dict]] = None, other_name: Optional[str] = None) -> List[dict]: + other_rows: Optional[Sequence[dict]] = None, other_name: Optional[str] = None, + repeats: Optional[Sequence[dict]] = None) -> List[dict]: """One section per window (see :func:`split_windows`): its rows, ``breakdown`` stats, the ``feature_buckets``, the ``horizon_table`` (with ``data``), the stop / - target ``grid`` (with ``do_grid`` and ``data``) and the other profile's stats - over the same window (with ``other_rows``).""" + target ``grid`` (with ``do_grid`` and ``data``), the other profile's stats + over the same window (with ``other_rows``) and the ``late_entry_table`` (with + ``repeats``).""" windows = split_windows(rows, split) others = split_windows(other_rows, split) if other_rows is not None else [(None, None)] * len(windows) sections = [] for (label, rws), (_, o_rws) in zip(windows, others): sec: Dict[str, Any] = {"label": label, "rows": rws, "stats": breakdown(rws), - "features": feature_buckets(rws), "horizons": None, "grid": None, "other": None} + "features": feature_buckets(rws), "horizons": None, "grid": None, "other": None, + "late": late_entry_table(rws, repeats) if repeats is not None else None} if data is not None: sec["horizons"] = horizon_table(rws, data) if do_grid: @@ -803,6 +877,23 @@ def _feature_lines(buckets: Sequence[dict], label: str) -> List[str]: return lines +def _late_lines(table: Sequence[dict], label: str) -> List[str]: + lines = ["", f"### Late entry: buying a signal on its Nth day on the list ({label})", "", + "Day 1 = the first report (the rows above). A signal stays listed while it is still CONFIRMED on later " + "scan days, and the report shows it again; each later listing is filled at the next open under the " + "same rules. Pairs = listings traded on both that day and day 1; the difference is day N minus day 1 " + "over those pairs, with its month-block interval.", "", + "| Day on the list | Listed | Traded | Hit rate | Mean R | Median R | 95% CI | Pairs | Day 1 mean R | " + "Day N mean R | Difference | 95% CI |", + "|---|---|---|---|---|---|---|---|---|---|---|---|"] + for t in table: + diff_ci = "-" if t["diff_ci_low"] is None else f"[{t['diff_ci_low']:+.2f}, {t['diff_ci_high']:+.2f}]" + lines.append(f"| {t['listed_day']} | {t['listed']} | {t['n']} | {_pct(t['hit_rate'])} | {_r(t['mean_r'])} | " + f"{_r(t['median_r'])} | {_ci(t)} | {t['pairs'] if t['pairs'] is not None else '-'} | " + f"{_r(t['pair_day1_mean_r'])} | {_r(t['pair_mean_r'])} | {_r(t['diff_mean_r'])} | {diff_ci} |") + return lines + + def render(rows: Sequence[dict], sections: Sequence[dict], days: int, horizon: int, universe: Optional[Mapping[str, Any]] = None, end: Optional[str] = None) -> str: """Markdown report (readable as a GitHub step summary): the universe, one block per window, every signal.""" @@ -834,6 +925,8 @@ def render(rows: Sequence[dict], sections: Sequence[dict], days: int, horizon: i lines += _horizon_lines(sec["horizons"], sec["label"]) if sec.get("features"): lines += _feature_lines(sec["features"], sec["label"]) + if sec.get("late"): + lines += _late_lines(sec["late"], sec["label"]) if sec["grid"]: lines += ["", f"### Stop / target variants (same signals, {sec['label']})", "", "Extra ATR = distance added below the reported stop (which already sits 0.25 ATR under the " @@ -951,29 +1044,32 @@ def main(argv: Optional[Sequence[str]] = None) -> int: json.dump({"days": args.days, "horizon": args.horizon, "bars": args.bars, "end": args.end, "universe": universe, "ablation": table}, fh, indent=2) return 0 - rows = walk_forward(data, args.days, args.horizon, **replay) - log.info("replay done in %.0fs: %d first-seen confirmed signals", time.time() - t0, len(rows)) + repeats: List[dict] = [] + rows = walk_forward(data, args.days, args.horizon, repeats=repeats, **replay) + log.info("replay done in %.0fs: %d first-seen confirmed signals, %d later listings", + time.time() - t0, len(rows), len(repeats)) other_rows = other_name = None if args.grid: other_name = "legacy" if scan.ACTIVE_PROFILE == "spec" else "spec" other_rows = profile_pass(data, args.days, args.horizon, other_name, **replay)["rows"] log.info("%s-profile pass done in %.0fs: %d first-seen confirmed signals", other_name, time.time() - t0, len(other_rows)) - sections = report_sections(rows, args.days, args.horizon, args.split, data, args.grid, other_rows, other_name) + sections = report_sections(rows, args.days, args.horizon, args.split, data, args.grid, other_rows, other_name, + repeats) print(render(rows, sections, args.days, args.horizon, universe, args.end)) if args.json: pooled = sections[0] - keep = ("label", "stats", "features", "horizons", "grid", "other") + keep = ("label", "stats", "features", "horizons", "grid", "other", "late") with open(args.json, "w", encoding="utf-8") as fh: json.dump({"days": args.days, "horizon": args.horizon, "profile": scan.ACTIVE_PROFILE, "min_score": scan.MIN_SCORE, "bars": args.bars, "split": args.split, "end": args.end, "universe": universe, "market_symbols": sorted(market_frames), "stats": pooled["stats"], "features": pooled["features"], "horizons": pooled["horizons"], - "grid": pooled["grid"], + "grid": pooled["grid"], "late": pooled["late"], "other_profile": ({"profile": other_name, "stats": pooled["other"]["stats"], "rows": other_rows} if other_rows is not None else None), "windows": [{k: s[k] for k in keep} for s in sections[1:]], - "rows": rows}, fh, indent=2) + "rows": rows, "repeats": repeats}, fh, indent=2) return 0