Skip to content

feat: backtest confidence intervals, drawdown, in-sample / out-of-sample split and --bars - #103

Merged
yanivil merged 2 commits into
mainfrom
feat/backtest-split-and-intervals
Sep 7, 2026
Merged

yanivil merged 2 commits into
mainfrom
feat/backtest-split-and-intervals

Conversation

@yanivil

@yanivil yanivil commented Sep 7, 2026

Copy link
Copy Markdown
Owner

Description

Implements #91 with #94 folded in.

  • Statistics per slice (overall, per pattern, per score bucket, other profile): median R, standard deviation, total R, deepest drawdown of the cumulative R curve (1 R per trade in scan order), a 95 % bootstrap interval of the mean R and the drawdown exceeded in 5 % of resamples. The bootstrap resamples scan months, not trades: signals cluster in time (33 in June 2025, 2 in March 2026), so a per-trade interval would be too narrow. No interval is printed with fewer than two months or five trades.
  • --split YYYY-MM-DD: every table is reported three times, all sessions, before the date, from the date, so a rule chosen on one window is judged on the other. Sections nest the stop / target grid and the other-profile comparison per window.
  • --bars N: history each scan sees, so --period 5y --days 500 --bars 500 reproduces two years of nightly runs instead of scans with ever-longer histories.
  • Fill at or below the stop is below_stop: not traded, counted with the gaps in a "Not traded" column, instead of a stop with an undefined R that was dropped from the mean (Backtest: a fill below the stop is counted as a stop with undefined R #94).
  • Workflow inputs period, bars, split; JSON output gains windows, bars, split and the new per-slice fields.
  • Wiki 04 documents the statistics, the not-traded rule and the two-year dispatch; wiki 03 states that every calibration to date was judged on one window and points at --split; changelog.

Verified locally: ruff clean, 120 tests pass, and a 20-ticker end-to-end run with --period 5y --days 250 --bars 500 --split 2026-03-01 --grid --json renders the three windows and writes the JSON.

Related Issues

Closes #91
Closes #94

Checklist

  • Code follows existing project style and type annotations
  • Threshold changes include a one-line "why" comment in scan.py (no thresholds changed)
  • Offline unit tests pass (python -m pytest -q)
  • Random-walk false-positive rate stays under 5% (detection logic untouched)
  • Documentation updated in docs/wiki/
  • CHANGELOG.md updated

🤖 Generated with Claude Code

yanivil and others added 2 commits September 7, 2026 22:10
…fill through the stop is not traded

Every backtest summary now carries median R, standard deviation, total R, the
deepest drawdown of the cumulative R curve, a 95 % bootstrap interval of the
mean R and the drawdown exceeded in 5 % of resamples. The bootstrap resamples
scan months rather than trades because signals cluster in time.

--split YYYY-MM-DD reports the sessions before and from a date as separate
windows so a rule chosen on one window is judged on the other; --bars N limits
what each scan sees so a long download replays the nightly's 2y window.
Workflow inputs period, bars and split.

A next open at or below the stop was scored as a stop with an undefined R,
counted in the hit rate but dropped from mean R; it is now below_stop, not
traded, counted with the gaps.

Closes #91, closes #94.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…in the wiki

The first CI run of this branch took 11 minutes for 63 sessions with --grid,
in line with every earlier push run (6 to 12 minutes), not the "about 2
minutes" the wiki claimed. At about 5 s per session per profile a two-year
--grid replay needs roughly 90 minutes, the previous job timeout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@yanivil
yanivil merged commit de743ef into main Sep 7, 2026
5 of 6 checks passed
@yanivil
yanivil deleted the feat/backtest-split-and-intervals branch September 7, 2026 19:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Backtest: a fill below the stop is counted as a stop with undefined R Backtest: confidence intervals, drawdown and an in-sample / out-of-sample split

1 participant