Skip to content

fix(cpr-ai): prompt v3 - drop the veto examples the backtest refutes - #194

Merged
DoRmAmMu1997 merged 1 commit into
mainfrom
feat/cpr-ai-prompt-v3
Oct 1, 2026
Merged

DoRmAmMu1997 merged 1 commit into
mainfrom
feat/cpr-ai-prompt-v3

Conversation

@DoRmAmMu1997

Copy link
Copy Markdown
Owner

Why

Prompt v2 gave Codex four example red flags for vetoing a trend-day candidate, but the backtest takes every candidate, so none of the four had ever been tested. I scored each on the 264 five-year candidates (SELL-opposite premium points after costs; baseline PF 1.69):

v2 veto example (as measured) Trades Result
Whipsaw: 4+ closes switching side of the running VWAP before the candidate 96 PF 1.71 vs 1.67 for the rest. Vetoing would give up 908 of 2,542 points
Steep push into R2/S2 or a prior-day extreme 6 Five won, about +103 points each: the best group
One climactic bar plus a rejection wick 7 PF 0.81 (−34 points)
Bullish, confluence exactly 2, expiry day 11 PF 1.01 (+4 points)

The size of the candidate bar points the same way. Candidates whose own bar spanned at least 0.25 × ATR5 had PF 2.98 (36 trades). Candidates whose bar alone lifted the range past 1.0 × ATR5 had PF 1.86, against 1.44 for the rest.

Today's paper entry (the first under v2, a single 60-point bar) is the kind of candidate the old climactic-bar example could have led Codex to veto.

What changes

  • cpr-trend-day-rider-v3:
    • drops the whipsaw and R2/S2 veto examples;
    • keeps the two rare, roughly break-even ones (the climactic one now says that the wick is the warning, not the size of the bar);
    • tells Codex explicitly not to veto for a big candidate bar, a choppy morning or a push into R2/S2 / a prior-day extreme;
    • quotes the evidence.
  • ADR-0019: a dated update (2026-10-01) with:
    • the measurement definitions and results;
    • the 29/30 Sep no-candidate sessions (range 0.90 / 0.75 × ATR5 by 13:30);
    • the 1 Oct entry.
  • README and LLD: version references.

Unchanged: the host gate, the VWAP stop, the SELL expression, Codex's veto-only role, and the premise-exit guidance from v2.

Verification

  • New test, written first: test_prompt_veto_examples_are_only_the_ones_the_backtest_does_not_refute. It failed against v2 and passes against v3.
  • pytest "Tests/Signal Generators/CPR AI Agent": 132 passed.
  • Static checks: ruff, mypy (82 files), compileall, and pre-commit on the changed files all pass.
  • Not run locally: the full master and other suites. They were skipped because the trading laptop is running a live paper session; CI runs them. This change is prompt text only, and nothing outside the CPR AI suite pins it.

🤖 Generated with Claude Code

Prompt v2 listed four example red flags for vetoing a trend-day candidate,
but the backtest takes every candidate, so none had been tested. Scored on
the 264 five-year candidates (SELL-opposite premium points, baseline PF 1.69):

- whipsaw (4+ closes switching side of VWAP before the candidate): 96 trades,
  PF 1.71 vs 1.67 for the rest -- vetoing would give up 908 of 2,542 points;
- a steep push into R2/S2 or a prior-day extreme: 6 trades, five won,
  about +103 points each -- the best group;
- one climactic bar plus a rejection wick: 7 trades, PF 0.81;
- a bullish expiry-day candidate scraping two confluence factors: 11 trades,
  PF 1.01.

Large candidate bars were also better, not worse: own bar >= 0.25 x ATR5
had PF 2.98 (36 trades), and a bar that alone lifted the range past
1.0 x ATR5 had PF 1.86 vs 1.44.

cpr-trend-day-rider-v3 drops the whipsaw and R2/S2 examples, keeps the two
rare break-even ones (the climactic one now says that the wick is the warning,
not the bar size), tells Codex not to veto for a big bar, a choppy morning or
a push into R2/S2, and quotes the evidence. The host gate, stop, and
veto-only role are unchanged. ADR-0019 gains a dated update. That update also
records the 29/30 Sep no-candidate sessions and the 1 Oct entry: the first
under v2, on a single 60-point bar.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@DoRmAmMu1997
DoRmAmMu1997 merged commit 2c3f4c7 into main Oct 1, 2026
7 checks passed
@DoRmAmMu1997
DoRmAmMu1997 deleted the feat/cpr-ai-prompt-v3 branch October 1, 2026 07:43
DoRmAmMu1997 added a commit that referenced this pull request Oct 1, 2026
At the operator's direction. CI measured 74.6% on main at 2c3f4c7 (PR #194) and
74.7% on this PR, identically on Python 3.12 and 3.13, so 74.0 keeps roughly
0.6pp of margin (about 175 statements and branches) -- the same deliberate
pressure the 73.0 floor applied at 0.8pp. The floor only moves up, and only on
a CI figure.

Updated everywhere it is pinned: pyproject fail_under (with its history), the
repository-policy assertion, the threshold script's docstring, CLAUDE.md,
AGENTS.md, README and the testing-and-ci LLD. Negative-tested: with pyproject
back at 73.0 the policy test fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant