Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
39 commits
Select commit Hold shift + click to select a range
98daf93
docs: design subscription-backed LLM integrations
ASRagab Aug 1, 2026
9ddfe82
feat: add subscription-backed LLM roles
ASRagab Sep 23, 2026
dfdc4d9
MAESTRO: capture environment for subscription live gate evidence
ASRagab Sep 27, 2026
c098c09
MAESTRO: Codex subscription live gate evidence - all tests PASS
ASRagab Sep 27, 2026
29180f3
MAESTRO: Complete Claude subscription live gate tests - 3 tests PASS …
ASRagab Sep 27, 2026
4dac04b
MAESTRO: Add judge canaries to subscription live gate evidence
ASRagab Sep 27, 2026
20568c9
MAESTRO: Record subscription backend negative auth test cases - 5 tes…
ASRagab Sep 27, 2026
3961f92
MAESTRO: Complete Phase-02 artifact isolation scan and populate Verdi…
ASRagab Sep 27, 2026
66f2732
test: record Codex and Claude subscription live gate evidence
ASRagab Sep 27, 2026
fd0ec53
MAESTRO: Mark Phase-02 Live Subscription Gates playbook tasks complete
ASRagab Sep 27, 2026
cae9368
MAESTRO: Design paid API fallback live trigger (fake subscription ada…
ASRagab Sep 27, 2026
f0d1723
MAESTRO: Add opt-in paid API fallback live gate tests and README gate…
ASRagab Sep 27, 2026
3521c3c
MAESTRO: Split paid fallback gate into its own README block and note …
ASRagab Sep 27, 2026
106cb0b
MAESTRO: Verify paid fallback gate skips offline and default CI stays…
ASRagab Sep 27, 2026
445de51
MAESTRO: Gate paid fallback live run on API keys reaching the agent
ASRagab Sep 27, 2026
25a14f6
MAESTRO: Re-gate paid fallback run on Anthropic key delivery via env …
ASRagab Sep 27, 2026
267bb4d
MAESTRO: Record paid fallback attempt 3 and re-gate on a working Anth…
ASRagab Sep 27, 2026
5c565e3
MAESTRO: Run paid API fallback live gate once for real (4 passed)
ASRagab Sep 27, 2026
927ec72
test: add opt-in paid API fallback live gate and record evidence
ASRagab Sep 27, 2026
8515b20
MAESTRO: Inventory subscription documentation gaps for U8 rollout docs
ASRagab Sep 27, 2026
54a5b46
MAESTRO: Document subscription versions, Claude scope, fallback, data…
ASRagab Sep 27, 2026
da8408d
MAESTRO: Document versioned generated evaluator runtime in evaluator-…
ASRagab Sep 27, 2026
e2b7724
MAESTRO: Document host backend flags in packaged command and skill do…
ASRagab Sep 27, 2026
4ac760c
MAESTRO: Record subscription live gate additions to local smoke-gates…
ASRagab Sep 27, 2026
778bc1b
MAESTRO: docs: document subscription backend versions, fallback, data…
ASRagab Sep 27, 2026
2f07d97
MAESTRO: docs: add subscription backends to Unreleased changelog section
ASRagab Sep 27, 2026
e4c430b
MAESTRO: Mark Phase-05 tasks 2-3 complete - branch review and pre-pus…
ASRagab Sep 27, 2026
6e4a8b2
MAESTRO: Record CI status in subscription backends audit document
ASRagab Sep 27, 2026
5987494
MAESTRO: Mark Phase-05 task 4 complete - PR opened with CI passing
ASRagab Sep 27, 2026
0b0a995
MAESTRO: Mark Phase-05 task 5 complete - CI watching and status recor…
ASRagab Sep 27, 2026
6bd3443
MAESTRO: Build R1-R14 subscription backends audit matrix
ASRagab Sep 27, 2026
f3eb01a
MAESTRO: Record U6 LiteLLM source-isolation check in audit
ASRagab Sep 27, 2026
4a9f78a
MAESTRO: Tighten audit evidence for R14 and the isolation search
ASRagab Sep 27, 2026
9978eec
MAESTRO: Record Verification Contract steps 1-5 offline gate results
ASRagab Sep 27, 2026
910a261
MAESTRO: Record Verification Contract step 6 CLI and evaluator checks
ASRagab Sep 27, 2026
1612d62
MAESTRO: Record Verification Contract step 7 leakage scan results
ASRagab Sep 27, 2026
279f022
MAESTRO: Close subscription backend audit gaps with offline coverage
ASRagab Sep 27, 2026
4fb98a7
MAESTRO: test: audit subscription backends U1-U7 and record offline g…
ASRagab Sep 27, 2026
21e5f39
chore: merge main into subscription backend branch
ASRagab Sep 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,9 @@ node_modules/

# miscellaneous files
seed.txt
docs/
docs/*
# Exception: allow verification evidence to be committed
!docs/verification/
smoke_outputs/
artifacts/
integration_runs/
Expand All @@ -30,3 +32,5 @@ PROMPT.md
run.log
ROADMAP*
runs/
.codegraph/
todos/
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Phase 01: Completion Audit and Offline Verification Gates

Verify that units U1-U7 of `docs/plans/2026-09-22-1841-feature-subscription-backends-plan.md` are actually complete on branch `feat/codex-claude-subscription`, and prove it by running the network-free half of the plan's Verification Contract (steps 1-7). Output is a requirement-by-requirement audit matrix (R1-R14 mapped to code + tests + evidence), a green offline suite, and fixes for any gap the audit surfaces. Nothing in this phase consumes subscription quota or needs credentials; every task runs on this machine with what is already installed. This phase is the foundation for the live gates in Phase 02: never run quota-consuming gates on a branch whose offline contract is not green.

## Tasks

<!-- MAESTRO:MODEL tier="high" effort="high" reason="Auditing whether 14 security-sensitive requirements are truly satisfied requires reading adapter code, tests, and the plan together and spotting silent gaps. A superficial pass here would let an incomplete adapter reach the live gates and burn quota on a known-bad branch." -->

- [x] Build the requirement audit matrix for U1-U7. Read the plan (`docs/plans/2026-09-22-1841-feature-subscription-backends-plan.md`, sections "Requirements", "Implementation Units", "Verification Contract") and the design spec (`docs/superpowers/specs/2026-08-01-codex-claude-subscription-backends-design.md`). For each requirement R1-R14, locate the implementing code under `src/optimize_anything/llm_backends/` (`base.py`, `schema.py`, `litellm_backend.py`, `codex_backend.py`, `claude_backend.py`, `coordination.py`, `fallback.py`, `factory.py`, `provenance.py`), `src/optimize_anything/evaluator_runtime.py`, `cli.py`, `cli_optimize.py`, `cli_tools.py`, `spec_loader.py`, `llm_judge.py`, `evaluator_generator.py`, and the matching test(s) in `tests/test_llm_backend_contract.py`, `test_codex_backend.py`, `test_claude_backend.py`, `test_llm_fallback.py`, `test_llm_coordination.py`, `test_llm_factory.py`, `test_evaluator_runtime.py`, `test_cli.py`, `test_spec_loader.py`, `test_plugin_regression.py`. Write the result to `docs/verification/subscription-backends-audit.md` with YAML front matter (`type: report`, `title: Subscription Backends U1-U7 Completion Audit`, `created: 2026-09-26`, `tags: [subscription-backends, audit, verification]`, `related: ['[[Subscription-Live-Evidence]]']`). One table row per R-ID: unit, implementing symbol(s) with `file:line`, covering test(s) with `file::test_name`, status (`covered` / `gap` / `partial`), notes. Below the table, a "Gaps" section listing each `gap`/`partial` row with the concrete missing behavior. Also check each unit's listed "Test scenarios" bullet in the plan against actual test names and note any scenario with no test.

- [x] Verify the U6 source-isolation claim mechanically: search `src/optimize_anything/` for `import litellm` and `litellm.` and confirm every hit lives in `llm_backends/litellm_backend.py` (generated-template strings in `evaluator_generator.py` count as a violation of R13 if they emit a LiteLLM import into judge/composite scripts; deterministic templates are exempt). Also confirm `codex_backend.py` and `claude_backend.py` never pass prompt text through argv (search for `subprocess`, `argv`, `args =` construction and confirm prompts go through stdin/request bodies). Append findings to the "Gaps" section of `docs/verification/subscription-backends-audit.md`; if isolation holds, record the exact search commands and hit list as evidence.

<!-- MAESTRO:MODEL tier="low" effort="low" reason="Running existing test and check commands and recording their output is mechanical. No design judgment is needed until a failure appears, and failures are handled by the fix task below." -->

- [x] Run Verification Contract steps 1-5 offline and record results:
- `uv run pytest tests/test_llm_backend_contract.py tests/test_codex_backend.py tests/test_claude_backend.py tests/test_llm_fallback.py tests/test_llm_coordination.py tests/test_llm_factory.py tests/test_evaluator_runtime.py -v` (step 1, focused unit tests)
- `uv run pytest -m "not integration"` (step 2, full offline suite; note count of passed/skipped/failed)
- `uv run python scripts/check.py --skip-smoke` (step 3)
- `uv run python scripts/smoke_harness.py --budget 1` (step 4)
- `uv run python scripts/score_check.py` (step 5)
- Save raw output under `.maestro/playbooks/Initiation/Working/offline-gates/` (one file per command) and add a "Offline Gate Results" section to `docs/verification/subscription-backends-audit.md` with command, exit code, and the decisive summary line for each. Any failure is a hard stop for this task: record it, do not mark the gate as passing.

- [x] Run Verification Contract step 6 (CLI help and generated-evaluator compilation/contract checks):
- `uv run optimize-anything optimize --help`, `score --help`, `analyze --help`, `validate --help`, `generate-evaluator --help`; confirm each shows its backend flags (`--proposer-backend`, `--judge-backend`, `--analysis-backend`, `--subscription-concurrency`, `--no-api-fallback`, `--openai-api-fallback-model`, `--anthropic-api-fallback-model`) and that `--help` output contains no provider secrets or account identity.
- Generate one `judge` and one `composite` evaluator via `uv run optimize-anything generate-evaluator` with `--backend codex`, `--backend claude`, and no backend flag; `python -m py_compile` each; assert none contain `import litellm` and each calls `optimize_anything.evaluator_runtime`. Generate one deterministic (command-style) evaluator and confirm it remains standalone (no `evaluator_runtime` import).
- Record commands and outcomes in the "Offline Gate Results" section.

- [x] Run Verification Contract step 7 (secret/prompt leakage assertions) using the offline fakes already in the suite: `uv run pytest tests/test_codex_backend.py tests/test_claude_backend.py tests/test_llm_coordination.py tests/test_llm_fallback.py -k "argv or leak or secret or sentinel or scrub or prompt or cache" -v`. Then run one fake-backed end-to-end optimize with `--run-dir` under `Working/leak-scan/` (use the existing echo evaluator `examples/evaluators/echo_score.sh` and a sentinel string such as `LEAKSENTINEL-9f3a` set as `OPENAI_API_KEY` and `ANTHROPIC_API_KEY`) and grep the run directory, any coordination state directory, `fitness_cache`, and stdout for the sentinel and for the objective/candidate text in cache keys. Record hit counts (expected: zero for the sentinel in any retained artifact or cache key) in the audit report.

<!-- MAESTRO:MODEL tier="high" effort="high" reason="Closing gaps in auth-class checks, isolation, or fallback eligibility changes security behavior. Fixes must follow the plan's stop conditions (never weaken isolation to make a test pass), which requires careful reasoning rather than a quick patch." -->

- [x] Fix every `gap` / `partial` row and every failed gate recorded in `docs/verification/subscription-backends-audit.md`. For each fix: write or extend the failing test first in the matching `tests/test_*.py` (test must encode which R-ID it protects in its docstring), then implement the minimum change in the owning module, matching existing style. Do not touch adjacent code. Do not relax isolation, auth-class verification, or fallback eligibility (`_ELIGIBLE` in `fallback.py` is `BackendUnavailable, AuthenticationError, RateLimitError, QuotaExceeded` and must stay that way; `Timeout`, `Cancelled`, `InvalidResponse`, `ConfigurationError` never fall back). If a gap cannot be closed without a product decision, leave it documented under a "Deferred" heading with the reason; do not silently mark it covered. If the audit found zero gaps, state that explicitly in the report and skip to the next task.

- 2026-09-27: Added offline coverage for R1-R12, routed all generated judge/composite evaluators through the installed runtime (R13), recorded API proposer provenance and aggregation, and improved the TOML conflict diagnostic. Offline pytest: 542 passed, 18 deselected. R4 GEPA cache identity and R6 invalid-TOML override semantics remain documented under Deferred in the audit pending product decisions.

- [x] Re-run the full offline contract after fixes and update the report: `uv run pytest -m "not integration"`, `uv run python scripts/check.py --skip-smoke`, `uv run mypy src` (if configured in `pyproject.toml`/pre-commit), and `uv run pre-commit run --all-files` if `.pre-commit-config.yaml` exists. Update every affected row in `docs/verification/subscription-backends-audit.md` to `covered` with the new test name, refresh the "Offline Gate Results" section with final exit codes, and add a one-line "Verdict" at the top: `Offline contract green: yes/no` with the date. Commit code fixes and the report with message `test: audit subscription backends U1-U7 and record offline gates` (do not commit the `Working/` scratch directory or the staged `.gitignore` change unless it is intentional; inspect `git diff --cached .gitignore` first and keep it only if it ignores generated artifacts).

- 2026-09-27: Final offline rerun passed: 542 offline tests, 88 focused backend tests, check.py, smoke, score, mypy, pre-commit, generated-evaluator compilation, and leakage checks. The report marks 12 requirements covered and keeps R4/R6 partial under Deferred pending product decisions.
32 changes: 32 additions & 0 deletions .maestro/playbooks/Initiation/Phase-02-Live-Subscription-Gates.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Phase 02: Live Subscription Gates (Codex and Claude)

Run the quota-consuming half of the Verification Contract (steps 8-10) on this machine, where both Codex CLI and Claude Code are already authenticated via subscription. Each gate runs with deliberately invalid API-key sentinels and `--no-api-fallback`, so a pass proves the request really went through the saved subscription and never silently fell back to a paid API. Evidence is recorded per the plan: versions, OS, auth class, requested/actual model, isolation assertions, pass/fail, and no account identity or artifact prompts. Do not weaken isolation or auth-class checks to make a live test pass; a fix must preserve R3, R8, R10, R11.

## Tasks

<!-- MAESTRO:MODEL tier="low" effort="low" reason="Capturing tool versions and OS details is a mechanical prerequisite for the evidence record. Nothing here needs judgment." -->

- [x] Capture the environment for the evidence record and start `docs/verification/subscription-live-evidence-2026-09-26.md` with YAML front matter (`type: report`, `title: Subscription Live Gate Evidence 2026-09-26`, `created: 2026-09-26`, `tags: [subscription-backends, live-gate, codex, claude, evidence]`, `related: ['[[Subscription-Backends-U1-U7-Completion-Audit]]']`). Record: `sw_vers` / `uname -srm`, `uv --version`, `uv run python --version`, `uv run python -c "import openai_codex, importlib.metadata as m; print(m.version('openai-codex'))"` (install the extra first with `uv sync --extra codex` if the import fails), `codex --version`, `claude --version`, `codex login status` (record only the auth class/method line; redact any email, account id, or plan name), and the Claude auth class as reported by the adapter's own preflight (`uv run python -c "from optimize_anything.llm_backends.claude_backend import ClaudeCliBackend; s=ClaudeCliBackend().preflight(); print(s.auth_class, s.auth_source)"`). Confirm the Claude CLI version meets the minimum stated in `install.md` (2.1.278 or newer). Never paste account identity into the report.

- [x] Run the Codex live gates (Verification Contract step 8) with API fallback disabled:
- `OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1 OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run pytest tests/test_subscription_live.py -k codex -v -s` (covers structured completion, seedless budget-1 proposer optimize, and generated judge evaluator).
- Then an explicit CLI run with a persisted run directory so artifacts can be inspected: `OPENAI_API_KEY=deliberately-invalid-live-gate uv run optimize-anything optimize --no-seed --objective "Write a concise friendly greeting." --budget 1 --proposer-backend codex --no-api-fallback --evaluator-command bash examples/evaluators/echo_score.sh --run-dir .maestro/playbooks/Initiation/Working/live-codex/`.
- Save stdout/stderr under `Working/live-codex/`. Add a "Codex" section to the evidence report with: each test name and pass/fail, requested vs actual backend and model from the provenance JSON, `auth_class` and `auth_source` (expected `subscription` / `chatgpt`), `fallback_used` (expected false), wall time, and usage if reported.

- [x] Run the Claude live gates (Verification Contract step 9) under the scrubbed child environment with API fallback disabled:
- `OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1 OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run pytest tests/test_subscription_live.py -k claude -v -s`.
- Then the explicit CLI run: `ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run optimize-anything optimize --no-seed --objective "Write a concise friendly greeting." --budget 1 --proposer-backend claude --no-api-fallback --evaluator-command bash examples/evaluators/echo_score.sh --run-dir .maestro/playbooks/Initiation/Working/live-claude/`.
- Note: this playbook itself may be executing inside a Claude Code session, so `CLAUDECODE` and related parent-agent variables will be set in the parent environment. The adapter is required (R10) to scrub them; a pass here is direct evidence of that scrub. If the gate fails with a nested-session or auth error, inspect `claude_backend.py`'s scrub list against the current `claude --help` output and the plan's R10 before changing anything, and treat any fix as a security change (test first, in `tests/test_claude_backend.py`).
- Save output under `Working/live-claude/`. Add a "Claude" section to the evidence report with the same fields as Codex (expected `auth_class=subscription`, `auth_source=claude_subscription`, `fallback_used=false`).

- [x] Run one built-in judge live canary per provider (Verification Contract step 10), only after the proposer-only gates above passed. Use the existing `score` subcommand so the judge role (not the proposer) is exercised: `OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run optimize-anything score examples/seed.txt --judge-backend codex --no-api-fallback --objective "Score clarity"` and the same with `--judge-backend claude` (pick any small existing artifact under `examples/` if `seed.txt` is absent; check `ls examples` first). Confirm the score JSON carries `llm_provenance` with `role=judge`, `actual_backend` matching the provider, and `auth_class=subscription`. Record both in a "Judge canaries" section of the evidence report.

- [x] Run the no-fallback negative auth cases with fakes, so the evidence file shows fail-closed behavior alongside the live passes: `uv run pytest tests/test_codex_backend.py tests/test_claude_backend.py -k "api_key or wrong_auth or logged_out or rejected or unavailable" -v` and `uv run pytest tests/test_llm_fallback.py -k "timeout or cancel or invalid or configuration" -v`. Record the test names and results under a "Negative cases (fakes)" section. If any expected scenario from the plan's U4/U5 test-scenario lists (logged-out preflight exits before a model request; API-key auth rejected as subscription; wrong auth class) has no test, add it in the matching test file with a fake client and re-run.

- [x] Perform the live artifact secret/prompt scan (Verification Contract step 7 applied to real runs). Over `Working/live-codex/`, `Working/live-claude/`, all saved stdout/stderr, and any coordination/state directory the run created (search `coordination.py` for the directory naming pattern and locate it under the run dir or temp), grep for: `deliberately-invalid-live-gate`, any `sk-` prefixed token, `@` email patterns, `account`, the objective text `Write a concise friendly greeting.` inside cache keys or coordination state files (it is expected in stdout as the run objective, but never in cache identity or coordination state), and `CLAUDECODE`. Record each pattern with its hit count and location in an "Isolation and leakage assertions" section. Expected: zero hits for secrets/identity in any retained artifact, coordination state, or cache key. Also confirm the Claude child ran with no tools/MCP by checking the adapter's argv construction assertions in the fake tests and noting the flag set actually detected on this machine's CLI version.

- [x] Finalize the evidence report: add a "Verdict" block at the top with one line per gate (Codex structured, Codex budget-1, Codex generated-evaluator, Claude structured, Claude budget-1, Claude generated-evaluator, Codex judge canary, Claude judge canary, leak scan) marked PASS/FAIL, followed by the plan's required fields summary (versions, OS, auth class, requested/actual model, isolation assertions). Re-read the report once and delete any line containing an email, account id, subscription plan name, or artifact prompt text beyond the fixed objective string. Commit the evidence report (and any test additions) as `test: record Codex and Claude subscription live gate evidence`. Do not commit `Working/`.

## Manual Follow-Up (not executed by Auto Run)

- If either live gate FAILED for reasons outside the adapter (Codex or Claude service outage, expired subscription login), re-authenticate with `codex login` or `claude auth login` and re-run this phase. The playbook must not attempt login itself (R3).
Loading
Loading