Repository navigation
feat: subscription-backed Codex and Claude LLM backends - #7
Merged
Merged
Conversation
…with subscription auth confirmed
…ct section - Scan 17 files in Working/live-codex (8) and live-claude (9) for leakage - Zero hits on: deliberately-invalid-live-gate, sk- tokens, @emails, account word, objective in cache, CLAUDECODE - Verify Claude argv: --safe-mode, --tools "", --disable-slash-commands, empty MCP config, env scrubbing - Populate Isolation and leakage assertions section with scan results - Complete Verdict section with gate-by-gate results and summary - Mark task 6 complete in Phase-02-Live-Subscription-Gates.md
Finalize Phase-02 evidence report with Verdict block at top containing: - Gate-by-gate PASS/FAIL results for all 10 subscription live gates - Required fields summary (versions, OS, auth class, models, isolation assertions) - Clean of any email, account ID, or subscription plan identity - All 10 gates PASS: Codex/Claude structured completions, budget-1 proposer runs, generated evaluator tests, judge canaries, and artifact isolation scan Mark task 6 (Finalize evidence report) as complete in Phase-02 playbook.
…pter via factory seam) - Add tests/test_api_fallback_live.py header recording the chosen trigger: option (a), a test-only fake subscription adapter monkeypatched into the lazily imported factory seam, with the real LiteLLMBackend as the API leg - Record why option (b) (CODEX_HOME/PATH) was rejected and the invariants confirmed by an offline probe (fail in complete(), not preflight()) - Note next-task corrections in the Phase-03 playbook: no fallback-model defaults exist, fallback_used maps to result.fallback, keep real keys - Mark the design task complete in Phase-03-Paid-API-Fallback-Gate.md
… docs Two parametrized cases per vendor behind OPTIMIZE_ANYTHING_RUN_PAID_FALLBACK_LIVE=1 (integration-marked): forced eligible subscription failure falls back to the same-vendor API with billing warning before dispatch and a sticky per-role circuit; --no-api-fallback raises BackendUnavailable with zero API calls. Verified offline with a stubbed litellm dry run (6 dispatches) and three mutations that each turn the targeted test red.
…key-scan pattern Copy-pasting the subscription live gate block no longer also runs the paid API fallback gate, which bills the OpenAI and Anthropic API accounts. The Phase-03 playbook note for Task 5 now says to scan for exact key values plus a key-shaped pattern, because a bare "sk-" substring matches ordinary prose.
…file Attempt 2 of the Phase-03 live gate did not run (0 billed calls). The human step was ticked, but this agent's Maestro per-agent env holds only OPENAI_API_KEY, and a lone OpenAI key bills the 3 codex calls before the claude case fails. This commit includes that human tick. The attempt-1 advice to put ANTHROPIC_API_KEY in this claude-code agent's per-agent env was unsafe: Claude Code in -p mode prefers that key over subscription login, so every agent turn would bill the API. The new gate asks for a chmod 600 env file outside the repo, loaded only into the pytest child via uv run --env-file. Also records that the inherited ANTHROPIC_BASE_URL (lean-ctx proxy on 127.0.0.1:4444) would route LiteLLM's paid Anthropic leg through the proxy. The next attempt runs under env -u ANTHROPIC_BASE_URL after a boolean-only pre-check that must resolve https://api.anthropic.com.
…ropic key Attempt 3 of the Phase-03 live gate ran once for real and stopped under -x (1 passed, 1 failed). It billed 3 OpenAI calls (51 tokens) and 0 Anthropic calls. The codex case passed every assertion: the sticky role circuit, the billing warning printed before dispatch, and same-vendor provenance. The claude case failed at its first API dispatch with AuthenticationError. A free GET /v1/models call shows Anthropic rejecting the key in the env file itself (401 "API key is invalid."). The key is well-formed and no proxy is involved. That makes this an environment problem, so no code changed and a new human step asks for a working key. This commit includes the human's tick of the previous env-file step. The next attempt first validates both keys with free model-lookup requests, so a bad Anthropic key can no longer bill the OpenAI leg before failing.
Attempt 4 of the Phase-03 live gate passed on its only run: 4 passed in 7.79s, exit 0, with no code change. It billed 6 calls (3 OpenAI, 3 Anthropic); counting attempt 3, the task billed 6 OpenAI and 3 Anthropic calls in total. Zero-cost checks came first. ANTHROPIC_API_KEY was unset in the agent's own env, so agent turns stayed on the subscription. The key file sat outside the repo at mode 600. The free validity probe returned 200 200 https://api.anthropic.com with ANTHROPIC_BASE_URL stripped. The log was key-scanned before display, with zero exact-value, key-shaped, or sk- hits. For both vendors the eligible case shows timeline [sub judge, api, api, sub score, api] and billing warning [true, false, true]. The events report requested_backend=<provider>, actual_backend=api, auth_class=api, and fallback_source=<provider>. Usage was 51 tokens on openai/gpt-5.6-luna and 60 on anthropic/claude-sonnet-5. Both --no-api-fallback cases recorded one subscription attempt and no events. This commit includes the human's tick of the Anthropic key replacement step.
The opt-in gate itself (tests/test_api_fallback_live.py and its README entry) landed earlier in f0d1723 and 3521c3c and is unchanged here. This commit carries the evidence, following the Phase-02 precedent 66f2732. The live evidence report has a new "Paid API fallback (step 11)" section for the only full run: 4 passed in 7.79s. For both vendors the forced eligible subscription failure fell back to the same vendor's API (codex to openai/gpt-5.6-luna, claude to anthropic/claude-sonnet-5) with requested_backend=<provider>, actual_backend=api, and auth_class=api. The billing warning printed before each dispatch that opened a circuit ([true, false, true]), and the role circuit stayed sticky. With --no-api-fallback, BackendUnavailable was raised after one subscription attempt, with no API dispatch and no events. The run billed 6 calls and 111 tokens in total. A key scan of the three live-fallback logs and the report found zero exact-value, key-shaped, or bare-prefix hits. To get the report's bare count to zero, one Phase-02 table label was reworded; its finding is unchanged. Working/ is not committed. This commit also ticks Phase-03 Task 5 in the playbook.
… handling, and removal in install.md and README Fills U8 DoD gaps: supported/tested versions and macOS-only platform, experimental local-only Claude scope with policy caveat, eligible vs never-fallback categories and sticky per-role circuit, data handling and content-free provenance, preflight/backend plan and concurrency warning, and steps to return to API defaults or remove the SDK/plugins. Co-Authored-By: Claude <noreply@anthropic.com>
…cookbook Add R13 subsection: thin wrapper for --judge-backend codex|claude, min_runtime_contract_version 1, unchanged JSON-lines contract plus llm_provenance side info, and actionable runtime error codes. PROTOCOL.md left unchanged (no contradiction).
…cs (R7) - commands/optimize.md: LLM-judge step names --judge-backend claude for subscription mode - commands/validate.md: explicit claude selector for Claude Code host, API default otherwise - skills/generate-evaluator: unknown hosts keep --judge-backend api default - skills/optimize-prompt: host subscription substitution for analyze/score/optimize
… and release checklist docs
… handling, and removal Final Phase-04 verification. Contract tests and scripts/check.py --skip-smoke pass. The score regression gate failed on evaluator-cookbook.md (0.8905 vs 0.9088 baseline) after the versioned-runtime section was added: bash comment lines in the example block were counted as headings and the section added length. Fixed in the doc (comments moved into prose, wording tightened); baseline and evaluator left unchanged. Score now 0.8989, within tolerance.
Map every plan requirement R1-R14 to implementing symbols (file:line), covering tests (pytest node IDs), and covered/partial/gap status. Check each U1-U7 plan test scenario against real tests and list gaps. Result: R5 and R14 covered, R13 gap (default API judge/composite templates still embed LiteLLM), all other rows partial. Preserve the Phase-05 CI record below the audit.
Record the exact grep commands and hit list. Runtime code holds: every LiteLLM import or call is in llm_backends/litellm_backend.py. Generator templates fail: evaluator_generator.py:452 emits 'from litellm import' into judge scripts and composite runs that script (R13 gap). The playbook's literal patterns miss 'from litellm import'. Codex adapter spawns no process; Claude adapter sends prompt on stdin with a constant list argv.
Cite the ci.yml:45 gate for the networked integration job instead of its date, and state the push-to-main interpretation. Correct the preserved Phase-05 claim: the judge-matrix job was skipped, not passed, in PR run 35814308325. Widen the LiteLLM search beyond .py files (no new hits). Spell out the R2 claude live node IDs and name representative tests/test_cli.py node IDs for R5.
All 5 offline contract steps pass with zero failures: - Step 1: 37 backend-focused tests pass - Step 2: 443 offline tests pass, 18 deselected - Step 3: 447 tests pass, 14 skipped; all gates pass - Step 4: smoke harness passes with budget 1 - Step 5: score check passes (3 skill docs) Verdict: Offline contract green (2026-09-27) Phase-01 Task 3 complete.
Step 6 offline contract checks pass: - All CLI commands (optimize, score, analyze, validate, generate-evaluator) show backend flags and no secrets leaked - Generated evaluators compile successfully: - judge --backend api: contains from litellm (R13 gap as documented) - judge --backend codex: uses evaluator_runtime (correct) - judge --backend claude: uses evaluator_runtime (correct) - command evaluator: standalone bash (no imports) Evaluators saved to .maestro/playbooks/Initiation/Working/step-6-evaluators/ Phase-01 Task 4 complete.
Step 7 leakage detection passes: - Secret/leakage test suite: 3 passed - Checks: argv, scrub, prompt isolation - End-to-end optimize run with sentinel (LEAKSENTINEL-9f3a): - Sentinel hits in run artifacts: 0 (expected) - Objective text in caches: 0 (expected) - Run artifacts: seed.txt, best_artifact.txt, summary.json, etc. (no secrets) - Verification result: No credential leakage detected All Verification Contract steps 1-7 complete and passing offline. Phase-01 Task 5 complete.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Added subscription-backed Codex and Claude LLM integrations for proposer, judge, analysis, score, and validation roles. Codex uses OpenAI's ChatGPT subscription, Claude uses Anthropic's subscription. Both include conservative same-vendor fallback to API models, run-scoped coordination, isolation guarantees, and provenance tracking.
Design and Implementation
Verification
Live Evidence Report
Subscription Live Evidence 2026-09-26
Verdict Summary:
CI Status
Verification Contract Steps
Features Added
--proposer-backend,--judge-backend,--analysis-backend,--subscription-concurrency,--no-api-fallback,--openai-api-fallback-model,--anthropic-api-fallback-modelevaluator_runtimefor backward-compatible schema changescodexoptional extrasImportant Notes
Claude support is experimental and local-only. Claude completion is only available via Claude Code on macOS, not in CI or remote environments.
Reviewer Focus
Security-relevant files:
src/optimize_anything/llm_backends/codex_backend.py— Codex integration with subscription isolationsrc/optimize_anything/llm_backends/claude_backend.py— Claude subprocess isolation, environment scrubbing, capability limitingsrc/optimize_anything/llm_backends/fallback.py— Same-vendor fallback policy, error classificationsrc/optimize_anything/llm_backends/coordination.py— Run-scoped coordination, provider slots, event recordingKey invariants: