Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -5,14 +5,14 @@
},
"metadata": {
"description": "Claude Code plugin for optimizing text artifacts using gepa",
"version": "0.5.1"
"version": "0.6.0"
},
"plugins": [
{
"name": "optimize-anything",
"source": "./",
"description": "Optimize any text artifact using gepa — prompts, code, configs, skills",
"version": "0.5.1",
"version": "0.6.0",
"homepage": "https://github.com/ASRagab/optimize-anything",
"repository": "https://github.com/ASRagab/optimize-anything",
"keywords": ["optimization", "gepa", "prompts", "evaluator", "llm-judge"],
Expand Down
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "optimize-anything",
"version": "0.5.1",
"version": "0.6.0",
"description": "Optimize any text artifact using gepa — prompts, code, configs, skills",
"keywords": ["optimization", "gepa", "prompts", "evaluator"],
"author": {
Expand Down
2 changes: 1 addition & 1 deletion .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "optimize-anything",
"version": "0.5.1",
"version": "0.6.0",
"description": "Optimize prompts and other text artifacts with measured evaluator feedback",
"author": {
"name": "optimize-anything contributors"
Expand Down
5 changes: 4 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,8 @@ node_modules/
# miscellaneous files
seed.txt
docs/*
# Exception: allow verification evidence to be committed
# Exceptions: track active smoke instructions and verification evidence
!docs/smoke-gates.md
!docs/verification/
smoke_outputs/
artifacts/
Expand All @@ -34,3 +35,5 @@ ROADMAP*
runs/
.codegraph/
todos/
.maestro
.cursor/
9 changes: 8 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,14 @@

## Unreleased

## v0.6.0 - 2026-09-27

### Plugin and documentation repairs
- Made shared skills use the bundled runtime in plugin installs, including the compatible Codex SDK extra
- Fixed the default installer on macOS Bash 3.2 and repaired dataset and LiteLLM evaluator examples
- Clarified subscription backend prerequisites, API fallback behavior, evaluator generation, and current smoke gates across active guides
- Aligned Python, Claude Code, and Codex plugin metadata at 0.6.0

### Subscription-backed LLM backends
- Added support for Codex and Claude as subscription-backed proposer, judge, analysis, score, and validation roles
- Integrated optional `codex` extras with secure auth isolation and no API key exposure on host
Expand All @@ -14,7 +22,6 @@
- Run-scoped coordination with same-vendor conservative fallback (Codex falls back to OpenAI, Claude to Anthropic)
- Provenance field (`llm_provenance`) in evaluation results capturing backend, auth class, and fallback decisions
- Opt-in live gates for subscription backend verification with no impact on default CI or offline workflows
- No version bump to pyproject.toml, plugin metadata, or Codex plugin in this release

## v0.5.1 - 2026-07-28

Expand Down
11 changes: 8 additions & 3 deletions EXAMPLES.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,14 @@
# EXAMPLES

Worked examples for optimize-anything v2.
Illustrative CLI examples for optimize-anything v2. Replace seed, evaluator,
and dataset paths with your files. These commands assume the global CLI installer
from [install.md](install.md); use `uv run optimize-anything` from a source
checkout. The shown API models can incur provider charges. `--budget` counts
evaluator calls, not dollars, and an iteration can exceed the requested count.

## Result JSON shape (current contract)

Optimization output examples should follow this structure:
An abbreviated optimization output has this structure:

```json
{
Expand Down Expand Up @@ -165,7 +169,8 @@ optimize-anything optimize strategy.md \
--evaluator-command bash eval_unbounded.sh \
--model openai/gpt-5.6-sol \
--objective "Maximize reward" \
--score-range any
--score-range any \
--budget 20
```

Use when your evaluator emits finite scores outside `[0,1]`.
11 changes: 8 additions & 3 deletions WALKTHROUGH.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,12 @@ Step-by-step v2 workflow for optimize-anything.

- Python 3.10+
- `uv`
- API key(s) for any LLM providers you plan to use
- API key(s) for the models in this walkthrough

These commands use API models and can incur provider charges. `--budget` limits
evaluator calls approximately, not dollars; GEPA may finish an iteration after
the limit is reached. For local subscription backends and billed fallback
controls, see [install.md](install.md).

## Step 1: Install

Expand Down Expand Up @@ -49,8 +54,8 @@ chmod +x evaluators/eval.sh
## Step 4: Test evaluator contract

```bash
echo '{"_protocol_version":2,"candidate":"test"}' | python evaluators/eval.py
# expected: JSON with required "score"
echo '{"_protocol_version":2,"candidate":"test"}' | uv run python evaluators/eval.py
# expected: JSON with required "score"; the default judge makes an API call
```

## Step 5: Baseline score
Expand Down
122 changes: 122 additions & 0 deletions docs/smoke-gates.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,122 @@
# Smoke Gates

Repeatable smoke checks for `RB-009` and `RB-010`.

## Prerequisites
- Dependencies installed: `uv sync`
- `OPENAI_API_KEY` for the default `openai/gpt-5.6-sol` proposer, or set `OPTIMIZE_ANYTHING_MODEL` and provide that model's API credential

## Run Smoke Harness (RB-009)

One command:

```bash
uv run python scripts/smoke_harness.py --budget 1
```

What it does:
- Creates a temporary seed file and temporary evaluator command script.
- Runs CLI optimize smoke (`optimize_anything.cli` path) with a temporary command evaluator.
- Asserts that the CLI summary has a non-empty `best_artifact`, positive `total_metric_calls`, `top_diagnostics`, and `score_summary`.
- Runs `generate-evaluator` and checks that it emits a script with a shebang.
- Runs `intake` and checks that its JSON includes `execution_mode`.
- Saves logs/artifacts under `smoke_outputs/smoke-<timestamp>/` by default.
- Exits non-zero on any assertion failure.

Use a custom output directory:

```bash
uv run python scripts/smoke_harness.py --budget 1 --output-dir /tmp/opt-anything-smoke
```

## Run Consecutive Smoke Gate (RB-010)

One command:

```bash
uv run python scripts/consecutive_smoke_gate.py --budget 1
```

What it does:
- Runs `scripts/smoke_harness.py` twice consecutively.
- Prints concise summary fields:
- `pass1`
- `pass2`
- `overall`
- Exits non-zero if either pass fails.
- Saves gate logs and per-pass outputs under `smoke_outputs/consecutive-<timestamp>/` by default.

Use a custom output directory:

```bash
uv run python scripts/consecutive_smoke_gate.py --budget 1 --output-dir /tmp/opt-anything-consecutive
```

## Subscription Live Gates (opt-in)

These gates exercise the Codex and Claude subscription backends against real
local logins. They are marked `pytest.mark.integration` and skip unless their
environment variable is set to `1`, so default CI and `uv run pytest` never
spend subscription quota or bill an API account. Run them on macOS with a saved
Codex login (`uv sync --extra codex`) and Claude Code 2.1.278 or newer.

| Gate | Opt-in variable | Spends | Test file |
|---|---|---|---|
| Subscription live (plan steps 8-10) | `OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1` | Local Codex / Claude subscription quota | `tests/test_subscription_live.py` |
| Paid API fallback (plan step 11) | `OPTIMIZE_ANYTHING_RUN_PAID_FALLBACK_LIVE=1` | OpenAI and Anthropic API accounts (billed) | `tests/test_api_fallback_live.py` |

### Subscription live gate

Invalid API-key sentinels prove the run cannot silently fall back to a paid
key. Run each provider separately:

```bash
OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1 OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run pytest tests/test_subscription_live.py -k codex -v -s
OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1 OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run pytest tests/test_subscription_live.py -k claude -v -s
```

Each covers structured completion, a seedless budget-1 proposer optimize, and a
generated judge evaluator. After both pass, run one built-in judge canary per
provider:

```bash
OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run optimize-anything score examples/seeds/sample_seed.txt --judge-backend codex --no-api-fallback --objective "Score clarity"
OPENAI_API_KEY=deliberately-invalid-live-gate ANTHROPIC_API_KEY=deliberately-invalid-live-gate uv run optimize-anything score examples/seeds/sample_seed.txt --judge-backend claude --no-api-fallback --objective "Score clarity"
```

The score JSON must carry `llm_provenance` with `actual_backend` matching the
provider, `auth_class=subscription`, and no `fallback_source` /
`fallback_reason` keys (those appear only when API fallback ran).

### Paid API fallback gate

This gate bills real API calls. A fake subscription adapter forces an eligible
failure so the real `FallbackBackend` and `LiteLLMBackend` run. Keep
`ANTHROPIC_API_KEY` out of your shell and pass it through an env file (outside
the repo, mode 600). Unset `ANTHROPIC_BASE_URL` so LiteLLM reaches
`https://api.anthropic.com` directly instead of a local proxy:

```bash
env -u ANTHROPIC_BASE_URL OPTIMIZE_ANYTHING_RUN_PAID_FALLBACK_LIVE=1 \
uv run --env-file "$HOME/.config/optimize-anything/paid-fallback.env" \
pytest tests/test_api_fallback_live.py -v -s
```

It asserts same-vendor fallback, the billing warning before dispatch, the sticky
per-role circuit, and zero API calls under `--no-api-fallback`.

### Evidence to record

Every live run records, without account identity, secrets, or artifact prompts:

- Versions: uv, Python, `openai-codex`, Codex CLI, Claude Code CLI (and pytest / litellm for the paid gate)
- OS and architecture
- Auth class and auth source per provider
- Requested and actual backend and model
- Whether fallback ran (`fallback_source` / `fallback_reason`) and `retry_count`; for the paid gate, fallback model, billed call count, and token usage
- Isolation assertions: no secret or identity leakage in retained files, coordination state, or cache keys; Claude child with tools and MCP disabled and a scrubbed environment
- Key scan result over logs and the report
- Pass/fail per gate

Latest evidence: [[Subscription-Live-Evidence-2026-09-26]]
(`docs/verification/subscription-live-evidence-2026-09-26.md`).
47 changes: 47 additions & 0 deletions docs/verification/release-0.6.0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# 0.6.0 release verification

## Live Codex skill optimization

**Run date:** 2026-09-27 PDT (2026-09-28 UTC)
**Decision:** Reject all generated candidates. No candidate changed the skill or `scores.json` during the run; a later release-review correction changed one skill example.

The post-packaging source of `skills/generate-evaluator/SKILL.md` scored **0.8876** with `evaluators/skill_clarity.sh` before the live run. Its SHA-256 was `36176f2956c44802031495bae173725c6843ec8d3b14391dbabcbf98deefa108`. The older `scores.json` entry records 0.8503; that historical value was not used as this run's baseline. The acceptance target was 0.9 plus the evaluator contract and plugin workflow checks.

Both runs used the direct CLI `optimize` path with `--proposer-backend codex`, `--no-api-fallback`, `--no-parallel`, one proposal per iteration, and the deterministic evaluator. Each requested six metric calls. `--evaluator-command bash evaluators/skill_clarity.sh` was the last flag. Raw prompts, candidate text, and run logs remain in ignored `integration_runs/release-0.6.0-u3/`.

| Run ID | Actual metric calls | Codex proposal calls | Finite proposal scores | Retained `num_candidates` | Best score |
|---|---:|---:|---|---:|---:|
| `94a607a28ce9408597ec49b8694be3e8` | 7 | 3 | 0.8275, 0.8728, 0.8473 | 1 | 0.8876 |
| `5e22321923c64eb98ac53d0ba6d6646a` | 7 | 3 | 0.8374, 0.8477, 0.8360 | 1 | 0.8876 |

The six proposals were distinct from the seed, with three distinct proposal texts per run. All six Codex calls reported `actual_backend=codex`, `auth_class=subscription`, and model `gpt-5.6-terra`; the backend plan disabled API fallback and reported no fallback cause or mixed backend. Total spend was **14 evaluator calls** and **6 Codex subscription calls** (57,914 reported tokens across the two runs). GEPA used one more evaluator call than each requested budget because it checks the budget between iterations. No billed API fallback was used.

`num_candidates=1` in each summary counts retained candidates, so it does not mean the run failed to evaluate proposals. The run logs record three evaluated proposals in each run. The best proposal scored 0.8728, a **-0.0148** delta from the frozen baseline; the retained best stayed at 0.8876, a **0.0** delta. The release plan's stricter `num_candidates > 1` verification condition was **not met**, and no candidate reached the 0.9 target. No candidate was applied, so candidate-specific Protocol v2, preflight, numeric score, evaluator-pattern, host/backend, no-fallback, and bundled-launcher acceptance checks were not reached. This is an explicit rejection under the 15-call cap, not an accepted optimization improvement.

Immediately after both runs, the skill still matched the frozen source byte for byte, so the applied optimization diff was empty. Release review later corrected the dataset-aware CLI example; the current skill has SHA-256 `dfbc4465dc5de7ca515520b65422039b08ea518e0396bbc7384b32e182c6e4e7` and scores 0.8875 with the same evaluator, 0.0001 below the frozen run seed. `scores.json` still has SHA-256 `ddfade5810c4323b38daede71a2ff10e9afb0e91df65b1c5c7d40038d177377a`; no score entries were changed.

## Premerge release-candidate checks

**Checked on:** 2026-09-27 PDT (2026-09-28 UTC). **Tested release-code commit:** `d9defd1ec5295223527b81fdf2c83ae5ef752a9c`. The full gate, focused contracts, Bash syntax, type check, and package build ran on that tree immediately before or after its commit. The earlier U2 clean-install check is identified separately below. Subsequent release-evidence edits need the final PR-head CI check; these local results do not claim it has passed.

| Gate | Status | Observed result |
|---|---|---|
| Version and lockfile | PASS | `uv lock` resolved 77 packages and changed the local project entry from 0.5.1 to 0.6.0. Python, Claude manifest and marketplace, and Codex manifest fields now read 0.6.0; the Codex marketplace resolves its version from the manifest. |
| Focused release contracts | PASS | `uv run pytest tests/test_prompt_plugin_contract.py tests/test_doc_contract.py tests/test_plugin_launcher.py`: 38 passed. The version-parity contract reads all active version fields; the installer test covers default Bash 3.2 and Codex paths. |
| Full project gate | PASS | `uv run python scripts/check.py`: 546 passed, 18 skipped; CLI smoke and all three tracked score baselines passed. Credentialed plugin regression was skipped because `--with-plugin` was not set. The reviewed skill's current deterministic score is 0.8875. |
| Type and shell checks | PASS | `uv run mypy src/optimize_anything`: no issues in 27 source files. `/bin/bash -n install.sh scripts/run-optimize-anything` and `git diff --check` passed. |
| Package build and wheel install | PASS | `uv build` made `optimize_anything-0.6.0.tar.gz` and `optimize_anything-0.6.0-py3-none-any.whl`. The wheel lists the CLI module, entry point, and 0.6.0 metadata. An isolated `uv tool run --from dist/optimize_anything-0.6.0-py3-none-any.whl optimize-anything --help` succeeded. |
| Plugin payload inspection | PASS | The 0.6.0 source archive lists both host manifests, the Claude marketplace and commands, the shared launcher, all four canonical skills, and the prompt skill's agent, reference, and evaluator resources. The wheel is the Python CLI package; plugin distribution uses the repository/source tree. |
| Clean host install (U2 result, reused) | PASS | A fresh copied plugin snapshot with no global CLI exposed four shared skills and nine Claude commands. Its bundled launcher ran `budget` with Codex SDK 0.156.0, generated a Codex/no-fallback evaluator, and passed child preflight. A separate isolated global install and Python 3.10.21 check passed after pinning the SDK to 0.156.0; U2 reported 56 focused tests passed. This check was not rerun during U4 package inspection. |
| Review and fixes | PASS | Compound Engineering review run `20260927-174849-e5d41434` completed with an independent Claude adversarial pass. Three validated findings were fixed: default Bash installer, dataset quick start, and bundled LiteLLM example/guard. The evaluator child's dependency flags were aligned. A platform limitation remains documented below. |
| PR head and post-merge gates | PENDING | Final PR-head CI results, merge SHA, main workflow, and tag/install results belong to the release operation. |

The plan's `num_candidates > 1` proxy is a failed check, not a publication gate: GEPA counts retained candidates there. The live logs instead show six distinct evaluated proposals with finite scores, actual Codex subscription provenance, and no API fallback. This satisfies the real-candidate part of R4. All scores were below the frozen baseline and the 0.9 acceptance target, so R5 rejected them; no optimization candidate changed the source or score history. The later CLI example correction is separate from that decision. No candidate improvement or cross-provider corroboration is claimed.

## Known audit findings and release disposition

- **GEPA evaluator cache identity (R4): follow-up, with an operating constraint.** `test_cache_fingerprint_uses_actual_route_without_exposing_input` covers the backend's safe fingerprint, but GEPA's evaluator `fitness_cache` key does not include the actual completion route after a subscription-to-API fallback. This affects opt-in `--cache`/`--cache-from` runs when an evaluator or judge can switch routes; it does not affect the default uncached path or the uncached Codex release run above. `install.md` now tells users to omit those cache flags for such runs. A route-aware cache policy or GEPA integration remains follow-up work; the audit's R4 status remains partial.
- **TOML table-over-scalar precedence (R6): follow-up.** One TOML document cannot define `model.proposer` as both a scalar and a table; standard parsing rejects it before precedence can apply. `test_scalar_table_model_conflict_names_both_keys` verifies the specific diagnostic, while existing CLI-over-spec tests verify the supported override path. Use one representation per role and a CLI flag for an override. Multi-file layering or new syntax requires a separate product decision; the audit's R6 status remains partial.
- **Plugin platforms without a Codex binary wheel: follow-up.** The bundled launcher selects the Codex extra for every plugin invocation. On an architecture without an `openai-codex-cli-bin` wheel, even API-only plugin commands cannot start. The release install check covers macOS arm64; other platforms remain untested. `install.md` directs API-only users on those platforms to the global CLI or source install without the Codex extra. A broader plugin runtime policy needs a separate compatibility decision.

The premerge evidence and changelog contain no credentials, account identity, raw prompts, or private candidate text. Detailed live artifacts remain in ignored `integration_runs/`; built packages and smoke outputs remain outside the committed release evidence.
2 changes: 1 addition & 1 deletion docs/verification/subscription-backends-audit.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ This report audits units U1-U7 of `docs/plans/2026-09-22-1841-feature-subscripti

- **Initial audited commit:** `0b0a99562e657d173b32c94f1f27062669e7b794` on `feat/codex-claude-subscription`.
- `src/` is identical to `origin/main` at `70e1fdf2905a1f44b5b447eaa027f5ecfb6dbcef`, because PR #6 merged the backends.
- PR #7 is still open: https://github.com/ASRagab/optimize-anything/pull/7
- PR #7 was open during the initial audit and merged on 2026-09-27: https://github.com/ASRagab/optimize-anything/pull/7
- **Post-fix baseline:** `279f022`; the final offline gates below ran on this code before the audit report update.
- **Live evidence:** Verification Contract steps 8-11 (the opt-in subscription and paid-fallback gates) are recorded in `docs/verification/subscription-live-evidence-2026-09-26.md`.
- This audit covers code and offline tests only.
Expand Down
Loading
Loading