Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 57 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,6 +178,56 @@ commands, expected report fields, acceptance criteria, and troubleshooting.
- `analyze`
- `validate`

## Codex and Claude subscription backends

Local Codex and Claude Code logins can power proposer and built-in evaluator
roles without API keys. Selection is explicit; omitting backend flags preserves
the existing LiteLLM API behavior.

```bash
# Install the pinned Codex SDK adapter, then authenticate with ChatGPT.
uv sync --extra codex
codex login

# Codex subscription: proposer plus built-in judge.
optimize-anything optimize seed.txt \
--proposer-backend codex --judge-backend codex \
--objective "Improve clarity"

# Claude Code subscription. Install/sign in to the claude CLI first.
optimize-anything optimize seed.txt \
--proposer-backend claude --judge-backend claude \
--objective "Improve clarity"
```

Subscription calls are serialized per provider by default. The CLI may switch
an eligible failure to a same-vendor API only when a matching key and fallback
model are available, and prints a billing warning first. Use
`--no-api-fallback` to prohibit that switch. Claude support is local-only and
experimental; neither adapter initiates login or copies credential contents.

Single-role commands use `--analysis-backend` (`analyze`) or
`--judge-backend` (`score`). `validate --providers` also accepts `codex`,
`codex:<model>`, `claude`, and `claude:<model>`. Structured TOML uses role
tables such as:

```toml
[model.proposer]
backend = "codex"
api_fallback = false

[model.judge]
backend = "claude"
api_fallback_model = "anthropic/claude-sonnet-5"
```

Opt-in live gates consume local subscription quota:

```bash
OPTIMIZE_ANYTHING_RUN_SUBSCRIPTION_LIVE=1 \
uv run pytest tests/test_subscription_live.py
```

## Agent Plugins

The Claude Code plugin and Codex plugin share the same `skills/` tree and
Expand Down Expand Up @@ -380,7 +430,7 @@ optimize-anything optimize seed.txt \

### `optimize` flags (complete)

Exactly one evaluator source is required: `--evaluator-command` OR `--evaluator-url` OR `--judge-model`.
Exactly one evaluator source is required: `--evaluator-command` OR `--evaluator-url` OR a built-in judge selected by `--judge-model`/`--judge-backend`.

| Flag | Description | Default |
|---|---|---|
Expand All @@ -397,7 +447,13 @@ Exactly one evaluator source is required: `--evaluator-command` OR `--evaluator-
| `--budget <int>` | Max evaluator calls | `100` |
| `--output, -o <file>` | Write best artifact to file | -- |
| `--model <model>` | Proposer model (or env fallback) | `OPTIMIZE_ANYTHING_MODEL`, then `openai/gpt-5.6-sol` |
| `--proposer-backend api\|codex\|claude` | Proposal backend | `api` |
| `--judge-model <model>` | Built-in LLM judge evaluator model | -- |
| `--judge-backend api\|codex\|claude` | Built-in judge backend | `api` |
| `--subscription-concurrency <int>` | Concurrent calls allowed per subscription provider | `1` |
| `--no-api-fallback` | Prohibit subscription-to-API fallback | `false` |
| `--openai-api-fallback-model <model>` | Same-vendor API fallback for Codex | -- |
| `--anthropic-api-fallback-model <model>` | Same-vendor API fallback for Claude | -- |
| `--judge-objective <text>` | Judge objective override | falls back to `--objective` |
| `--api-base <url>` | Override LiteLLM API base | -- |
| `--diff` | Print unified diff (seed vs best) to stderr | `false` |
Expand Down
13 changes: 13 additions & 0 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,19 @@ evaluation, invoke `$optimize-prompt` in Claude Code or the namespaced
`$optimize-anything:optimize-prompt` in Codex. It uses the bundled runtime,
keeps search output outside the source, and applies only an accepted result.

## Reuse the Host Subscription

When this skill runs in Codex, pass `--proposer-backend codex` and use
`--judge-backend codex` for built-in judging. When it runs in Claude Code, use
the corresponding `claude` values. For `analyze`, pass `--analysis-backend`;
for `validate`, use the reserved provider selector `codex` or `claude`.
Do not infer a backend in the Python runtime or for an unknown host.

Tell the user before running that subscription calls are serialized by default.
Eligible availability, authentication, rate-limit, or quota failures may switch
to a same-vendor API model only when a matching key and fallback model are
available; pass `--no-api-fallback` when billed fallback is not acceptable.

## Available Skills

- **optimize-prompt** — Optimize inline prompts, files, embedded regions, or independent batches with fast prompt-quality or rigorous task-output evidence
Expand Down
7 changes: 7 additions & 0 deletions commands/analyze.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@
name: analyze
description: Discover quality dimensions for an artifact and objective
---
When running in Codex, pass `--analysis-backend codex`; in Claude Code, pass
`--analysis-backend claude`. A subscription backend may omit `--judge-model`.
Unknown hosts use the existing API model. Announce possible same-vendor billed
API fallback, or pass `--no-api-fallback` when it is not acceptable.

# analyze

Expand All @@ -21,6 +25,9 @@ Use an LLM to discover relevant quality dimensions for a given artifact and opti
--objective "Score for clarity and persuasiveness"
```

In Codex or Claude Code, replace `--judge-model ...` with
`--analysis-backend codex` or `--analysis-backend claude`, respectively.

## Next Step: optimize with discovered dimensions
After dimension discovery, proceed directly to optimization using the returned `intake_json`:

Expand Down
4 changes: 4 additions & 0 deletions commands/compare.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@
name: compare
description: Side-by-side comparison of two artifacts using composed score calls
---
When comparison uses LLM scoring, select the current host explicitly:
`--judge-backend codex` in Codex or `--judge-backend claude` in Claude Code.
Unknown hosts keep API defaults. Announce possible same-vendor billed API
fallback, or add `--no-api-fallback` to prohibit it.
Compare two artifacts with the same scoring setup by composing existing `score` calls.

## Usage
Expand Down
9 changes: 8 additions & 1 deletion commands/optimize.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,12 @@ For inline prompts, embedded prompt regions, independent prompt batches, or
task-output evaluation with representative examples, invoke the shared
`$optimize-prompt` skill instead.

When executing this Claude Code command, add `--proposer-backend claude` and,
when using the built-in judge, `--judge-backend claude`. Announce that
subscription calls are serialized and that
eligible failures may use billed same-vendor API fallback; honor a request to
disable it with `--no-api-fallback`.

## Step 1: Identify the artifact
- If the user provided a file argument, use it directly.
- Otherwise ask: **"What file should I optimize?"**
Expand All @@ -33,7 +39,8 @@ Present these options and ask the user to choose one unless they already specifi
- If evaluator is already specified, proceed.
- If no evaluator is specified:
1. Run `analyze` first:
- `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model <model> --objective "<objective>"`
- Subscription mode: `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --analysis-backend claude --objective "<objective>"`
- API mode: `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model <model> --objective "<objective>"`
2. If analyze fails (API key missing, model unavailable): ask the user for their preferred model, or suggest using `--evaluator-command` with a custom script instead.
3. Ask: **"Should we use LLM judge directly, or do you want a custom evaluator?"**
4. If custom evaluator is needed, invoke evaluator generation workflow.
Expand Down
10 changes: 8 additions & 2 deletions commands/quick.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,15 +4,21 @@ description: Zero-config one-shot optimization for fast improvements
---
Run a no-questions-asked fast optimization.

Use Claude Code's subscription explicitly with `--analysis-backend claude`,
`--proposer-backend claude`, and `--judge-backend claude`. Announce serialized
subscription use and possible same-vendor billed API fallback before running.

## Usage
`/optimize-anything:quick <file> "<objective>"`

## Behavior (do not ask follow-up questions)
1. Run analysis to discover dimensions:
- `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model openai/gpt-5.6-luna --objective "<objective>"`
- Subscription mode: `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --analysis-backend claude --objective "<objective>"`
- API mode: `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model openai/gpt-5.6-luna --objective "<objective>"`
- If analyze fails, skip dimension discovery and run optimize with `--judge-model` directly using the objective as-is.
2. Run optimization using LLM judge with:
- `--judge-model openai/gpt-5.6-luna`
- Subscription mode: `--proposer-backend claude --judge-backend claude` (omit API model strings)
- API mode: `--judge-model openai/gpt-5.6-luna`
- `--budget 50`
- `--diff`
- `--early-stop`
Expand Down
4 changes: 4 additions & 0 deletions commands/score.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@
name: score
description: Score a single artifact with an evaluator
---
For an LLM judge, use `--judge-backend codex` in Codex or
`--judge-backend claude` in Claude Code; the subscription model may be omitted.
Do not add a judge backend for command or HTTP evaluators. Announce possible
same-vendor billed API fallback, or add `--no-api-fallback` to prohibit it.

# score

Expand Down
4 changes: 4 additions & 0 deletions commands/validate.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@
name: validate
description: Cross-validate an artifact with multiple LLM judge providers
---
The `--providers` list accepts `codex`, `codex:<model>`, `claude`, and
`claude:<model>` before ordinary LiteLLM model strings. Prefer the selector for
the current host when subscription reuse is requested. Announce possible
same-vendor billed API fallback, or add `--no-api-fallback` to prohibit it.
Use multiple LLM judges to verify that a quality improvement is not provider-specific.

## When to use
Expand Down
Loading
Loading