Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
{
"name": "optimize-anything",
"interface": {
"displayName": "Optimize Anything"
},
"plugins": [
{
"name": "optimize-anything",
"source": {
"source": "local",
"path": "./"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Productivity"
}
]
}
24 changes: 24 additions & 0 deletions .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"name": "optimize-anything",
"version": "0.5.0",
"description": "Optimize prompts and other text artifacts with measured evaluator feedback",
"author": {
"name": "optimize-anything contributors"
},
"repository": "https://github.com/ASRagab/optimize-anything",
"license": "MIT",
"keywords": ["optimization", "prompts", "evaluator", "gepa"],
"skills": "./skills/",
"interface": {
"displayName": "Optimize Anything",
"shortDescription": "Optimize prompts with measured feedback",
"longDescription": "Build an evaluation contract, improve prompts with the bundled optimize-anything runtime, and apply only accepted results.",
"developerName": "optimize-anything contributors",
"category": "Productivity",
"capabilities": ["Interactive", "Write"],
"defaultPrompt": [
"Optimize this prompt and verify the result.",
"Improve the prompt in this file safely."
]
}
}
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,13 +8,24 @@
- Centralized project-owned model defaults across the CLI, evaluator generation, judge flows, and operational scripts
- Preserved proposer precedence as explicit CLI model, then `OPTIMIZE_ANYTHING_MODEL`, then the project default

### Prompt optimization workflow
- Added the `optimize-prompt` skill for inline prompts, files, embedded prompt regions, and independent batches
- Separated fast prompt-quality scoring from rigorous task-output evaluation with deterministic hard gates and held-out acceptance
- Kept repository sources unchanged during search and limited accepted edits to the recorded prompt destination

### Plugin distribution
- Added a self-locating locked-runtime launcher used by every packaged Claude command
- Added native Codex plugin and marketplace metadata over the same canonical skill tree
- Aligned Python, Claude, and Codex release metadata at 0.5.0

### Provider compatibility
- Let providers apply their own sampling defaults unless a judge temperature is explicitly supplied
- Removed hard-coded sampling parameters from generated judge and composite evaluators

### Documentation and verification
- Modernized runnable examples, command guidance, protocol docs, skills, and integration tooling for current model identifiers
- Added drift coverage for model defaults, CLI help and resolution, generated evaluators, and judge request payloads
- Added offline launcher, evaluator, workflow-fixture, documentation, manifest, and release-version contracts
- Synchronized package, plugin, and marketplace release versions

## v0.4.0 - 2026-07-27
Expand Down
86 changes: 74 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,9 +163,59 @@ commands, expected report fields, acceptance criteria, and troubleshooting.
- `analyze`
- `validate`

## Claude Code Plugin
## Agent Plugins

optimize-anything is also a Claude Code plugin with guided slash commands and skills.
The Claude Code plugin and Codex plugin share the same `skills/` tree and
bundled locked runtime. Plugin users need `uv` and Python 3.10 or newer, but do
not need a global `optimize-anything` command. The standalone CLI remains a
separate installation choice.

### Prompt Optimization Workflow

Claude Code users invoke `$optimize-prompt`; Codex users invoke the namespaced
`$optimize-anything:optimize-prompt`. Both accept prompt text, a standalone
file, an embedded prompt region, or a list of independent prompt files.

| Evidence mode | Evaluation | What it proves |
|---|---|---|
| **Fast mode** | Scores the prompt text for clarity, constraints, and task fitness | Prompt-quality evidence only |
| **Rigorous mode** | Runs the candidate on representative inputs, then scores task outputs | Task-performance evidence for the tested examples |
| **Composite mode** | Runs deterministic hard constraints before the rigorous judge | No subjective score can override a failed gate |

Rigorous datasets use JSONL records with `input`, optional `expected`,
`criteria`, and `hard_constraints`. See
[`prompt-execution-dataset.md`](skills/optimize-prompt/references/prompt-execution-dataset.md)
for representative examples, the default system-prompt adapter, custom adapter
guidance, and expected cost controls. Use explicit proposer, target, and judge
models; bound calls with a small dataset, budget, and early stopping.

The workflow captures the baseline and writes search output to a temporary or
run directory. It accepts only a positive comparable score delta with all hard
constraints satisfied, plus required held-out acceptance in rigorous mode.
Independent prompt files receive separate decisions. Coupled components require
an explicit structured-candidate adapter and are not optimized as unrelated
files.

Inline example:

```text
$optimize-prompt Improve this prompt in fast mode and return the accepted result:
"Summarize this."
```

The response includes the complete accepted prompt, evidence mode, and score
delta; it does not write a repository file.

Embedded repository example:

```text
$optimize-prompt Optimize SYSTEM_PROMPT in src/agent.py with the examples in
evals/prompt.jsonl. Apply it only if held-out acceptance passes.
```

The workflow optimizes a captured copy, replaces only `SYSTEM_PROMPT`, preserves
the source representation, and runs the cheapest relevant parse or targeted
test. Rejected candidates leave the file unchanged.

### Plugin Regression Workflow

Expand All @@ -182,25 +232,36 @@ uv run python scripts/check.py --with-plugin
Requirements:
- `claude` CLI installed and authenticated
- `OPENAI_API_KEY` set in the shell that launches the command
- `ANTHROPIC_API_KEY` set in the shell that launches the command
- `ANTHROPIC_API_KEY` for `validate` or the full scenario set

The harness runs three real scenarios (`analyze`, `validate`, `quick`), saves Claude JSON outputs plus stderr logs, and fails if Claude does not execute the expected workflow or the optimized artifact is not written.
The harness runs existing CLI scenarios plus bounded prompt inline-return and
repository-apply scenarios. `--dry-run` verifies prompt and artifact wiring
without credentials or model calls.

### Installation
### Claude Code Plugin

```bash
# In Claude Code
/plugin install ASRagab/optimize-anything
/plugin marketplace add ASRagab/optimize-anything
/plugin install optimize-anything@optimize-anything
```

Or clone and install locally:
For a local clone, replace the repository name in the first command with its
absolute path. Restart Claude Code after installation or update, then invoke
`$optimize-prompt` or an existing slash command.

### Codex Plugin

```bash
git clone https://github.com/ASRagab/optimize-anything.git
cd optimize-anything
/plugin install .
codex plugin marketplace add ASRagab/optimize-anything
codex plugin add optimize-anything@optimize-anything
```

For local development, pass the clone path to `codex plugin marketplace add`.
Start a new Codex thread after install or update, then invoke
`$optimize-anything:optimize-prompt` so skill discovery refreshes.
See [install.md](install.md) for update, removal, verification, and the optional
standalone CLI path.

### Slash Commands

| Command | Description |
Expand All @@ -217,8 +278,9 @@ cd optimize-anything

### Skills

The plugin includes three skills that Claude Code can invoke automatically:
Both plugins include four skills that the host can invoke:

- **optimize-prompt** — Build the rubric, choose fast or rigorous evidence, optimize outside the source, and return or safely apply an accepted prompt
- **optimization-guide** — Full workflow walkthrough covering modes, configuration, budget, and result interpretation
- **generate-evaluator** — Choose the right evaluator pattern (judge, command, composite) and generate a script
- **evaluator-patterns** — Library of ready-to-use evaluator templates for prompts, code, docs, and agent instructions
Expand Down
6 changes: 6 additions & 0 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,14 @@ Or skip evaluator setup entirely — the guided workflow handles it:
/optimize-anything:quick my-prompt.txt "make it clearer and more specific"
```

For inline prompts, repository prompt regions, or representative-example
evaluation, invoke `$optimize-prompt` in Claude Code or the namespaced
`$optimize-anything:optimize-prompt` in Codex. It uses the bundled runtime,
keeps search output outside the source, and applies only an accepted result.

## Available Skills

- **optimize-prompt** — Optimize inline prompts, files, embedded regions, or independent batches with fast prompt-quality or rigorous task-output evidence
- **generate-evaluator** — Choose the right evaluator pattern (judge, command, composite) and generate a script
- **optimization-guide** — Full workflow walkthrough covering optimization modes, configuration, budget, and result interpretation
- **evaluator-patterns** — Library of ready-to-use evaluator templates for prompts, code, docs, and agent instructions
Expand Down
6 changes: 3 additions & 3 deletions commands/analyze.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,13 @@ Use an LLM to discover relevant quality dimensions for a given artifact and opti
## Usage

```bash
optimize-anything analyze SEED_FILE --judge-model openai/gpt-5.6-luna --objective "Quality"
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze SEED_FILE --judge-model openai/gpt-5.6-luna --objective "Quality"
```

## Example

```bash
optimize-anything analyze my-prompt.txt \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze my-prompt.txt \
--judge-model openai/gpt-5.6-luna \
--objective "Score for clarity and persuasiveness"
```
Expand All @@ -25,7 +25,7 @@ optimize-anything analyze my-prompt.txt \
After dimension discovery, proceed directly to optimization using the returned `intake_json`:

```bash
optimize-anything optimize my-prompt.txt \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" optimize my-prompt.txt \
--judge-model openai/gpt-5.6-luna \
--objective "Score for clarity and persuasiveness" \
--intake-json '<paste intake_json from analyze>' \
Expand Down
4 changes: 2 additions & 2 deletions commands/budget.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,13 @@ Analyze a seed artifact and recommend an appropriate iteration budget based on i
## Usage

```
optimize-anything budget SEED_FILE
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" budget SEED_FILE
```

## Example

```
optimize-anything budget my-prompt.txt
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" budget my-prompt.txt
```

See [README](../README.md) for full flag documentation.
2 changes: 1 addition & 1 deletion commands/compare.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Compare two artifacts with the same scoring setup by composing existing `score`

## Procedure
1. Score the original artifact:
- `optimize-anything score <original> --judge-model <model> --objective "..." [--intake-json ...]`
- `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" score <original> --judge-model <model> --objective "..." [--intake-json ...]`
2. Score the optimized artifact with the exact same evaluator setup.
3. Present side-by-side:
- Overall score for each artifact
Expand Down
4 changes: 2 additions & 2 deletions commands/explain.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,13 @@ Display the optimization plan that would be executed for the given seed artifact
## Usage

```
optimize-anything explain SEED_FILE --evaluator-command bash eval.sh --objective "your goal"
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" explain SEED_FILE --evaluator-command bash eval.sh --objective "your goal"
```

## Example

```
optimize-anything explain prompt.txt \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" explain prompt.txt \
--evaluator-command bash evaluators/clarity.sh \
--budget 50
```
Expand Down
4 changes: 2 additions & 2 deletions commands/intake.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,13 +10,13 @@ Normalize and validate an intake specification, filling in defaults for quality
## Usage

```
optimize-anything intake --intake-json '{"artifact_class": "prompt"}'
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" intake --intake-json '{"artifact_class": "prompt"}'
```

## Example

```
optimize-anything intake --intake-file intake.json
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" intake --intake-file intake.json
```

See [README](../README.md) for full flag documentation.
6 changes: 5 additions & 1 deletion commands/optimize.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,10 @@ description: Guided optimization workflow with mode selection and evaluator setu
---
Run optimization using a deterministic, guided workflow.

For inline prompts, embedded prompt regions, independent prompt batches, or
task-output evaluation with representative examples, invoke the shared
`$optimize-prompt` skill instead.

## Step 1: Identify the artifact
- If the user provided a file argument, use it directly.
- Otherwise ask: **"What file should I optimize?"**
Expand All @@ -29,7 +33,7 @@ Present these options and ask the user to choose one unless they already specifi
- If evaluator is already specified, proceed.
- If no evaluator is specified:
1. Run `analyze` first:
- `optimize-anything analyze <file> --judge-model <model> --objective "<objective>"`
- `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model <model> --objective "<objective>"`
2. If analyze fails (API key missing, model unavailable): ask the user for their preferred model, or suggest using `--evaluator-command` with a custom script instead.
3. Ask: **"Should we use LLM judge directly, or do you want a custom evaluator?"**
4. If custom evaluator is needed, invoke evaluator generation workflow.
Expand Down
2 changes: 1 addition & 1 deletion commands/quick.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Run a no-questions-asked fast optimization.

## Behavior (do not ask follow-up questions)
1. Run analysis to discover dimensions:
- `optimize-anything analyze <file> --judge-model openai/gpt-5.6-luna --objective "<objective>"`
- `"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" analyze <file> --judge-model openai/gpt-5.6-luna --objective "<objective>"`
- If analyze fails, skip dimension discovery and run optimize with `--judge-model` directly using the objective as-is.
2. Run optimization using LLM judge with:
- `--judge-model openai/gpt-5.6-luna`
Expand Down
8 changes: 4 additions & 4 deletions commands/score.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,17 +10,17 @@ Score a single artifact file using a command evaluator, HTTP evaluator, or LLM j
## Usage

```
optimize-anything score SEED_FILE --evaluator-command bash eval.sh
optimize-anything score SEED_FILE --judge-model openai/gpt-5.6-luna --objective "Score clarity"
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" score SEED_FILE --evaluator-command bash eval.sh
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" score SEED_FILE --judge-model openai/gpt-5.6-luna --objective "Score clarity"
```

## Example

```
optimize-anything score my-prompt.txt \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" score my-prompt.txt \
--evaluator-command bash evaluators/clarity.sh

optimize-anything score my-prompt.txt \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" score my-prompt.txt \
--judge-model openai/gpt-5.6-luna \
--objective "Score for persuasiveness"
```
Expand Down
2 changes: 1 addition & 1 deletion commands/validate.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Use multiple LLM judges to verify that a quality improvement is not provider-spe

## Usage
```bash
optimize-anything validate <file> \
"${CLAUDE_PLUGIN_ROOT}/scripts/run-optimize-anything" validate <file> \
--providers openai/gpt-5.6-luna anthropic/claude-sonnet-5 gemini/gemini-3.6-flash \
--objective "Score for clarity and constraint adherence"
```
Expand Down
Loading
Loading