Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,3 +107,17 @@ The intake specification (`--intake-json` / `--intake-file`) normalizes these fi
- `evaluation_pattern` — scoring strategy (`verification`, `judge`, `simulation`, `composite`)
- `execution_mode` — evaluator transport (`command`, `http`)
- `evaluator_cwd` — working directory for command evaluators

## Agent skills

### Issue tracker

Issues and PRDs live as GitHub issues in `ASRagab/optimize-anything`, managed via the `gh` CLI. See `docs/agents/issue-tracker.md`.

### Triage labels

Five canonical triage roles, with `needs-triage`→`triage`, `ready-for-agent`→`dev`, `ready-for-human`→`review` remapped. See `docs/agents/triage-labels.md`.

### Domain docs

Single-context: one `CONTEXT.md` + `docs/adr/` at the repo root. See `docs/agents/domain.md`.
6 changes: 5 additions & 1 deletion EXAMPLES.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,10 +89,14 @@ optimize-anything optimize prompt.txt \
--dataset data/train.jsonl \
--model openai/gpt-4o-mini \
--budget 120 \
--parallel --workers 6 \
--workers 6 \
--cache --run-dir runs
```

Optimization runs evaluator calls in parallel by default. Use `--workers` to cap
concurrency, or `--no-parallel` when an evaluator writes shared temp files,
depends on process-global state, or must stay below strict provider rate limits.

With validation set:

```bash
Expand Down
26 changes: 23 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ optimize-anything optimize seed.txt \
--objective "Improve clarity and specificity" \
--model openai/gpt-4o-mini \
--budget 20 \
--parallel --workers 4 \
--workers 4 \
--cache \
--run-dir runs \
--output result.txt
Expand Down Expand Up @@ -60,7 +60,7 @@ optimize-anything optimize prompt.txt \
--dataset data/train.jsonl \
--valset data/val.jsonl \
--model openai/gpt-4o-mini \
--budget 120 --parallel --workers 6 --cache --run-dir runs
--budget 120 --workers 6 --cache --run-dir runs
```

### Multi-provider validation
Expand Down Expand Up @@ -91,6 +91,9 @@ optimize-anything optimize --no-seed \

- Early stop is auto-enabled when `--budget > 30` (or force with `--early-stop`)
- Reuse prior evaluator cache with `--cache-from` (requires `--cache` + `--run-dir`)
- Evaluator calls run in parallel by default; pass `--no-parallel` for evaluators
that write shared temp files, depend on process-global state, or need strict
provider rate-limit control.

```bash
optimize-anything optimize seed.txt \
Expand All @@ -115,6 +118,22 @@ optimize-anything optimize seed.txt \
--score-range any
```

### Optimization observability benchmark

Use the maintainer benchmark when you want evidence that an optimization loop
improved a useful artifact and was not just a seed-only or scorer-gaming run.

```bash
uv run python scripts/optimization_observer.py setup
uv run python scripts/optimization_observer.py report \
--run-dir integration_runs/optimization-observability/evaluator-generation-guidance/runs/run-YYYYMMDD-HHMMSS \
--benchmark examples/optimization-observability/evaluator-generation-benchmark.json \
--strict
```

See `examples/optimization-observability/README.md` for the full benchmark
commands, expected report fields, acceptance criteria, and troubleshooting.

## CLI Subcommands

- `optimize`
Expand Down Expand Up @@ -288,7 +307,8 @@ Exactly one evaluator source is required: `--evaluator-command` OR `--evaluator-
| `--api-base <url>` | Override LiteLLM API base | -- |
| `--diff` | Print unified diff (seed vs best) to stderr | `false` |
| `--run-dir <path>` | Save run artifacts in timestamped run dir | -- |
| `--parallel` | Enable parallel evaluator calls | `false` |
| `--parallel` | Explicitly enable parallel evaluator calls | `true` |
| `--no-parallel` | Run evaluator calls serially | -- |
| `--workers <int>` | Max workers for parallel evaluation | -- |
| `--cache` | Enable evaluator cache | `false` |
| `--cache-from <run-dir>` | Copy prior `fitness_cache` into new run | -- |
Expand Down
10 changes: 7 additions & 3 deletions WALKTHROUGH.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,19 +105,23 @@ uv run optimize-anything optimize seed.txt \
--budget 120 --cache --run-dir runs
```

## Step 9.5: Speed up with parallel workers
## Step 9.5: Tune parallel workers

For expensive evaluators, add parallel execution:
Optimization runs evaluator calls in parallel by default. For expensive
evaluators, cap concurrency with workers:

```bash
uv run optimize-anything optimize seed.txt \
--evaluator-command bash evaluators/eval.sh \
--objective "Improve quality" \
--model openai/gpt-4o-mini \
--budget 100 \
--parallel --workers 8
--workers 8
```

Use `--no-parallel` if the evaluator writes shared temp files, depends on
process-global state, or needs strict provider rate-limit control.

## Step 10: Accept or rerun

If improved, copy best artifact back into source. Otherwise adjust objective/intake and rerun.
Expand Down
7 changes: 6 additions & 1 deletion commands/optimize.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,12 @@ Mode guidance:

If optimization fails: check error message. Common issues: missing API key, model quota exceeded, evaluator script error.

For larger budgets (especially >50), suggest `--parallel` (and optionally `--workers`) when evaluator setup can support concurrency.
Optimization runs evaluator calls in parallel by default. For larger budgets
(especially >50), suggest `--workers` when evaluator setup can support
concurrency. Suggest `--no-parallel` for evaluators that write shared temp
files, depend on process-global state, or need strict provider rate-limit
control. `--parallel` remains valid when users want to enable parallel mode
explicitly.

## Step 5: Present results clearly
After completion, provide:
Expand Down
143 changes: 143 additions & 0 deletions evaluators/evaluator_generation_training.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
#!/usr/bin/env bash
# Deterministic training scorer for the evaluator-generation benchmark.
# Input: {"candidate": "<SKILL.md text>"}
# Output: {"score": <0..1>, dimension scores, "feedback": [...]}

set -euo pipefail

input_file=$(mktemp)
trap 'rm -f "$input_file"' EXIT
cat > "$input_file"

python3 - "$input_file" <<'PY'
from __future__ import annotations

import json
import math
import re
import sys


def clamp(value: float) -> float:
return max(0.0, min(1.0, value))


def contains_any(text: str, options: list[str]) -> bool:
return any(option in text for option in options)


def ratio(hits: list[bool]) -> float:
return sum(1 for hit in hits if hit) / max(len(hits), 1)


def length_score(text: str) -> float:
length = len(text)
if length < 700:
return clamp(length / 700.0)
if length <= 5200:
return 1.0
if length <= 7500:
return clamp(1.0 - ((length - 5200) / 3500.0))
return 0.25


try:
with open(sys.argv[1], "r", encoding="utf-8") as fh:
payload = json.load(fh)
except json.JSONDecodeError:
print(json.dumps({"score": 0.0, "error": "Input must be valid JSON"}))
raise SystemExit(0)

if "candidate" not in payload or not isinstance(payload["candidate"], str):
print(json.dumps({"score": 0.0, "error": "Payload must include string field 'candidate'"}))
raise SystemExit(0)

candidate = payload["candidate"]
if not candidate.strip():
print(json.dumps({"score": 0.0, "error": "Candidate must not be empty"}))
raise SystemExit(0)

lower = candidate.lower()
feedback: list[str] = []
code_blocks = re.findall(r"```[a-zA-Z0-9_-]*\n(.*?)```", candidate, re.DOTALL)

contract_completeness = ratio([
"candidate" in lower,
"json" in lower and contains_any(lower, ["stdin", "post body", "http post"]),
'"score"' in candidate or "`score`" in candidate or "score output" in lower,
contains_any(lower, ["float", "numeric", "finite", "[0,1]", "[0, 1]"]),
contains_any(lower, ["feedback", "diagnostic", "dimension", "reasoning"]),
contains_any(lower, ["dataset", "example"]),
])
if contract_completeness < 0.85:
feedback.append("Clarify candidate input, numeric score output, score range, and diagnostic side fields.")

runnable_examples = ratio([
bool(code_blocks),
"optimize-anything generate-evaluator" in lower,
"echo" in lower and "candidate" in lower,
contains_any(lower, ["python3", "bash", "--evaluator-command"]),
"--objective" in lower,
])
if runnable_examples < 0.8:
feedback.append("Add copy-pasteable commands that generate an evaluator and test it with a JSON payload.")

feedback_quality = ratio([
contains_any(lower, ["feedback", "diagnostic", "reasoning"]),
contains_any(lower, ["dimension", "quality_dimensions", "subscore"]),
contains_any(lower, ["reflection", "weak", "improve"]),
contains_any(lower, ["hard constraint", "constraint failure", "safety gate"]),
contains_any(lower, ["customize", "scoring logic", "rubric"]),
])
if feedback_quality < 0.8:
feedback.append("Explain how evaluators return actionable diagnostics for reflection, not just a scalar score.")

calibration_guidance = ratio([
contains_any(lower, ["0.3-0.7", "0.3 to 0.7", "0.85"]),
contains_any(lower, ["discrimination", "non-discriminating", "too easy"]),
contains_any(lower, ["baseline", "threshold", "score range"]),
contains_any(lower, ["test", "validate"]),
contains_any(lower, ["edge case", "malformed", "empty"]),
])
if calibration_guidance < 0.65:
feedback.append("Add calibration guidance for weak or non-discriminating evaluators.")

structure = ratio([
candidate.lstrip().startswith("---"),
lower.count("\n## ") >= 4,
lower.count("\n### ") >= 2,
candidate.count("- ") >= 6,
len(code_blocks) >= 2,
])
if structure < 0.8:
feedback.append("Use frontmatter, section headings, lists, and fenced examples for scanability.")

conciseness = length_score(candidate)
if conciseness < 0.8:
feedback.append("Keep the guidance concise enough to remain usable as a skill file.")

score = (
contract_completeness * 0.28
+ runnable_examples * 0.22
+ feedback_quality * 0.20
+ calibration_guidance * 0.16
+ structure * 0.08
+ conciseness * 0.06
)
score = clamp(score)
if not math.isfinite(score):
score = 0.0
if not feedback:
feedback.append("Strong baseline; look for sharper calibration, held-out validation, or scorer-debugging guidance.")

print(json.dumps({
"score": round(score, 4),
"contract_completeness": round(contract_completeness, 4),
"runnable_examples": round(runnable_examples, 4),
"feedback_quality": round(feedback_quality, 4),
"calibration_guidance": round(calibration_guidance, 4),
"structure": round(structure, 4),
"conciseness": round(conciseness, 4),
"feedback": feedback[:5],
}))
PY
109 changes: 109 additions & 0 deletions evaluators/evaluator_generation_validation.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Deterministic held-out scorer for the evaluator-generation benchmark.
# Uses a different rubric from the training scorer to discourage scorer gaming.

set -euo pipefail

input_file=$(mktemp)
trap 'rm -f "$input_file"' EXIT
cat > "$input_file"

python3 - "$input_file" <<'PY'
from __future__ import annotations

import json
import re
import sys


def clamp(value: float) -> float:
return max(0.0, min(1.0, value))


def any_of(text: str, phrases: list[str]) -> bool:
return any(phrase in text for phrase in phrases)


def score_hits(hits: list[bool]) -> float:
return sum(1 for hit in hits if hit) / max(len(hits), 1)


try:
with open(sys.argv[1], "r", encoding="utf-8") as fh:
payload = json.load(fh)
except json.JSONDecodeError:
print(json.dumps({"score": 0.0, "error": "Input must be valid JSON"}))
raise SystemExit(0)

if "candidate" not in payload or not isinstance(payload["candidate"], str):
print(json.dumps({"score": 0.0, "error": "Payload must include string field 'candidate'"}))
raise SystemExit(0)

candidate = payload["candidate"]
lower = candidate.lower()
blocks = re.findall(r"```[a-zA-Z0-9_-]*\n(.*?)```", candidate, re.DOTALL)
feedback: list[str] = []

contract_safety = score_hits([
any_of(lower, ["invalid json", "malformed", "empty"]),
any_of(lower, ["score range", "[0,1]", "[0, 1]", "finite"]),
any_of(lower, ["stderr", "exit", "return code", "timeout"]),
any_of(lower, ["candidate", "example"]),
])
if contract_safety < 0.5:
feedback.append("Held-out check: include failure handling for malformed payloads and invalid scores.")

pattern_coverage = score_hits([
"judge" in lower,
"command" in lower,
"http" in lower,
"composite" in lower,
"dataset" in lower or "valset" in lower,
])
if pattern_coverage < 0.8:
feedback.append("Held-out check: cover judge, command, HTTP, composite, and dataset-aware evaluators.")

reviewability = score_hits([
any_of(lower, ["baseline", "acceptance", "threshold"]),
any_of(lower, ["discrimination", "too easy", "0.85"]),
any_of(lower, ["validate", "cross-check", "held-out"]),
len(blocks) >= 2,
])
if reviewability < 0.5:
feedback.append("Held-out check: explain how maintainers decide whether an evaluator is good enough.")

workflow_fit = score_hits([
candidate.lstrip().startswith("---"),
lower.count("\n## ") >= 4,
"--objective" in lower,
"optimize-anything" in lower,
])
if workflow_fit < 0.75:
feedback.append("Held-out check: preserve skill structure and optimize-anything command context.")

brevity = 1.0
if len(candidate) > 7000:
brevity = 0.5
feedback.append("Held-out check: guidance is long; tighten before accepting.")
elif len(candidate) < 700:
brevity = 0.4
feedback.append("Held-out check: guidance is too short to be operational.")

score = clamp(
contract_safety * 0.25
+ pattern_coverage * 0.25
+ reviewability * 0.20
+ workflow_fit * 0.20
+ brevity * 0.10
)

print(json.dumps({
"score": round(score, 4),
"contract_safety": round(contract_safety, 4),
"pattern_coverage": round(pattern_coverage, 4),
"reviewability": round(reviewability, 4),
"workflow_fit": round(workflow_fit, 4),
"brevity": round(brevity, 4),
"feedback": feedback[:5],
}))
PY
Loading
Loading