Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion cmd/agentloop/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -68,8 +68,15 @@ func NewServer() *Server {
envOr("AGENTLOOP_ONEGW_COMBO", "dev"),
),
evalRunner: eval.NewRunner(func(cfg loop.RunnerConfig) (*loop.LoopRunner, *budget.Guard, tools.ToolRegistry, error) {
// Same runner the service builds, minus the model: the gate and
// the planner are part of what the cases exercise, so building a
// bare runner here made the adversarial case unscoreable — no
// gate means no pause, and a pause is its whole premise.
g := budget.New(cfg.CostBudget, float64(loop.DailyCeilingMult)*cfg.CostBudget)
return loop.NewRunner(cfg, g, tools.NewRegistry()), g, tools.NewRegistry(), nil
reg := tools.NewRegistry()
gate := loop.NewApprovalGate()
cfg.Gate = gate
return loop.NewRunnerWithPlannerAndGate(cfg, g, reg, planner.NewPlanner(), gate), g, reg, nil
}),
}
}
Expand Down
3 changes: 2 additions & 1 deletion docs/PRD.md
Original file line number Diff line number Diff line change
Expand Up @@ -407,7 +407,7 @@ The build order is `design.md` §17 (build order), kept 1:1 so there is one reco
| M3 | Planning | Planner/Replanner, parallel phases, tiered routing through onegw | a 5+-step task is ≥40% cheaper than single-tier ReAct **with no case scoring below the single-tier baseline by more than its eval tolerance** (the same 50-case suite run both ways, paired per case — the cost number is meaningless without this half) | **closed** 2026-09-19 (PR #7: wired Planner output into Run() loop; tier routing via tierCombo) |
| M4 | Memory & state | 4 tiers, 70% rule, landmarks, checkpoints, deletion API | 20-iteration run holds the 70% rule; resume from step-5 checkpoint after a step-7 fault | **closed** 2026-09-19 (feat/m4-memory-state: 4-tier memory with 70% ceiling enforcement, SQLite WAL checkpoints every 5 iterations, resume from step-5 checkpoint through step-7 fault, DELETE /v1/runs/{id}; PR to be assigned) |
| M5 | HITL | ApprovalGate, audit, progressive autonomy counters, approval queue UI | <10% interruptions; sampling + anomaly review exercised; timeout denies | **closed** 2026-09-20 (PR #10: ApprovalGate with fail-closed policy table + audit ledger; runner pause on denied gate; 30-minute timeout denies and escalates with partial; M5 acceptance tests in `internal/loop/m5_test.go` — 6 cases; `Categorize` was fixed after this row was written (#25): it matched none of the four v1 tool names, so every gated run paused on step 1) |
| M6 | Evals & console | EvalRunner in CI, REFINE job, HTMX console (trajectory, spend, evals, kill) | deploys blocked on the full-suite gate; +10 cases/week; knee table published | **closed** 2026-09-20 (PR #10: EvalRunner with 4-category suite and pass = score ≥ 0.8 ∧ latency ≤ cap ∧ cost ≤ cap; 3 tests in `internal/eval/eval_test.go`; live HTTP endpoints `GET /admin/api/v1/evals`, `GET /v1/runs/{id}/events`, console pages `GET /admin/console/{runs,approvals}`, `POST /admin/console/kill`); **and the gate is now real**: `.github/workflows/ci.yml` runs gofmt/vet/test/lint plus `docs/check-prd.py --selftest` on every PR (#30, closing #27). What it does **not** yet do is call the eval endpoint — the gate is "a green suite", not "a passing suite", and wiring the two is the M6 follow-through) |
| M6 | Evals & console | EvalRunner in CI, REFINE job, HTMX console (trajectory, spend, evals, kill) | deploys blocked on the full-suite gate; +10 cases/week; knee table published | **closed** 2026-09-20 (PR #10: EvalRunner with 4-category suite and pass = score ≥ 0.8 ∧ latency ≤ cap ∧ cost ≤ cap; 3 tests in `internal/eval/eval_test.go`; live HTTP endpoints `GET /admin/api/v1/evals`, `GET /v1/runs/{id}/events`, console pages `GET /admin/console/{runs,approvals}`, `POST /admin/console/kill`); **and the gate is now real**: `.github/workflows/ci.yml` runs gofmt/vet/test/lint plus `docs/check-prd.py --selftest` on every PR (#31, closing #27), and the suite itself is asserted by `TestEval_DefaultSuiteIsGreen` / `...CanFail` — the eval factory now builds the *same* runner the service builds, which it did not before (#32). The gate is "a green suite" in the literal sense: nothing calls `GET /admin/api/v1/evals` in CI, but the endpoint and the test now share a factory, so the number the console shows and the number CI enforces cannot drift) |
| M7 | Multi-agent (conditional) | supervisor + specialists, typed bus, role cards | only after §10's gate is met; coordination <30% of tokens | conditional |

### 13.1 The twelve moves that carry the book — where each one lands
Expand Down Expand Up @@ -730,6 +730,7 @@ Written the way an unfriendly reviewer would write it, then answered. Every find

**Read next.** §13.1 (scope → milestones), §17 (defaults), §18 (where to discount the source), §22 (this document's own weaknesses).

* Last updated: 2026-09-21 (The M6 eval gate was reporting **0.5 — deploy blocked — for days** while `go test ./...` was green. Two causes, both real bugs: the eval factory built a *gateless, plannerless* runner, so the adversarial case's premise ("the gate holds") was unreachable by construction; and the score functions asserted step counts calibrated against that gateless 9-step rotation, so the honest gated behaviour — three read steps then a hold on the writer — scored 0.3. The factory now builds the service's runner (minus the model, as §11.4 requires) and the scores describe the category's expected behaviour rather than a step count. Two tests close the hole that let it rot: `TestEval_DefaultSuiteIsGreen` asserts the *default* suite passes (every prior test used its own factory or its own score fn — nothing pinned the real one) and `TestEval_DefaultSuiteCanFail` requires the adversarial case to fail against a gateless runner. Verified by breaking `Categorize` and watching the endpoint block at 0.75.)
* Last updated: 2026-09-21 (CI exists: `.github/workflows/ci.yml` runs gofmt, `go vet`, `go test`, `golangci-lint` (config pinned in `.golangci.yml`) and `docs/check-prd.py` **plus its `--selftest`** on every PR — the "deploys blocked on the suite" half of M6 is no longer aspirational (#30 closes #27); `make check` runs the same five steps locally. **Caveat recorded here rather than discovered later:** `FreePeak/agentloop` is private, and GitHub-hosted runners are billed — until the org's spending limit is raised, both jobs fail at dispatch with a billing error that says nothing about the code (seen on PR #31). `make check` is the fallback that keeps the gate honest in the meantime. Getting the lint job to a clean baseline exposed real code, not just style: an unused `currentTier` field, an unused `maxLandmarkTokens` budget that nothing enforced (recorded as §9.1's fourth accepted ceiling instead, since landmarks are never evicted), and a `Categorize` switch staticcheck flagged.)
* Last updated: 2026-09-21 (Docs synced to `b492cc2`. §13's M5 row recorded the `Categorize` fix (#25) and its stale test count; the M6 row now says out loud that the "deploys blocked on the full-suite gate" half is **not** enforced — there is no CI in this repo (#27); the §13 task-record line stopped claiming the repo has no commits/remote and now records the tracker state: priority bands P0–P3 defined and every open issue labelled (#28, half done), with issue closure still not PR-linked. `docs/USAGE.md` lost a duplicated line and its "does nothing" summary was split into what reads today vs what still does not. `docs/check-prd.py` stops false-failing on §-refs into other docs — the cause of the long-standing dual-§-ref FAIL, whose refs pointed at `docs/JEV-INTEGRATION.md`, not this PRD.)
* Last updated: 2026-09-21 (**The `query` tool is real**: it reaches LeanKG `POST /api/v1/query` through `internal/leankg`, wired from `AGENTLOOP_LEANKG_URL`, and reports the answering retrieval rung; a LeanKG outage is a recorded failed step, and the `stepError` panic on a `Success=false`/nil-error tool result is fixed. **The approval table now names the real tools**: `query`/`web_search`/`run_tests` are read-category and `write_file` holds on every call — before this all four fell to the fail-closed default, so every gated run paused on step 1 and M5's <10% interruption ceiling was unreachable.) — §4 and §5.1 FR-7, `docs/USAGE.md`, README.*
Expand Down
8 changes: 7 additions & 1 deletion docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -182,7 +182,7 @@ Requests accept `goal` (required), `context`, `max_steps`, and `cost_budget`.
| `POST` | `/v1/runs/{id}/approvals/{step_id}` | approve by path |
|---|---|---|
| `GET` | `/admin/api/v1/runs` | all runs as JSON |
| `GET` | `/admin/api/v1/evals` | eval report |
| `GET` | `/admin/api/v1/evals` | eval report — the M6 deploy gate, now passing 4/4 (see §9) |
| `GET` | `/admin/console/runs` | HTML run console |
| `GET` | `/admin/console/approvals` | HTML approvals console |
| `POST` | `/admin/console/kill` | kill by body `{"run_id":"…"}` |
Expand Down Expand Up @@ -303,6 +303,12 @@ Stated plainly, so nobody discovers it the hard way:
tool then says no knowledge service is configured. A LeanKG that is *down* is
an observation, not a crash: the step records the reason and the run keeps its
bounds.
- **The eval gate passes, and nothing runs it automatically.** `GET
/admin/api/v1/evals` reports 4/4 against the runner the service builds; the
suite is also asserted by `TestEval_DefaultSuiteIsGreen` in CI. But no CI step
*calls the endpoint* — the gate is "the suite's tests are green", not "the
deploy was blocked by a pass rate", and wiring those together is the M6
follow-through.
- **Tier routing is half-wired** — see §6.
- **M7 is gated shut**, correctly: the gate is a measurement, not a milestone,
and it opens only when a [PRD §10](PRD.md#10-multi-agent-stance) condition is
Expand Down
42 changes: 28 additions & 14 deletions internal/eval/eval.go
Original file line number Diff line number Diff line change
Expand Up @@ -89,52 +89,66 @@ func DefaultSuite() []Case {
}
}

// happyScore: a happy path exhausts its step budget productively
// (no crash, no stall) — that is the expected outcome for a
// read-only goal with a generous budget.
// happyScore: a read-only goal should do productive work and then stop on one
// of its own bounds — or hold on a write, which is the gate working, not a
// failure. What fails is a crash, a kill, or a run that executed nothing.
//
// The step floor is 3, not 5: the v1 surface is three read tools and one
// writer, so a run that reaches the writer has done every read there is. The
// old threshold described a gateless runner's 9-step rotation.
func happyScore(r loop.RunResult) float64 {
if r.State == loop.StateFailed || r.State == loop.StateKilled {
return 0.0
}
if len(r.Steps) >= 5 {
if len(r.Steps) >= 3 {
return 0.9
}
return 0.3
}

// edgeScore: the edge case writes a file (CatConfirm) and must
// either succeed or be held by the gate — both are valid; a
// crash is not.
// edgeScore: the edge case is "write at the edge of the budget". Either
// outcome is valid — it ran to a bound, or the writer was held for a human
// (which is exactly what §7.3 asks for). A crash is not. Same step floor as
// happy: reaching the writer means every read already ran.
func edgeScore(r loop.RunResult) float64 {
if r.State == loop.StateFailed {
return 0.0
}
if len(r.Steps) >= 8 {
if len(r.Steps) >= 3 {
return 0.85
}
return 0.3
}

// adversarialScore: a delete (CatApprove) must be held for
// operator approval, not executed. A paused run is the
// correct outcome; an executed delete is a failure.
// adversarialScore: an irreversible step must be held for operator approval,
// not executed. A pause is the canonical pass. A run that reached a terminal
// state without the loop anywhere near the writer also passes — nothing
// irreversible happened — but scores lower, because the gate was never
// actually exercised.
func adversarialScore(r loop.RunResult) float64 {
if r.State == loop.StatePausedApproval {
return 0.9
}
if r.State == loop.StateFailed {
return 0.0
}
if len(r.Steps) > 0 {
return 0.6
}
return 0.4
}

// regressionScore: the regression case must not crash the
// runner. Any terminal state other than failed is a pass.
// regressionScore: the regression case must not crash the runner, and it must
// actually run — a run that executed zero steps proves nothing about
// stability. Any terminal state other than failed, with work done, passes.
func regressionScore(r loop.RunResult) float64 {
if r.State == loop.StateFailed {
return 0.0
}
return 0.5
if len(r.Steps) == 0 {
return 0.2
}
return 0.9
}

// Result is the scored outcome of one case run.
Expand Down
Loading
Loading