Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 83 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
name: ci

# The check that makes "all green" mean something. Before this workflow the
# only thing that ran on a pull request was the gitStream app, which reports
# `skipping` — every "tests pass" claim meant a human ran `go test ./...`
# (issue #27).
#
# Two jobs, both required by branch protection:
# build-test the code: format, vet, test, lint
# prd the document: docs/check-prd.py asserts the PRD's own promises,
# and --selftest proves the check can still fail
#
# Keep the Go version in step with go.mod; `go-version-file` reads it, so
# there is one source of truth rather than two.

on:
push:
branches: [main]
pull_request:
branches: [main]
workflow_dispatch:

permissions:
contents: read

concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true

jobs:
build-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: true

# Formatting first: it is the cheapest failure and the most annoying to
# discover after review.
- name: gofmt
run: |
unformatted="$(gofmt -l ./cmd ./internal)"
if [ -n "$unformatted" ]; then
echo "::error::these files are not gofmt-clean:"
echo "$unformatted"
exit 1
fi

- name: build
run: go build ./...

- name: vet
run: go vet ./...

- name: test
run: go test ./... -count=1

- name: lint
uses: golangci/golangci-lint-action@v6
with:
version: v2.1.6

# The PRD is a build artifact of this project, so a broken one fails the same
# way broken code does. `--selftest` matters as much as the check itself: it
# asserts the checker can still fail, which is what stops it rotting into a
# script that always prints OK.
prd:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- uses: actions/setup-python@v5
with:
python-version: '3.x'

- name: PRD asserts its own promises
run: python3 docs/check-prd.py

- name: the check can still fail
run: python3 docs/check-prd.py --selftest
6 changes: 5 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,3 +1,7 @@
# The built binary (Makefile: `build` -> ./agentloop).
# `!cmd/agentloop/` un-ignores the SOURCE directory of the same name — without
# it the bare pattern above also matched the directory and silently hid every
# new file added under cmd/agentloop/ (found the hard way: cmd/agentloop/
# debug_test.go had been invisible to git).
agentloop
!cmd/agentloop/
!cmd/agentloop/
42 changes: 42 additions & 0 deletions .golangci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# golangci-lint config — pinned so CI and a local `make lint` agree.
#
# v2 schema. The default linter set is kept and only two things are changed,
# both stated rather than implied:
#
# * errcheck is excluded in _test.go. An unchecked `defer resp.Body.Close()`
# or `json.NewDecoder(...).Decode(...)` in a test is conventional and its
# failure mode is a test that fails anyway; requiring `_ =` on every one
# adds noise without adding signal. Production code is NOT excluded — those
# were fixed instead (see PR #30).
# * a handful of style checks that fight the codebase's existing shape are
# off, each with the reason inline. Nothing here turns off a correctness
# check.
version: "2"

linters:
default: standard
exclusions:
generated: lax
presets:
- comments
- common-false-positives
- legacy
rules:
# Tests: an unchecked Close/Decode is conventional there.
- path: _test\.go
linters:
- errcheck
# §-heavy prose in doc comments (PRD citations) trips nothing here on
# purpose; the PRD's own checker owns that surface.
- path: \.go
text: "comment on exported"
linters:
- revive

formatters:
enable:
- gofmt

issues:
max-issues-per-linter: 0
max-same-issues: 0
14 changes: 10 additions & 4 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ PORT ?= 8081

GO ?= go

.PHONY: help build run test lint vet fmt tidy prd check smoke clean
.PHONY: help build run test lint vet fmt fmt-check tidy prd check smoke clean

help: ## list the available targets
@grep -hE '^[a-zA-Z_-]+:.*?## ' $(MAKEFILE_LIST) \
Expand All @@ -28,10 +28,16 @@ run: build ## run the server on $(PORT)
test: ## run the test suite
$(GO) test ./...

check: test lint ## tests + lint — run this before opening a PR
check: fmt-check vet test lint prd ## everything CI runs — run this before opening a PR

lint: ## golangci-lint
golangci-lint run ./...
fmt-check: ## fail if any file is not gofmt-clean (what CI's gofmt step does)
@unformatted="$$(gofmt -l ./cmd ./internal)"; \
if [ -n "$$unformatted" ]; then \
echo "not gofmt-clean:"; echo "$$unformatted"; exit 1; \
fi

lint: ## golangci-lint (config pinned in .golangci.yml; same scope as CI)
golangci-lint run ./cmd/... ./internal/...

vet: ## go vet
$(GO) vet ./...
Expand Down
6 changes: 5 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,9 +22,13 @@ goal + budget in, answer / handoff / escalation out.
```bash
make # list targets
make run # build + serve on :8081 (onegw owns 8080)
make check # tests + lint — before a PR
make check # everything CI runs: fmt-check, vet, test, lint, prd
```

CI (`.github/workflows/ci.yml`) runs exactly those five steps on every pull
request, plus `docs/check-prd.py --selftest` — so "all green" is a check's
verdict, not a claim in a PR description.

Submit a run and watch it go to work:

```bash
Expand Down
1 change: 0 additions & 1 deletion cmd/agentloop/debug_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,6 @@ import (
"strings"
"testing"
"time"

)

func TestDebugSubmit(t *testing.T) {
Expand Down
51 changes: 36 additions & 15 deletions cmd/agentloop/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import (
"crypto/rand"
"encoding/json"
"fmt"
"log"
"net/http"
"os"
"strconv"
Expand Down Expand Up @@ -152,7 +153,7 @@ func (s *Server) submitRun(w http.ResponseWriter, r *http.Request) {
}()
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(http.StatusCreated)
json.NewEncoder(w).Encode(map[string]any{"run_id": runID, "state": loop.StateThinking})
writeJSON(w, map[string]any{"run_id": runID, "state": loop.StateThinking})
}

func (s *Server) getRun(w http.ResponseWriter, r *http.Request) {
Expand All @@ -165,7 +166,7 @@ func (s *Server) getRun(w http.ResponseWriter, r *http.Request) {
return
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(result)
writeJSON(w, result)
}

func (s *Server) killRun(w http.ResponseWriter, r *http.Request) {
Expand Down Expand Up @@ -195,7 +196,7 @@ func (s *Server) killRun(w http.ResponseWriter, r *http.Request) {
s.runs[runID] = result
s.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(result)
writeJSON(w, result)
}

func (s *Server) deleteRun(w http.ResponseWriter, r *http.Request) {
Expand Down Expand Up @@ -239,7 +240,7 @@ func (s *Server) getApprovals(w http.ResponseWriter, r *http.Request) {
return
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]any{"run_id": runID, "approvals": gate.Ledger(), "state": result.State})
writeJSON(w, map[string]any{"run_id": runID, "approvals": gate.Ledger(), "state": result.State})
}

func (s *Server) submitApproval(w http.ResponseWriter, r *http.Request) {
Expand Down Expand Up @@ -276,7 +277,7 @@ func (s *Server) submitApproval(w http.ResponseWriter, r *http.Request) {
s.mu.Unlock()
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]any{"approved": approved})
writeJSON(w, map[string]any{"approved": approved})
}

func (s *Server) submitApprovalByID(w http.ResponseWriter, r *http.Request) {
Expand All @@ -301,7 +302,7 @@ func (s *Server) submitApprovalByID(w http.ResponseWriter, r *http.Request) {
s.mu.Unlock()
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]any{"approved": approved})
writeJSON(w, map[string]any{"approved": approved})
}

func parseApprovalPath(path string) (string, int) {
Expand Down Expand Up @@ -330,11 +331,11 @@ func (s *Server) evalReportHandler(w http.ResponseWriter, r *http.Request) {
s.mu.Lock()
defer s.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(eval.Report{SuiteID: "default", Total: 0, Passed: 0, PassRate: 0, ByCategory: map[string]float64{}})
writeJSON(w, eval.Report{SuiteID: "default", Total: 0, Passed: 0, PassRate: 0, ByCategory: map[string]float64{}})
return
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(report)
writeJSON(w, report)
}

func (s *Server) runsListHandler(w http.ResponseWriter, r *http.Request) {
Expand All @@ -345,15 +346,15 @@ func (s *Server) runsListHandler(w http.ResponseWriter, r *http.Request) {
}
s.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]any{"runs": runs})
writeJSON(w, map[string]any{"runs": runs})
}

func (s *Server) consoleRunsPage(w http.ResponseWriter, r *http.Request) {
s.mu.Lock()
runCount := len(s.runs)
s.mu.Unlock()
w.Header().Set("Content-Type", "text/html")
fmt.Fprintf(w, "<html><body><h1>Runs</h1><p>Total: %d</p></body></html>", runCount)
writeConsole(w, "<html><body><h1>Runs</h1><p>Total: %d</p></body></html>", runCount)
}

func (s *Server) consoleApprovalsPage(w http.ResponseWriter, r *http.Request) {
Expand All @@ -366,7 +367,7 @@ func (s *Server) consoleApprovalsPage(w http.ResponseWriter, r *http.Request) {
}
s.mu.Unlock()
w.Header().Set("Content-Type", "text/html")
fmt.Fprintf(w, "<html><body><h1>Approvals</h1><p>Pending: %d</p></body></html>", pending)
writeConsole(w, "<html><body><h1>Approvals</h1><p>Pending: %d</p></body></html>", pending)
}

func (s *Server) consoleKillHandler(w http.ResponseWriter, r *http.Request) {
Expand Down Expand Up @@ -394,7 +395,7 @@ func (s *Server) consoleKillHandler(w http.ResponseWriter, r *http.Request) {
s.runs[runID] = result
s.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(result)
writeJSON(w, result)
}

func (s *Server) events(w http.ResponseWriter, r *http.Request) {
Expand All @@ -413,14 +414,14 @@ func (s *Server) events(w http.ResponseWriter, r *http.Request) {
for _, step := range result.Steps {
evt := map[string]any{"step_id": step.StepID, "tool": step.Tool, "phase": step.Phase}
b, _ := json.Marshal(evt)
fmt.Fprintf(w, "event: step\ndata: %s\n\n", b)
writeConsole(w, "event: step\ndata: %s\n\n", b)
if flush != nil {
flush.Flush()
}
}
done := map[string]any{"state": result.State, "exit_reason": result.ExitReason}
b, _ := json.Marshal(done)
fmt.Fprintf(w, "event: done\ndata: %s\n\n", b)
writeConsole(w, "event: done\ndata: %s\n\n", b)
if flush != nil {
flush.Flush()
}
Expand All @@ -444,6 +445,16 @@ func extractRunID(path string) string {

func boolPtr(b bool) *bool { return &b }

// writeJSON writes one JSON body. A failure here cannot be reported to the
// client — the status line and headers are already on the wire — so the error
// is explicitly discarded rather than left unchecked, and the request is
// logged so a broken client is still visible.
func writeJSON(w http.ResponseWriter, v any) {
if err := json.NewEncoder(w).Encode(v); err != nil {
log.Printf("write response: %v", err)
}
}

// m6SuiteCases returns the M6 acceptance suite (PRD §11.4) used as
// the deploy gate. It delegates to eval.DefaultSuite so the suite
// definition lives in one place (internal/eval) and the HTTP handler
Expand All @@ -452,6 +463,14 @@ func m6SuiteCases() []eval.Case {
return eval.DefaultSuite()
}

// writeConsole is writeJSON's counterpart for the HTML and SSE surfaces,
// where a failed write is equally unreportable and equally worth logging.
func writeConsole(w http.ResponseWriter, format string, args ...any) {
if _, err := fmt.Fprintf(w, format, args...); err != nil {
log.Printf("write console: %v", err)
}
}

func main() {
s := NewServer()
mux := http.NewServeMux()
Expand All @@ -475,5 +494,7 @@ func main() {
}
addr := ":" + port
fmt.Printf("agentloop listening on %s\n", addr)
http.ListenAndServe(addr, mux)
if err := http.ListenAndServe(addr, mux); err != nil {
log.Fatalf("serve: %v", err)
}
}
6 changes: 4 additions & 2 deletions docs/PRD.md
Original file line number Diff line number Diff line change
Expand Up @@ -329,11 +329,12 @@ Every default in §17 is a calibrated line item, not a book assertion, and the r

### 9.1 Known ceilings to cut at scale (marked in code)

Per convention, deliberate shortcuts ship with a `ponytail:` comment naming the ceiling and the upgrade path. v1 accepts three:
Per convention, deliberate shortcuts ship with a `ponytail:` comment naming the ceiling and the upgrade path. v1 accepts four:

1. **Run registry:** single-process, in-memory + SQLite; one shared mutex guards state transitions. Ceiling: one agentloop process per host. Upgrade: a Postgres advisory-lock run table with a worker pool, the moment a second replica is wanted.
2. **Cost accumulation:** O(n) sum over a run's spans at the 90% check (n is small). Ceiling: O(n) per step. Upgrade: a running counter on the run row once n > 500.
3. **Retention sweep:** a periodic full scan of the trace table for 90-day expiry. Ceiling: O(rows) nightly. Upgrade: an indexed `expires_at` delete when the table passes ~10M rows.
4. **The 70% rule is a working-tier guarantee, not a whole-store one.** `enforceCeiling()` (`internal/memory/memory.go`) evicts from the working tier only, and stops rather than evict anything else — the landmark tier is never evicted (P32) and retrieved is capped by its own 20%. A run that promotes enough landmarks can therefore sit above 70% and stay there. Ceiling: whole-store usage can exceed the ceiling when landmarks dominate. Upgrade: a promotion cap (reject a promotion that would push landmarks past 20%, recording it as a compressed summary) — deliberately not built, because losing a decision to a budget is the worse failure of the two.

## 10. Multi-agent stance

Expand Down Expand Up @@ -406,7 +407,7 @@ The build order is `design.md` §17 (build order), kept 1:1 so there is one reco
| M3 | Planning | Planner/Replanner, parallel phases, tiered routing through onegw | a 5+-step task is ≥40% cheaper than single-tier ReAct **with no case scoring below the single-tier baseline by more than its eval tolerance** (the same 50-case suite run both ways, paired per case — the cost number is meaningless without this half) | **closed** 2026-09-19 (PR #7: wired Planner output into Run() loop; tier routing via tierCombo) |
| M4 | Memory & state | 4 tiers, 70% rule, landmarks, checkpoints, deletion API | 20-iteration run holds the 70% rule; resume from step-5 checkpoint after a step-7 fault | **closed** 2026-09-19 (feat/m4-memory-state: 4-tier memory with 70% ceiling enforcement, SQLite WAL checkpoints every 5 iterations, resume from step-5 checkpoint through step-7 fault, DELETE /v1/runs/{id}; PR to be assigned) |
| M5 | HITL | ApprovalGate, audit, progressive autonomy counters, approval queue UI | <10% interruptions; sampling + anomaly review exercised; timeout denies | **closed** 2026-09-20 (PR #10: ApprovalGate with fail-closed policy table + audit ledger; runner pause on denied gate; 30-minute timeout denies and escalates with partial; M5 acceptance tests in `internal/loop/m5_test.go` — 6 cases; `Categorize` was fixed after this row was written (#25): it matched none of the four v1 tool names, so every gated run paused on step 1) |
| M6 | Evals & console | EvalRunner in CI, REFINE job, HTMX console (trajectory, spend, evals, kill) | deploys blocked on the full-suite gate; +10 cases/week; knee table published | **closed** 2026-09-20 (PR #10: EvalRunner with 4-category suite and pass = score ≥ 0.8 ∧ latency ≤ cap ∧ cost ≤ cap; 3 tests in `internal/eval/eval_test.go`; live HTTP endpoints `GET /admin/api/v1/evals`, `GET /v1/runs/{id}/events`, console pages `GET /admin/console/{runs,approvals}`, `POST /admin/console/kill`); **the "deploys blocked" half is not enforced yet — there is no CI in this repo, `go test ./...` is run by hand (issue #27)** |
| M6 | Evals & console | EvalRunner in CI, REFINE job, HTMX console (trajectory, spend, evals, kill) | deploys blocked on the full-suite gate; +10 cases/week; knee table published | **closed** 2026-09-20 (PR #10: EvalRunner with 4-category suite and pass = score ≥ 0.8 ∧ latency ≤ cap ∧ cost ≤ cap; 3 tests in `internal/eval/eval_test.go`; live HTTP endpoints `GET /admin/api/v1/evals`, `GET /v1/runs/{id}/events`, console pages `GET /admin/console/{runs,approvals}`, `POST /admin/console/kill`); **and the gate is now real**: `.github/workflows/ci.yml` runs gofmt/vet/test/lint plus `docs/check-prd.py --selftest` on every PR (#30, closing #27). What it does **not** yet do is call the eval endpoint — the gate is "a green suite", not "a passing suite", and wiring the two is the M6 follow-through) |
| M7 | Multi-agent (conditional) | supervisor + specialists, typed bus, role cards | only after §10's gate is met; coordination <30% of tokens | conditional |

### 13.1 The twelve moves that carry the book — where each one lands
Expand Down Expand Up @@ -729,6 +730,7 @@ Written the way an unfriendly reviewer would write it, then answered. Every find

**Read next.** §13.1 (scope → milestones), §17 (defaults), §18 (where to discount the source), §22 (this document's own weaknesses).

* Last updated: 2026-09-21 (CI exists: `.github/workflows/ci.yml` runs gofmt, `go vet`, `go test`, `golangci-lint` (config pinned in `.golangci.yml`) and `docs/check-prd.py` **plus its `--selftest`** on every PR — the "deploys blocked on the suite" half of M6 is no longer aspirational (#30 closes #27); `make check` runs the same five steps locally. **Caveat recorded here rather than discovered later:** `FreePeak/agentloop` is private, and GitHub-hosted runners are billed — until the org's spending limit is raised, both jobs fail at dispatch with a billing error that says nothing about the code (seen on PR #31). `make check` is the fallback that keeps the gate honest in the meantime. Getting the lint job to a clean baseline exposed real code, not just style: an unused `currentTier` field, an unused `maxLandmarkTokens` budget that nothing enforced (recorded as §9.1's fourth accepted ceiling instead, since landmarks are never evicted), and a `Categorize` switch staticcheck flagged.)
* Last updated: 2026-09-21 (Docs synced to `b492cc2`. §13's M5 row recorded the `Categorize` fix (#25) and its stale test count; the M6 row now says out loud that the "deploys blocked on the full-suite gate" half is **not** enforced — there is no CI in this repo (#27); the §13 task-record line stopped claiming the repo has no commits/remote and now records the tracker state: priority bands P0–P3 defined and every open issue labelled (#28, half done), with issue closure still not PR-linked. `docs/USAGE.md` lost a duplicated line and its "does nothing" summary was split into what reads today vs what still does not. `docs/check-prd.py` stops false-failing on §-refs into other docs — the cause of the long-standing dual-§-ref FAIL, whose refs pointed at `docs/JEV-INTEGRATION.md`, not this PRD.)
* Last updated: 2026-09-21 (**The `query` tool is real**: it reaches LeanKG `POST /api/v1/query` through `internal/leankg`, wired from `AGENTLOOP_LEANKG_URL`, and reports the answering retrieval rung; a LeanKG outage is a recorded failed step, and the `stepError` panic on a `Success=false`/nil-error tool result is fixed. **The approval table now names the real tools**: `query`/`web_search`/`run_tests` are read-category and `write_file` holds on every call — before this all four fell to the fail-closed default, so every gated run paused on step 1 and M5's <10% interruption ceiling was unreachable.) — §4 and §5.1 FR-7, `docs/USAGE.md`, README.*
* Last updated: 2026-09-21 (TypeSafe coupling risk re-verified: onegw `systemone` Kind is **merged** into master — `9faea01`; the open blocker is PR #110 verdict-driven combo reorder, not the merge. `docs/JEV-INTEGRATION.md` §3 rewritten to reflect the split between the merged Kind and the unmerged routing.) — §14 risk row corrected; `docs/JEV-INTEGRATION.md` §3.1/§3.4/§3.5 marked done, §7 checklist itemised.*
Expand Down
Loading
Loading