Skip to content

Feat: agent mode - #2

Merged
abhijitkrm merged 3 commits into
mainfrom
feat/agent-mode
Oct 6, 2026
Merged

abhijitkrm merged 3 commits into
mainfrom
feat/agent-mode

Conversation

@tundak

@tundak tundak commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

Agent mode: end-to-end (PLAN.md §5)

Finishes agent mode from PLAN.md §5. One agent session now drives every front-end (chat TUI, line REPL, ask, and a new local web chat), with the same approval policy, redaction, and audit trail in each.

What's new

Agent core (internal/agent)

  • Streaming for Anthropic and OpenAI-compatible providers (OpenAI, Groq, Ollama, OpenRouter); --no-stream / agent.no_stream to disable.
  • Event stream (agent.Event: delta, text, tool_start, tool_result, approval) replaces the per-UI callbacks; every front-end just renders it.
  • Approval policy: ops (observe/diagnose auto, local-change confirm, on-chain confirm) and readonly (mutating tools hidden from the model and hard-blocked). Local-change can be autopiloted; on-chain can never be, enforced in code and rejected in config.
  • Shared slash commands: /mode, /approve, /model, /runbook, /tools, /profile, /audit, /reset, the same in the TUI, REPL, and web.
  • Live snapshot built from monitor.Collect (version, height, peers, signing, jail, disk, EVM drift), cached with a 60s TTL and capped at 1 KB.
  • Replay: cometcli audit replay <id> feeds recorded LLM turns and tool output back through the real loop. No LLM or node access.

Safety

  • Redaction now covers user input and the snapshot, not just tool output. Previously the raw prompt went to the LLM.
  • Mnemonic detection checks words against the BIP-39 wordlist. The old regex matched any 12 lowercase words, which would have mangled ordinary questions.
  • Per-profile redact_hosts / redact_endpoints mask hostnames and IPs.
  • COMETCLI_OFFLINE=1 disables the agent; all CLI commands and cometcli mcp keep working.
  • Local-change approvals show a config diff (net add-peer / rm-peer, snap statesync --apply) via the new toolkit.Diff.

Audit

  • Events carry a session id; LLM turns (kind: llm), the exact tool text the model saw (plus a digest), and signed tx bytes at broadcast are now logged.
  • New cometcli audit sessions.

Front-ends

  • cometcli serve [--open]: local web chat. Loopback only, one-time token link plus an HttpOnly cookie, Host-header check (prevents DNS rebinding), JSON-only POSTs (blocks cross-site form posts). Approval modals show diffs and tx details.
  • Bare cometcli opens the chat TUI when a profile exists and stdin is a terminal; otherwise it shows help.
  • TUI: deltas stream into one block, the approval modal renders diffs, and settings commands are refused while a turn is running (avoids a data race).

Bug fixes found while testing against local docker validators

  • sec.exposure reported "0 critical" on macOS. It parsed nothing (no ss, and BSD netstat prints *.26657) and treated empty as clean. It now reads docker port for docker profiles, parses BSD netstat, and errors instead of passing when it can't read any sockets. On the local network it now correctly reports 6 public ports per validator.
  • Crash on SSH dial failure: a typed-nil *SSH was closed in Context.Close. This predates this PR.
  • Groq rejected every request: ObjSchema emitted "required": null / "properties": null. Now omitted or {}, with a regression test over every registered tool.
  • ANSI color codes are stripped from tool output before it reaches the model.

Config additions (agent: in a profile)

mode, autopilot, redact_hosts, redact_endpoints, no_stream. All optional; existing profiles behave as before (ops mode).

Docs

  • TESTING.md (new): step-by-step local testing against docker validators, plus free LLM options (Groq, Ollama, OpenRouter).
  • docs/COMMANDS.md: new commands, flags, session commands, agent config.
  • .gitignore: /bin/.

Testing

  • go test ./... and go vet ./... pass; agent, serve, and TUI tests also pass under -race.
  • New tests: streaming SSE parsing (Anthropic and OpenAI against httptest), input/snapshot redaction, policy (on-chain never auto), readonly hiding and blocking, event ordering, commands, a replay round-trip that matches the live transcript, the serve HTTP surface (auth, rebinding, CSRF, on-chain approve round-trip, disconnect never hangs), diff, exposure parsers, and schema nulls.
  • Manual runs of the real binary against 4 local primium-evm docker validators with a fake streaming LLM: ask with tool calls, redaction verified in the outgoing request, audit replay with the LLM down, serve over HTTP, and a headless browser render.

Verified live

  • Groq: agent runs end to end against a local validator (streaming, tool calls, readonly mode).

Not verified

  • No live run against the Anthropic or OpenAI APIs (their streaming was tested against recorded-format SSE only).
  • TUI streaming was unit-tested but not exercised in a real terminal.
  • Diffs are attached only to the two config-editing tools; node.service and upgrade.prepare keep their existing approval details.

@abhijitkrm

Copy link
Copy Markdown
Owner

Review — solid architecture, one gap before merge

Ran the branch locally: go build, go test ./..., go vet, -race on internal/agent + internal/serve — all green. The event-stream unification (TUI/REPL/serve render the same agent.Event) is the right move, and the safety posture improved: readonly filters and hard-denies, --autopilot accepts only local-change, parseTier fails closed.

Highlights

  • serve security model is careful: loopback-only bind, constant-time token compare, SameSite=Strict HttpOnly cookie, Host-header rebinding check, JSON-only POSTs, deny-on-disconnect approvals. The JS escapes before formatting — no raw model HTML.
  • Replay is genuinely useful: deterministic re-run from audit with divergence reporting.
  • Real fixes: sec.exposure no longer silently passes on macOS (docker port + BSD netstat + error-not-clean), SSH typed-nil Close, BIP-39-verified redaction (prose no longer redacted), ObjSchema null fix.

Blocking

  1. Long-running tools are still advertised. mon.watch, mon.alerts, upgrade.watch appear in toolDefs, and toolCtx gives them WithCancel with no deadline — a model calling mon.watch hangs the turn until manual cancel. One-line fix:
for _, t := range a.Reg.All() {
    if toolkit.IsLongRunning(t) { continue }  // watchers can't complete in a turn
    if !a.Policy.Allows(t.Tier()) { continue }
    ...
}

Nits (non-blocking)

  • audit.Logger.Close(): l.sink.file panics on a zero-value Logger (l.sink == nil) — add it to the guard.
  • /model switches provider cleanly, but front-end status lines cache the old provider name — cosmetic.

Also noted: this PR fixes the same Groq ObjSchema null-schema issue found on main — dedupe whichever lands second.

Great PR overall — fix the watcher leak and I'm happy to merge.

mon.watch, mon.alerts, upgrade.watch were still advertised to the model
and run under WithCancel (no deadline) — a call would hang the turn.
Now filtered from toolDefs AND refused in execCall so even a scripted
provider cannot block a turn; regression tests cover both layers.

Also: audit.Log/Path/Close nil-sink guards on zero-value Logger, and
explicit errcheck discards in serve.go (Shutdown/Encode/Write).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@abhijitkrm

Copy link
Copy Markdown
Owner

Follow-up on the review — both items are now addressed in d62d5fe:

Blocking: long-running tools in agent turns — fixed in two layers:

  1. toolDefs() now skips toolkit.IsLongRunning(t) — mon.watch, mon.alerts, upgrade.watch are no longer advertised to the model.
  2. execCall additionally refuses long-running tools outright ("refused: … is a long-running watcher and cannot finish inside an agent turn") — so a provider that names one anyway gets a clean tool error instead of a turn parked on WithCancel forever.

Regression tests in loop_test.go cover both layers: TestLoop_LongRunningNotAdvertised (watcher absent, normal tools present) and TestLoop_LongRunningCallRefused (scripted call refused, watcher body never runs, turn completes).

CI lint (errcheck) — fixed the three findings in internal/serve/serve.go: explicit discards on srv.Shutdown, json.NewEncoder().Encode, and w.Write. golangci-lint 2.12.2 (same version as CI) reports 0 issues locally.

Also landed the nit from the earlier comment: audit.Logger Log/Path/Close now all guard l.sink == nil, so a zero-value Logger is fully safe.

Verified locally: go test ./... clean, go test -race ./internal/agent ./internal/serve clean, go vet ./... clean.

@abhijitkrm
abhijitkrm merged commit bd648d0 into main Oct 6, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants