Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,24 @@

All notable changes to GopherAgent are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/); versions follow [Semantic Versioning](https://semver.org/) — pre-1.0, breaking API changes only require a minor bump.

## [v0.43.0] — 2026-08-17

### Added

- **`agent.WithMaxParallelToolCalls` — the loop no longer dispatches a whole wave at once.** Tool calls inside a dependency wave were started with one goroutine per call and no ceiling, so the concurrency of a turn was decided entirely by the model: a reply asking for forty calls opened forty tool executions at the same instant. What that costs depends on the tools, which is exactly why the loop is the wrong place to leave it unbounded — forty concurrent HTTP fetches are fine, forty concurrent database connections exhaust a pool, and forty subprocesses are a different kind of incident. The cap is per wave, and it delays rather than drops: a call over the ceiling waits for a slot, so the wave's result set is byte-identical with or without it and only peak resource use changes. That distinction is what separates it from `MaxToolCallsPerTurn`, which bounds a turn by discarding the excess and telling the model why. `0` restores unlimited dispatch for callers who want the old behaviour. The semaphore is allocated only when it can bind — an unlimited setting, or a wave already at or below the ceiling, returns nil and skips the channel and its send/receive pair — so the ordinary small wave pays nothing for the option existing. The dependency scheduler and `<output_of:...>` argument substitution are untouched; the ceiling applies within each wave the scheduler produces, not across them. (`pkg/agent/loop_execute.go`, `pkg/agent/options.go`)
- **`tools.ErrTimeout`, `tools.DeadlineCause`, and `tools.TimedOut` — a layer can now tell its own expired budget from someone else's.** Every deadline in the tool path was a plain `context.WithTimeout`, and an expired context reports only `context.DeadlineExceeded` — the same value whether the deadline that fired was this layer's, an enclosing layer's, or a turn the caller cancelled. A layer that sets a budget cannot report one honestly without that distinction, and comparing `ctx.Err()` to `context.DeadlineExceeded` is not a test of ownership even though it reads like one. `DeadlineCause(what, d)` builds the cause to hand to `context.WithTimeoutCause`, naming the budget in its message while satisfying `errors.Is(err, ErrTimeout)`; `TimedOut(ctx)` is the ownership predicate, false when an enclosing context expired or was cancelled first because that context's own cause stays in place. One sentinel with a per-site cause rather than one exported sentinel per layer: callers get a single `errors.Is` handle, and the message still says which budget elapsed, so a model-facing report can name it without the reporting site holding the duration. The sentinel lives in `pkg/tools` because it is a tool-layer concept; placing it in `pkg/agent` would have made `pkg/tools` depend on it. `tools.WithTimeout` now wraps a failure caused by its own deadline so the error names the elapsed budget, and passes an outer deadline or cancellation through untouched. The wrapped error still unwraps to whatever the tool returned, so existing `errors.Is(err, context.DeadlineExceeded)` matches keep working. (`pkg/tools/errors.go`, `pkg/tools/middleware.go`)
- **`agent.WithRequestInvariant` and `agent.RequestViolationFunc` — the request pipeline can be checked against the history it claims to derive from.** What reaches a provider is not stored history: the token-budget policy rewrites it, and five further stages layer on the soft-landing hint, memory notes, the tool-chaining hint, the plan-mode hint and dynamic context, plus a prompt-cache stamp. Two properties keep that honest — stored history is read-only along the path, since every stage must return a derived slice rather than write through to the caller's messages, and the derivation is a pure function, so what a request carried stays reconstructable. Both were held by comments alone, and the first has been violated before: the cache stamp copies before writing precisely because it once leaked into the caller's session-loaded slice. Enabled, the option snapshots history before the call and afterwards checks that nothing wrote through, then re-derives from that snapshot and checks that every derived non-system message reached the provider in order and unchanged. System messages are excluded as the declared framing surface — four stages legitimately reshape them, one by prepending — and a non-system message absent from the derivation must carry the dynamic-context marker or it is content reaching the model that no re-derivation accounts for. Comparison is field-wise rather than reflective: media parts can carry megabytes and are compared by count, and the cache stamp is applied to the request copy by design. Off by default, and nil means no snapshot is taken at all, so the cost when unused is one nil comparison per iteration. The handler runs synchronously on the loop goroutine and the turn continues either way, which makes a handler that fails the build the intended use in development and staging. (`pkg/agent/request_invariant.go`, `pkg/agent/options.go`)

### Changed (breaking defaults)

- **Tool calls within a wave now run at most 8 at a time, where previously there was no limit.** Results are unaffected — the ceiling delays calls rather than dropping them, so the same tools run with the same arguments and the wave returns the same set — but a turn that previously issued twenty simultaneous requests now issues them eight at a time, and wall-clock for such a turn rises accordingly. Waves are normally far smaller than eight, so the change is invisible outside wide fan-outs, which are the case it exists for. Callers who genuinely want unbounded dispatch pass `WithMaxParallelToolCalls(0)`. Choosing a real default over preserving the old behaviour is deliberate: unbounded concurrency dictated by model output is a defect rather than a feature, and a ceiling nobody sets protects nobody. (`pkg/agent/loop_execute.go`)

### Fixed

- **A turn aborted mid-wave no longer saves an assistant message whose tool calls have no replies, and no longer discards the results that already completed.** When the anti-loop detector trips, it does so inside one call's goroutine while its siblings are still running or already finished. The abort path saved history and returned without ever draining the per-wave results, which produced two separate failures from one omission. The saved transcript was invalid — an assistant message carrying tool calls with no matching tool messages, a shape providers reject outright — and it survived only because a provider adapter repaired history on every call, so the defect was masked at the layer least able to explain it. Meanwhile the results that *had* completed were dropped on the floor: a query that ran, a fetch that returned, a subprocess that finished, all paid for and all invisible to the next turn, which then had every reason to request them again. Completed results are now preserved and every remaining call receives a reply naming the abort, so the model reads that those calls did not run rather than inferring silence meant success, and the transcript leaves the loop balanced by construction rather than by downstream repair. (`pkg/agent/loop_iteration.go`, `pkg/agent/loop_state.go`)
- **Cancelling a turn no longer reports the code interpreter or video generation as having exceeded their own timeouts.** Both gated their timeout message on `ctx.Err() == context.DeadlineExceeded`, which an enclosing deadline sets identically — so pressing stop told the user their code had timed out after thirty seconds when the code was fine and they had cancelled, and told them video generation had run five minutes when it had not. The message was not merely imprecise; it named the wrong cause and pointed at the wrong fix. Both now test ownership of the deadline, and an enclosing cancellation surfaces as cancellation. (`pkg/tools/builtin/code_interpreter.go`, `pkg/tools/builtin/generate_video.go`)
- **A SQL statement that exceeds its query budget now tells the model what happened instead of handing it the raw driver error.** The failure arrived as `context deadline exceeded` with nothing naming the budget or suggesting a remedy, which is unactionable at the point it matters: the model cannot distinguish it from a transient fault and retries the identical statement, spending the budget again to reach the same place. The result now names the elapsed budget and says to narrow the query, add a limit, or filter on an indexed column. An enclosing cancellation is still reported verbatim, because it is not the statement's fault and the model has nothing to fix. (`pkg/tools/builtin/sql_agent.go`)

## [v0.42.0] — 2026-08-15

### Added
Expand Down Expand Up @@ -646,6 +664,7 @@ Multi-user, long-running, audit-friendly chat surface — the foundation for sid
- README section on the permission flow — documents `RequiresConfirmation` × `ConfirmHITL` × `Permissions` interaction.
- Enum struct tag support in `tools.SchemaFor[T]()` — emit values into JSON-Schema's `enum` array so providers reject invalid values upstream.

[v0.43.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.43.0
[v0.42.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.42.0
[v0.41.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.41.0
[v0.40.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.40.0
Expand Down
Loading