agentkit's agent package has a small, deliberately host-neutral core. Learn
these five types and the rest is mechanics.
agent reasons over exactly two shapes:
llm.Message— the wire format sent to the model (role+content).agent.Entry— one durable conversation record, the shape your storage persists and the shaper reasons over.
It never imports your storage or event model. You map your own rows onto
Entry and back. Fields agentkit doesn't interpret (Tag, Origin) are
round-tripped verbatim, so you can smuggle host provenance through without
agentkit knowing what it means.
type Entry struct {
ID string
Kind EntryKind // user | assistant | tool_call | tool_result | compaction | notification
Content string
ToolCallID string // correlates a tool_call with its tool_result
ToolName string
Tag string // opaque display label for a notification (e.g. "deploy")
Origin string // opaque host provenance; agentkit never reads it
CreatedAt int64 // ns; the ordering key
}EntryKind drives how an entry renders into an llm.Message:
| Kind | renders as |
|---|---|
KindUser |
user message |
KindAssistant |
assistant message |
KindToolResult |
tool message (named by entry ID) |
KindCompaction |
user message prefixed [compacted] |
KindNotification (and anything unrecognized) |
user message prefixed [<Tag>] |
Six methods. No host event types leak across it.
type Store interface {
ClaimPending(ctx, sessionID, at int64) (int, error) // mark pending inbox arrivals shown; return how many
Append(ctx, sessionID, Entry) error // persist one entry
Context(ctx, sessionID) ([]Entry, error) // all non-subsumed entries (any order)
Compact(ctx, sessionID, Compaction) error // write a summary marker + flag subsumed rows, atomically
}ClaimPending returning a count is how the loop knows something arrived while
it was idle — the entries themselves surface through Context. Your inbox /
publish helpers (how a message becomes pending) are yours; they aren't part
of the interface. See examples/agentkit-demo/store.go for a complete
in-memory implementation you can copy.
You construct a Session per unit of work and call Turn:
sess := &agent.Session{
SessionID: "s1",
System: "…",
Store: store,
Runner: client, // agent.LLMRunner; *llm.Client satisfies it
Tools: tools, // []llm.ToolDef advertised to the model
Dispatch: dispatch, // agent.ToolDispatcher
ChatOpts: &llm.ChatOpts{}, // optional: tool_choice, grammar, response_format
// optional seams:
Build: shaper.Build, // ContextBuilder; nil → DefaultContextBuilder (verbatim)
Preparer: preparer, // pre-turn notification revalidation
OnAssistantToken: onToken, // streamed content for SSE / live UI
OnCompaction: onCompaction, // fires when the Shaper folds history mid-turn
OnUsage: onUsage, // running token tally each round
ForcedTerminalTool: "", // name the one tool that's the session's only exit
MaxTurns: 100,
Tracer: tracer, // optional spans
}
res, err := sess.Turn(ctx) // res.Reply, res.Compactions, res.Usage{Total, Active}Turn is the loop: claim the inbox → prepare notifications → build context →
stream a completion → persist the reply → dispatch tool calls → feed results
back → repeat, until the model stops calling tools (or a terminal tool fires,
or MaxTurns).
Session.Inject(ctx, Entry) appends to this session's own log so the next
Turn renders it — the self-inbox injection primitive.
The turn boundary is a coalescing point — and this is the most useful thing the model buys you. Whatever accumulates while the model is away — a resolved async (lifted) tool result, user messages that queued, and live notifications — is delivered as one merged context on the next turn, not one turn per item.
This isn't three features bolted together; it's a single seam. Every iteration
of Turn does the same thing at the top:
ClaimPending— mark everything that queued as shown,- the
Preparerhook — drop notifications whose condition already resolved, build()— render every non-subsumedEntry, sorted byCreatedAt.
So batching (many queued messages), injection (a pushed notification),
and lifting (a tool result that arrived out-of-band, keyed by its
ToolCallID) all flow through the same Store → ClaimPending → build path and
converge into one chronological transcript the model answers in a single pass.
Concretely: a tool call parks (the dispatcher returns PendingResult); the turn
ends; while the session is idle the upstream finishes and the host injects the
real result, the user sends another message, and an MCP notification fires. The
next Turn renders all of it together — the model sees the completed job, the
new question, and the notice as one coherent context and addresses them at once.
The converge demo runs exactly this.
Merging happens at the turn boundary, not between individual tool calls of one
batch: within an iteration all tool calls dispatch back-to-back, and anything
that arrives mid-dispatch is picked up at the next iteration's ClaimPending.
A Session.Build is any ContextBuilder. The default renders history
verbatim. Shaper.Build adds three phases on top:
- pristine tail — the last N messages + M tool exchanges are always kept verbatim, regardless of size.
- LOD truncation — older oversized entries render as a short stub (an
event_idpointer + head). Pure render-time; the stored entry is untouched. - compaction — if LOD alone can't fit, summarize the oldest contiguous
prefix into a
KindCompactionmarker viaStore.Compact, then re-check.
Policy is per-model:
type ShaperPolicy struct {
BudgetTokens int
PreserveLastMessages int
PreserveLastToolCalls int
LODTruncateAboveChars int
LODHeadroomTokens int // runway kept below budget (0 → ~10k)
}Headroom, not eagerness. Every LOD/compaction rewrites the prompt prefix,
which invalidates the server's KV cache from that point — expensive. So the
Shaper leaves the prefix alone (just appends to the tail, cache intact) until
the context would cross BudgetTokens − LODHeadroomTokens, then reshapes in one
decisive pass below that line — buying ~headroom tokens of runway before the
next event, instead of re-truncating a little every turn. LOD is tried first (it
fits within headroom → done); compaction is the escalation only when LOD can't.
Compaction is surfaced, transparently. A compaction happens inside the
turn — the same turn then continues straight to the model's reply, no restart.
The event is reported as CompactionInfo{Summary, SubsumedCount, TokensBefore, TokensAfter} on TurnResult.Compactions and the OnCompaction callback, so a
host can persist the summary (a hidden field on the turn) and show what folded.
Token estimation is a pluggable TokenEstimator (default: a conservative
chars/4 heuristic). agent.Budget(contextTokens, reservePct) computes a
budget that leaves room for the response. Each Turn also reports
TurnResult.Usage (and fires OnUsage): Total = cumulative prompt+
completion tokens billed this session, Active = the size of the current
(compacted + LOD) window the model sees now.
agentkit gives you a Session and the loop. It does not decide which session
runs, when, or why, nor does it model roles or a task graph. That's your
harness. The line is intentional: a client that couldn't run a tool loop would
force every host to re-wrap one; a client that owned scheduling would force
every host to adopt its orchestration model.