Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,8 @@ provider usage. Every agent is visible.
or agent IDs. Other ticks reuse the current directives without a planner call.
2. **Workers resolve directives through TypeSafe Jev** using compact observations
and opaque, engine-legal action candidates. Workers share a frozen pre-action
world; their calls currently run sequentially under one tick deadline.
world; a bounded adapter runs individual calls under one tick deadline.
Jev remains sequential by default; the seam also accepts keyed native batches.
3. **The deterministic world engine validates and resolves actions** in seeded
order. A complete tick commits atomically; cancellation commits no world
changes. Provider failures retain safe attempt records and use explicit
Expand Down Expand Up @@ -89,7 +90,7 @@ are in [the screenshot guide](docs/assets/README.md).
- **Inspection:** switch between Live and Agents while the same execution
controller stays mounted. Inspect Zero strategy, worker directives, reflex
choices, validation outcomes, territory, and safe activity records.
- **Research exports:** generate compact or pretty schema-v13 JSON, download
- **Research exports:** generate compact or pretty schema-v14 JSON, download
it, or manually save the exact generated artifact to local SQLite. Exports
include bounded safe tick and provider-attempt records, including attempts
that did not produce a committed tick.
Expand Down Expand Up @@ -164,7 +165,7 @@ probe. Neither that probe nor `compare:live` runs in default tests or CI.
- [ADR 0033](docs/adr/0033-retire-legacy-multi-agent-architecture.md) — retirement
of the previous architecture

Current code reads only schema-v13 exports. Pre-swarm scenarios, snapshots, and
Current code reads only schema-v14 exports. Pre-swarm scenarios, snapshots, and
exports require an older Git revision. Historical ADRs and experiment reports
remain as decision history.

Expand Down
13 changes: 11 additions & 2 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,15 @@ call) is retained as the ablation control for the swarm comparisons, not as a
second production architecture. Historical milestones below remain as
implementation history.

## Reflex-provider follow-up — batch-capable seam

The local PR 3 change adds optional native-batch worker selection, independent
result validation, and a capped individual-provider adapter. Production Jev
retains cap 1; raising production concurrency is a separate rollout decision.
Schema-v14 exports identify shared batch dispatches and preserve aggregate
billing without per-worker allocations. No native production provider, Laya,
training, planner-cadence changes, or player features are added. See ADR 0035.

## Agent Zero planner

Agent Zero is the sole generative planner. It makes an OpenRouter planning call
Expand Down Expand Up @@ -171,14 +180,14 @@ paths or SQL, recovery, scheduling, MCP, and archive authority remain deferred.

Persistent short- and long-term objectives, compact memories, plan revision, summaries, and longer simulation runs.

_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 13. See ADRs 0033 and 0034._
_Note: per-agent strategic goals, the compact memory ledger, and the Behavior Trace introduced in this milestone were subsequently removed. The SQLite experiment archive (pre-PR-5 observability slice) remains current, updated to schema version 14. See ADRs 0033–0035._

## PR 6 — Persistent autonomous world

Scheduled turns, snapshots, replay, retries, idempotency, durable budget/attempt
ledgers, failure recovery, and operation without the World Lab browser being
open. The current process-local attempt and credit-admission ceilings are an
operator safety boundary, and their schema-v13 safe ledger can be exported to
operator safety boundary, and their schema-v14 safe ledger can be exported to
the analysis archive even when no turn committed. This is not active runtime
persistence, restart recovery, or provider-account balance enforcement.

Expand Down
61 changes: 61 additions & 0 deletions apps/game-api/src/attempt-accounting.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,24 @@ import { describe, expect, it } from 'vitest';
import { AttemptAccounting } from './attempt-accounting';

describe('AttemptAccounting', () => {
it('earmarks only available retry capacity without poisoning admission or recording unstarted work', () => {
const accounting = new AttemptAccounting(10, '0.03', '0.01');
expect(accounting.reserve(2)).toBe(true);
expect(accounting.reserveUpTo(7)).toBe(1);
expect(accounting.reserveUpTo(1)).toBe(0);
expect(accounting.snapshot()).toMatchObject({
reservedPermits: 3,
attemptsStarted: 0,
exhausted: false,
});
accounting.releaseReservations();
expect(accounting.snapshot()).toMatchObject({
reservedPermits: 0,
committedCreditExposure: '0',
exhausted: false,
});
expect(accounting.ledger()).toEqual([]);
});
it('reserves atomically and finalizes known and unknown cost exactly once', () => {
const accounting = new AttemptAccounting(3, '0.03', '0.01');
expect(accounting.reserve(3)).toBe(true);
Expand Down Expand Up @@ -60,6 +78,49 @@ describe('AttemptAccounting', () => {
});
});

it('records one billed attempt for a batch and retains all participant attribution', () => {
const accounting = new AttemptAccounting(1, '0.01', '0.01');
expect(accounting.reserve(1)).toBe(true);
const first = '11111111-1111-4111-8111-111111111111' as never;
const second = '22222222-2222-4222-8222-222222222222' as never;
const permit = accounting.startReserved({
agentId: first,
intendedTurnNumber: 1,
intendedTickNumber: 1,
kind: 'initial',
startedAt: '2026-08-13T12:00:00.000Z',
modelId: 'jev-1.13.0' as never,
reasoningProfile: 'provider-default' as never,
batch: {
id: '018f3f38-6b7d-7db7-8e95-751b4ce2681f',
members: [
{ agentId: first, intendedTurnNumber: 1 },
{ agentId: second, intendedTurnNumber: 2 },
],
},
})!;
accounting.finalize(permit, {
outcome: 'completed',
completedAt: '2026-08-13T12:00:01.000Z',
provider: {
provider: 'typesafe',
model: 'jev-1.13.0' as never,
latencyMs: 20,
costCredits: 0.007,
},
});
expect(accounting.ledger()).toHaveLength(1);
expect(accounting.ledger()[0]).toMatchObject({
batch: { members: [{ agentId: first }, { agentId: second }] },
actualCostCredits: '0.007',
});
expect(accounting.snapshot()).toMatchObject({
attemptsStarted: 1,
attemptsFinalized: 1,
knownFinalizedCostCredits: '0.007',
});
});

it('reports explicit unlimited capacity', () => {
const accounting = new AttemptAccounting(null);
expect(accounting.startAdditional()).not.toBeNull();
Expand Down
36 changes: 36 additions & 0 deletions apps/game-api/src/attempt-accounting.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ import {
type ModelAttempt,
type ModelId,
type ProviderAttemptRecord,
type ProviderAttemptBatch,
type ProviderFailure,
type ProviderMetadata,
type ReflexDecision,
Expand Down Expand Up @@ -42,6 +43,7 @@ export interface AttemptStart {
startedAt: string;
modelId: ModelId;
reasoningProfile: ReasoningProfile;
batch?: ProviderAttemptBatch;
}

export interface AttemptCompletion {
Expand Down Expand Up @@ -200,6 +202,40 @@ export class AttemptAccounting {
return permitId;
}

/**
* Earmark available retry capacity without exhausting admission on a partial
* allocation. The caller owns a fixed slot per eligible job and must never
* reassign unused slots; all reservations are released at tick completion.
*/
reserveUpTo(count: number): number {
if (!Number.isInteger(count) || count < 0)
throw new Error('Retry reservation count must be a nonnegative integer.');
if (this.#exhaustionReason !== null) return 0;
let available = 0;
while (available < count) {
const next = available + 1;
if (
this.limit !== null &&
this.#started + this.#reserved + next > this.limit
)
break;
if (
this.creditLimit !== null &&
compare(
add(
this.#committedExposure,
multiply(this.reservationCreditsPerAttempt, this.#reserved + next),
),
this.creditLimit,
) > 0
)
break;
available = next;
}
if (available > 0) this.reserve(available);
return available;
}

startAdditional(details?: AttemptStart): number | null {
if (!this.reserve(1)) return null;
return this.startReserved(details);
Expand Down
48 changes: 48 additions & 0 deletions apps/game-api/src/experiment-export.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ import {
agentIdSchema,
eventIdSchema,
h3CellSchema,
providerAttemptRecordSchema,
type AgentId,
type H3Cell,
type WorldAction,
Expand Down Expand Up @@ -72,6 +73,53 @@ function rejectedMove(tickNumber: number, agentId: AgentId, to: H3Cell) {
}

describe('movement-pattern metrics', () => {
it('counts shared batch usage once in aggregate and excludes it from per-agent attribution', () => {
const batchAttempt = providerAttemptRecordSchema.parse({
id: '018f3f38-6b7d-7db7-8e95-751b4ce2681e',
agentId: agentX,
intendedTurnNumber: 1,
intendedTickNumber: 1,
kind: 'initial',
startedAt: '2026-08-13T12:00:00.000Z',
completedAt: '2026-08-13T12:00:01.000Z',
outcome: 'completed',
modelId: 'jev-1.13.0',
reasoningProfile: 'provider-default',
reservedCredits: '0.01',
actualCostCredits: '0.006',
provider: {
provider: 'typesafe',
model: 'jev-1.13.0',
latencyMs: 12,
promptTokens: 30,
completionTokens: 2,
costCredits: 0.006,
},
batch: {
id: '018f3f38-6b7d-7db7-8e95-751b4ce2681f',
members: [
{ agentId: agentX, intendedTurnNumber: 1 },
{ agentId: agentY, intendedTurnNumber: 2 },
],
},
});
const metrics = calculateExperimentMetrics(
[],
[agentX, agentY],
[batchAttempt],
);
expect(metrics.aggregate).toMatchObject({
modelCalls: 1,
knownCostCredits: 0.006,
tokens: { promptTokens: 30, completionTokens: 2 },
});
expect(metrics.byAgent.map(({ metrics: value }) => value)).toEqual(
expect.arrayContaining([
expect.objectContaining({ modelCalls: 0, knownCostCredits: 0 }),
]),
);
});

it('walks each agent path separately for direction streaks and revisits', () => {
const { a, b, direction, opposite } = straightLine();
const yTarget = neighbors(origin).find(
Expand Down
26 changes: 20 additions & 6 deletions apps/game-api/src/experiment-export.ts
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ import {
} from './geographic-direction';

export interface ExperimentSource {
schemaVersion: 13;
schemaVersion: 14;
id: ExperimentId;
startedAt: string;
providerMode: 'openrouter' | 'scripted-test';
Expand Down Expand Up @@ -203,7 +203,13 @@ export function createExperimentExport(
const initialAgentsById = new Map(
source.initialAgents.map((agent) => [agent.id, agent]),
);
const selectedAgents = selectedAgentIds
const selectedProfileIds = new Set<AgentId>([
...selectedAgentIds,
...providerAttempts.flatMap(
(attempt) => attempt.batch?.members.map(({ agentId }) => agentId) ?? [],
),
]);
const selectedAgents = [...selectedProfileIds]
.map(
(agentId) =>
currentAgentsById.get(agentId) ?? initialAgentsById.get(agentId),
Expand Down Expand Up @@ -390,8 +396,12 @@ function filterProviderAttempts(
selected: Set<AgentId>,
tickNumbers: Set<number> | 'all',
): ProviderAttemptRecord[] {
let attempts = source.providerAttempts.filter(({ agentId }) =>
selected.has(agentId),
let attempts = source.providerAttempts.filter(
({ agentId, batch }) =>
selected.has(agentId) ||
Boolean(
batch?.members.some(({ agentId: memberId }) => selected.has(memberId)),
),
);
if (tickNumbers !== 'all')
attempts = attempts.filter(
Expand Down Expand Up @@ -521,7 +531,9 @@ function attemptMetrics(attempts: readonly ProviderAttemptRecord[]) {
).length;
const groups = new Map<string, ProviderAttemptRecord[]>();
for (const attempt of attempts) {
const key = `${attempt.agentId}:${attempt.intendedTickNumber ?? attempt.intendedTurnNumber}`;
const key = attempt.batch
? `batch:${attempt.batch.id}`
: `${attempt.agentId}:${attempt.intendedTickNumber ?? attempt.intendedTurnNumber}`;
groups.set(key, [...(groups.get(key) ?? []), attempt]);
}
const retriedGroups = [...groups.values()].filter((group) =>
Expand Down Expand Up @@ -745,7 +757,9 @@ export function calculateExperimentMetrics(
agentId,
metrics: metricCountsFor(
resolvedActions.filter((action) => action.agentId === agentId),
providerAttempts.filter((attempt) => attempt.agentId === agentId),
providerAttempts.filter(
(attempt) => !attempt.batch && attempt.agentId === agentId,
),
controlChanges,
[agentId],
agentId,
Expand Down
Loading
Loading