From b0085d6562d01c2ce615e9728200793d67bb09e4 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 14:47:30 +0800 Subject: [PATCH 01/51] Add Docker runtime observability foundation --- CONTRIBUTING.md | 8 + contracts/agents-api/runtime-observability.md | 73 ++++++++ .../internal/runtimeobs/identity.go | 19 +++ .../internal/runtimeobs/resolver.go | 71 ++++++++ .../internal/runtimeobs/resolver_test.go | 73 ++++++++ .../agents-api/internal/runtimeobs/sample.go | 55 ++++++ .../agents-api/internal/runtimeobs/service.go | 66 ++++++++ .../internal/runtimeobs/service_test.go | 91 ++++++++++ .../agents-api/internal/runtimeobs/source.go | 13 ++ .../internal/sandbox/docker/provider_test.go | 11 +- .../internal/sandbox/docker/resources.go | 91 ++++++++++ .../internal/sandbox/docker/resources_test.go | 158 ++++++++++++++++++ 12 files changed, 728 insertions(+), 1 deletion(-) create mode 100644 contracts/agents-api/runtime-observability.md create mode 100644 services/agents-api/internal/runtimeobs/identity.go create mode 100644 services/agents-api/internal/runtimeobs/resolver.go create mode 100644 services/agents-api/internal/runtimeobs/resolver_test.go create mode 100644 services/agents-api/internal/runtimeobs/sample.go create mode 100644 services/agents-api/internal/runtimeobs/service.go create mode 100644 services/agents-api/internal/runtimeobs/service_test.go create mode 100644 services/agents-api/internal/runtimeobs/source.go create mode 100644 services/agents-api/internal/sandbox/docker/resources.go create mode 100644 services/agents-api/internal/sandbox/docker/resources_test.go diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 58b311ac6..c979f5e62 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -197,6 +197,14 @@ starts the Runtime and its daemon authenticates and initiates the Core connectio Core verifies principal ownership and the exact Environment binding. These are management responsibilities, not separate execution architectures. +Runtime telemetry uses a separate read-only service boundary documented in +[`contracts/agents-api/runtime-observability.md`](contracts/agents-api/runtime-observability.md). +Resolve durable Session, Environment and Runtime-instance identity before selecting +a provider source. Observation never extends a lease or changes compute lifecycle. +Keep observed zero, unavailable data and unsupported Runtime modes distinct. Metrics +may inform operators, but automatic suspension requires durable Core-owned activity +state and must not use a monitoring backend as lifecycle authority. + In V1, our daemon fills the user-side executor role. Users deploy daemon, the selected harness, local tools and workspace together. Do not require Codex `exec-server`, a service-side harness, registry/Noise transport or remote tool diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md new file mode 100644 index 000000000..61d64ccc5 --- /dev/null +++ b/contracts/agents-api/runtime-observability.md @@ -0,0 +1,73 @@ +# Runtime observability contract + +This document defines the internal Runtime observation boundary. It does not add +an Agents API resource or change the pinned public protocol. + +## Ownership and identity + +Runtime telemetry is attributed to durable Core identity before it is sampled: + +```text +managed: tenant_id -> session_id -> environment_id -> runtime_allocation_id +self-hosted: tenant_id -> session_id -> environment_id -> device_id + connection_generation +none: tenant_id -> session_id (no Session-owned Runtime instance) +``` + +The managed allocation's persisted `provider_key` selects exactly one configured +observation source. A provider must independently verify the allocation labels or +equivalent ownership data. A Session, daemon connection, process, container, and +native harness Session are different identities and must not be substituted for +one another. + +The first implementation supports managed Docker allocations. `self_hosted` and +`none` are recognized but explicitly unsupported. A future self-hosted source must +use authenticated daemon telemetry fenced by the current connection generation. +Core must not attribute shared host statistics to an `environment:none` Session. + +## Sample semantics + +One sample contains: + +- `observed_at`, the provider observation time; +- `started_at`, the current compute incarnation start time; +- cumulative CPU usage in seconds; +- configured CPU capacity in cores, when known; +- current memory usage in bytes; and +- configured memory limit in bytes, when known. + +Measurements are optional. A present pointer with value zero means the provider +observed zero. An absent measurement means it was unavailable and must never be +rendered or aggregated as zero. A whole observation has one of three states: +`observed`, `unsupported`, or `unavailable`. Provider and permission failures are +errors, not ordinary unavailability. + +Docker reports cumulative cgroup CPU time and current cgroup memory usage. CPU and +memory capacity come from the inspected container configuration. Inspect and Stats +are read-only; observation must not renew, restart, create, or stop the container. +The Docker `StartedAt` value defines current compute uptime and resets after a +container restart. + +## Duration boundaries + +These durations answer different questions and must remain separate: + +- allocation age: `runtime_allocations.created_at` through `released_at` or now; +- compute uptime: provider `started_at` through `observed_at`; and +- busy Turn duration: `turns.started_at` through `completed_at` or now. + +This phase supplies compute uptime evidence and retains the existing durable +allocation and Turn timestamps. It does not infer idle time. CPU quietness, +heartbeat age, connection status, and `kept_at` are not authoritative idle state. + +Future automatic suspension requires a separate durable control model, including +an activity revision and timestamps such as `idle_since` and +`shutdown_requested_at`. Metrics, an in-memory cache, or a monitoring backend must +not become the lifecycle authority. + +## First-phase boundary + +The first phase adds no migration, public endpoint, Web view, metrics backend, +token aggregation, Kubernetes/E2B source, or automatic lifecycle action. The +internal source interface is intended to admit those providers without changing +Session attribution or the existing sandbox lifecycle interface. + diff --git a/services/agents-api/internal/runtimeobs/identity.go b/services/agents-api/internal/runtimeobs/identity.go new file mode 100644 index 000000000..7e33f4bb9 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/identity.go @@ -0,0 +1,19 @@ +package runtimeobs + +// Instance is one provider-owned Runtime incarnation. AllocationID is present +// for managed compute. DeviceID and ConnectionGeneration are reserved for a +// future authenticated self-hosted telemetry source. +type Instance struct { + AllocationID string + ProviderKey string + DeviceID string + ConnectionGeneration string +} + +// Target binds telemetry to durable Core identity. A Session is not itself a +// process or sandbox, so callers must retain the complete binding. +type Target struct { + TenantID, SessionID, EnvironmentID string + Mode Mode + Instance Instance +} diff --git a/services/agents-api/internal/runtimeobs/resolver.go b/services/agents-api/internal/runtimeobs/resolver.go new file mode 100644 index 000000000..7b1047c0c --- /dev/null +++ b/services/agents-api/internal/runtimeobs/resolver.go @@ -0,0 +1,71 @@ +package runtimeobs + +import ( + "context" + "encoding/json" + "errors" + "fmt" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" +) + +var ErrUnavailable = errors.New("Runtime observation unavailable") + +type sessionStore interface { + GetSession(context.Context, string, string) (store.Session, error) + GetRuntimeAllocation(context.Context, string, string) (store.RuntimeAllocation, error) +} + +type Resolver struct{ store sessionStore } + +func NewResolver(s sessionStore) (*Resolver, error) { + if s == nil { + return nil, errors.New("Runtime observation store is required") + } + return &Resolver{store: s}, nil +} + +func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Target, error) { + session, err := r.store.GetSession(ctx, tenantID, sessionID) + if err != nil { + return Target{}, fmt.Errorf("resolve Runtime Session: %w", err) + } + var configuration struct { + Environment *struct { + Type string `json:"type"` + } `json:"environment"` + } + if err := json.Unmarshal(session.Configuration, &configuration); err != nil || configuration.Environment == nil { + return Target{}, errors.New("invalid stored Runtime environment configuration") + } + target := Target{TenantID: session.TenantID, SessionID: session.ID, Mode: Mode(configuration.Environment.Type)} + switch target.Mode { + case ModeNone: + if session.Environment != nil { + return Target{}, errors.New("environment:none unexpectedly has a durable Environment") + } + return target, nil + case ModeSelfHosted: + if session.Environment == nil { + return Target{}, errors.New("self-hosted Session is missing its Environment") + } + target.EnvironmentID = session.Environment.ID + return target, nil + case ModeManaged: + if session.Environment == nil { + return Target{}, errors.New("managed Session is missing its Environment") + } + target.EnvironmentID = session.Environment.ID + allocation, err := r.store.GetRuntimeAllocation(ctx, tenantID, target.EnvironmentID) + if errors.Is(err, store.ErrNotFound) { + return target, ErrUnavailable + } + if err != nil { + return Target{}, fmt.Errorf("resolve Runtime allocation: %w", err) + } + target.Instance = Instance{AllocationID: allocation.ID, ProviderKey: allocation.ProviderKey, DeviceID: allocation.DeviceID} + return target, nil + default: + return Target{}, errors.New("invalid stored Runtime environment type") + } +} diff --git a/services/agents-api/internal/runtimeobs/resolver_test.go b/services/agents-api/internal/runtimeobs/resolver_test.go new file mode 100644 index 000000000..738c79f2c --- /dev/null +++ b/services/agents-api/internal/runtimeobs/resolver_test.go @@ -0,0 +1,73 @@ +package runtimeobs + +import ( + "context" + "errors" + "testing" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" +) + +type resolverStore struct { + session store.Session + allocation store.RuntimeAllocation + allocationErr error +} + +func (s resolverStore) GetSession(context.Context, string, string) (store.Session, error) { + return s.session, nil +} + +func (s resolverStore) GetRuntimeAllocation(context.Context, string, string) (store.RuntimeAllocation, error) { + return s.allocation, s.allocationErr +} + +func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { + r, err := NewResolver(resolverStore{ + session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment"}}, + allocation: store.RuntimeAllocation{ID: "allocation", ProviderKey: "provider", DeviceID: "device"}, + }) + if err != nil { + t.Fatal(err) + } + target, err := r.Resolve(t.Context(), "tenant", "session") + if err != nil { + t.Fatal(err) + } + if target.TenantID != "tenant" || target.SessionID != "session" || target.EnvironmentID != "environment" || target.Mode != ModeManaged || target.Instance.AllocationID != "allocation" || target.Instance.ProviderKey != "provider" || target.Instance.DeviceID != "device" { + t.Fatalf("incorrect managed identity binding: %+v", target) + } +} + +func TestResolverKeepsUnsupportedModesDistinct(t *testing.T) { + for _, tc := range []struct { + mode string + environment *store.Environment + }{ + {mode: "none"}, + {mode: "self_hosted", environment: &store.Environment{ID: "environment"}}, + } { + r, err := NewResolver(resolverStore{session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"` + tc.mode + `"}}`), Environment: tc.environment}}) + if err != nil { + t.Fatal(err) + } + target, err := r.Resolve(t.Context(), "tenant", "session") + if err != nil || string(target.Mode) != tc.mode { + t.Fatalf("mode %s was not resolved accurately: %+v %v", tc.mode, target, err) + } + } +} + +func TestResolverReportsManagedAllocationAsUnavailable(t *testing.T) { + r, err := NewResolver(resolverStore{ + session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment"}}, + allocationErr: store.ErrNotFound, + }) + if err != nil { + t.Fatal(err) + } + target, err := r.Resolve(t.Context(), "tenant", "session") + if !errors.Is(err, ErrUnavailable) || target.EnvironmentID != "environment" || target.Mode != ModeManaged { + t.Fatalf("allocation absence was not preserved: %+v %v", target, err) + } +} diff --git a/services/agents-api/internal/runtimeobs/sample.go b/services/agents-api/internal/runtimeobs/sample.go new file mode 100644 index 000000000..bf653c174 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/sample.go @@ -0,0 +1,55 @@ +// Package runtimeobs resolves durable Session identity to provider-owned Runtime +// observations. It is telemetry only and never owns Runtime lifecycle decisions. +package runtimeobs + +import ( + "errors" + "time" +) + +type Mode string + +const ( + ModeNone Mode = "none" + ModeSelfHosted Mode = "self_hosted" + ModeManaged Mode = "openai_hosted" +) + +type Status string + +const ( + StatusObserved Status = "observed" + StatusUnsupported Status = "unsupported" + StatusUnavailable Status = "unavailable" +) + +// Sample contains provider-neutral cumulative counters and current gauges. +// Pointer fields distinguish an observed zero from an unavailable measurement. +type Sample struct { + ObservedAt time.Time + StartedAt *time.Time + + CPUUsageSecondsTotal *float64 + CPUCapacityCores *float64 + MemoryUsageBytes *uint64 + MemoryLimitBytes *uint64 +} + +func (s Sample) validate(now time.Time) error { + if s.ObservedAt.IsZero() || s.ObservedAt.After(now) { + return errors.New("invalid Runtime observation time") + } + if s.StartedAt != nil && (s.StartedAt.IsZero() || s.StartedAt.After(s.ObservedAt)) { + return errors.New("invalid Runtime start time") + } + if s.CPUUsageSecondsTotal != nil && *s.CPUUsageSecondsTotal < 0 { + return errors.New("invalid Runtime CPU usage") + } + if s.CPUCapacityCores != nil && *s.CPUCapacityCores <= 0 { + return errors.New("invalid Runtime CPU capacity") + } + if s.MemoryLimitBytes != nil && *s.MemoryLimitBytes == 0 { + return errors.New("invalid Runtime memory limit") + } + return nil +} diff --git a/services/agents-api/internal/runtimeobs/service.go b/services/agents-api/internal/runtimeobs/service.go new file mode 100644 index 000000000..e05f7c678 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/service.go @@ -0,0 +1,66 @@ +package runtimeobs + +import ( + "context" + "errors" + "fmt" + "time" +) + +type Observation struct { + Target Target + Status Status + Sample *Sample + Reason string +} + +type Service struct { + resolver TargetResolver + sources map[string]Source + now func() time.Time +} + +func NewService(resolver TargetResolver, sources map[string]Source) (*Service, error) { + if resolver == nil { + return nil, errors.New("Runtime observation resolver is required") + } + copySources := make(map[string]Source, len(sources)) + for key, source := range sources { + if key == "" || source == nil { + return nil, errors.New("invalid Runtime observation source") + } + copySources[key] = source + } + return &Service{resolver: resolver, sources: copySources, now: time.Now}, nil +} + +func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string) (Observation, error) { + target, err := s.resolver.Resolve(ctx, tenantID, sessionID) + if errors.Is(err, ErrUnavailable) { + return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_allocation_unavailable"}, nil + } + if err != nil { + return Observation{}, err + } + if target.Mode == ModeNone || target.Mode == ModeSelfHosted { + return Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable"}, nil + } + if target.Mode != ModeManaged || target.Instance.AllocationID == "" || target.Instance.ProviderKey == "" { + return Observation{}, errors.New("invalid managed Runtime observation target") + } + source, ok := s.sources[target.Instance.ProviderKey] + if !ok { + return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_source_unavailable"}, nil + } + sample, err := source.Observe(ctx, target) + if errors.Is(err, ErrUnavailable) { + return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_sample_unavailable"}, nil + } + if err != nil { + return Observation{}, fmt.Errorf("observe Runtime: %w", err) + } + if err := sample.validate(s.now()); err != nil { + return Observation{}, err + } + return Observation{Target: target, Status: StatusObserved, Sample: &sample}, nil +} diff --git a/services/agents-api/internal/runtimeobs/service_test.go b/services/agents-api/internal/runtimeobs/service_test.go new file mode 100644 index 000000000..bff6bf33d --- /dev/null +++ b/services/agents-api/internal/runtimeobs/service_test.go @@ -0,0 +1,91 @@ +package runtimeobs + +import ( + "context" + "errors" + "testing" + "time" +) + +type fixedResolver struct { + target Target + err error +} + +func (r fixedResolver) Resolve(context.Context, string, string) (Target, error) { + return r.target, r.err +} + +type fixedSource struct { + sample Sample + err error + calls int +} + +func (s *fixedSource) Observe(context.Context, Target) (Sample, error) { + s.calls++ + return s.sample, s.err +} + +func TestServiceDoesNotCallSourcesForUnsupportedModes(t *testing.T) { + for _, mode := range []Mode{ModeNone, ModeSelfHosted} { + source := &fixedSource{} + service, err := NewService(fixedResolver{target: Target{Mode: mode}}, map[string]Source{"provider": source}) + if err != nil { + t.Fatal(err) + } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusUnsupported || source.calls != 0 { + t.Fatalf("unsupported mode touched a source: %+v %v calls=%d", observation, err, source.calls) + } + } +} + +func TestServicePreservesUnavailableAndObservedZero(t *testing.T) { + target := Target{Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider"}} + service, err := NewService(fixedResolver{target: target}, nil) + if err != nil { + t.Fatal(err) + } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusUnavailable || observation.Sample != nil { + t.Fatalf("missing source was not unavailable: %+v %v", observation, err) + } + + zeroCPU := float64(0) + zeroMemory := uint64(0) + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + source := &fixedSource{sample: Sample{ObservedAt: now, CPUUsageSecondsTotal: &zeroCPU, MemoryUsageBytes: &zeroMemory}} + service, err = NewService(fixedResolver{target: target}, map[string]Source{"provider": source}) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + observation, err = service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusObserved || observation.Sample == nil || observation.Sample.CPUUsageSecondsTotal == nil || observation.Sample.MemoryUsageBytes == nil { + t.Fatalf("observed zero was lost: %+v %v", observation, err) + } +} + +func TestServiceMapsOnlyDeclaredUnavailability(t *testing.T) { + target := Target{Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider"}} + for _, tc := range []struct { + err error + wantError bool + }{ + {err: ErrUnavailable}, + {err: errors.New("Docker permission denied"), wantError: true}, + } { + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": &fixedSource{err: tc.err}}) + if err != nil { + t.Fatal(err) + } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if (err != nil) != tc.wantError { + t.Fatalf("wrong error classification: %+v %v", observation, err) + } + if !tc.wantError && observation.Status != StatusUnavailable { + t.Fatalf("declared unavailability was not mapped: %+v", observation) + } + } +} diff --git a/services/agents-api/internal/runtimeobs/source.go b/services/agents-api/internal/runtimeobs/source.go new file mode 100644 index 000000000..4044456ab --- /dev/null +++ b/services/agents-api/internal/runtimeobs/source.go @@ -0,0 +1,13 @@ +package runtimeobs + +import "context" + +// Source reads one provider-owned Runtime instance. Implementations must verify +// ownership before returning data and must not renew, restart, or stop compute. +type Source interface { + Observe(context.Context, Target) (Sample, error) +} + +type TargetResolver interface { + Resolve(context.Context, string, string) (Target, error) +} diff --git a/services/agents-api/internal/sandbox/docker/provider_test.go b/services/agents-api/internal/sandbox/docker/provider_test.go index c99ef6387..7216208f9 100644 --- a/services/agents-api/internal/sandbox/docker/provider_test.go +++ b/services/agents-api/internal/sandbox/docker/provider_test.go @@ -12,6 +12,7 @@ import ( "testing" "time" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" "github.com/containerd/errdefs" "github.com/google/uuid" @@ -66,7 +67,8 @@ func TestDockerProviderLifecycle(t *testing.T) { t.Fatal(e) } defer c.Close() - p, e := New(c, Config{InstallationID: uuid.NewString(), Image: image, Network: "bridge", Seccomp: string(seccomp)}) + installationID := uuid.NewString() + p, e := New(c, Config{InstallationID: installationID, Image: image, Network: "bridge", Seccomp: string(seccomp)}) if e != nil { t.Fatal(e) } @@ -111,6 +113,13 @@ func TestDockerProviderLifecycle(t *testing.T) { if info.State != "running" || info.ProviderID == "" { t.Fatalf("bad compute observation: %+v", info) } + resources, e := p.Observe(ctx, runtimeobs.Target{ + TenantID: b.TenantID, SessionID: b.SessionID, EnvironmentID: b.EnvironmentID, Mode: runtimeobs.ModeManaged, + Instance: runtimeobs.Instance{AllocationID: b.AllocationID, ProviderKey: installationID, DeviceID: b.DeviceID}, + }) + if e != nil || resources.StartedAt == nil || resources.CPUUsageSecondsTotal == nil || resources.MemoryUsageBytes == nil || resources.CPUCapacityCores == nil || resources.MemoryLimitBytes == nil { + t.Fatalf("bad resource observation: %+v %v", resources, e) + } inspected, e := p.inspect(ctx, b.Reference) if e != nil { t.Fatal(e) diff --git a/services/agents-api/internal/sandbox/docker/resources.go b/services/agents-api/internal/sandbox/docker/resources.go new file mode 100644 index 000000000..2f444c961 --- /dev/null +++ b/services/agents-api/internal/sandbox/docker/resources.go @@ -0,0 +1,91 @@ +package docker + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" + "github.com/moby/moby/api/types/container" + "github.com/moby/moby/client" +) + +var _ runtimeobs.Source = (*Provider)(nil) + +// Observe is read-only. Inspect verifies allocation ownership before Docker +// statistics are requested; it never renews or changes the container. +func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runtimeobs.Sample, error) { + if target.Mode != runtimeobs.ModeManaged || target.Instance.AllocationID == "" { + return runtimeobs.Sample{}, sandbox.ErrInvalid + } + if target.Instance.ProviderKey != p.config.InstallationID { + return runtimeobs.Sample{}, sandbox.ErrOwnership + } + reference := sandbox.Reference{TenantID: target.TenantID, EnvironmentID: target.EnvironmentID, AllocationID: target.Instance.AllocationID} + inspected, err := p.inspect(ctx, reference) + if errors.Is(err, sandbox.ErrNotFound) { + return runtimeobs.Sample{}, runtimeobs.ErrUnavailable + } + if err != nil { + return runtimeobs.Sample{}, err + } + if inspected.Container.State == nil || !inspected.Container.State.Running { + return runtimeobs.Sample{}, runtimeobs.ErrUnavailable + } + result, err := p.client.ContainerStats(ctx, inspected.Container.ID, client.ContainerStatsOptions{Stream: false, IncludePreviousSample: false}) + if err != nil { + return runtimeobs.Sample{}, fmt.Errorf("read Docker Runtime statistics: %w", err) + } + defer result.Body.Close() + var stats dockerStatsResponse + if err := json.NewDecoder(result.Body).Decode(&stats); err != nil { + return runtimeobs.Sample{}, fmt.Errorf("decode Docker Runtime statistics: %w", err) + } + return sampleFromDocker(inspected.Container, stats) +} + +type dockerStatsResponse struct { + Read time.Time `json:"read"` + CPUStats *struct { + CPUUsage *struct { + TotalUsage *uint64 `json:"total_usage"` + } `json:"cpu_usage"` + } `json:"cpu_stats"` + MemoryStats *struct { + Usage *uint64 `json:"usage"` + } `json:"memory_stats"` +} + +func sampleFromDocker(inspected container.InspectResponse, stats dockerStatsResponse) (runtimeobs.Sample, error) { + if inspected.State == nil || inspected.HostConfig == nil || !inspected.State.Running || stats.Read.IsZero() { + return runtimeobs.Sample{}, errors.New("incomplete Docker Runtime observation") + } + startedAt, err := time.Parse(time.RFC3339Nano, inspected.State.StartedAt) + if err != nil || startedAt.IsZero() || startedAt.After(stats.Read) { + return runtimeobs.Sample{}, errors.New("invalid Docker Runtime start time") + } + sample := runtimeobs.Sample{ + ObservedAt: stats.Read, + StartedAt: &startedAt, + } + if stats.CPUStats != nil && stats.CPUStats.CPUUsage != nil && stats.CPUStats.CPUUsage.TotalUsage != nil { + cpuUsage := float64(*stats.CPUStats.CPUUsage.TotalUsage) / float64(time.Second) + sample.CPUUsageSecondsTotal = &cpuUsage + } + if stats.MemoryStats != nil && stats.MemoryStats.Usage != nil { + memoryUsage := *stats.MemoryStats.Usage + sample.MemoryUsageBytes = &memoryUsage + } + if inspected.HostConfig.NanoCPUs > 0 { + capacity := float64(inspected.HostConfig.NanoCPUs) / 1_000_000_000 + sample.CPUCapacityCores = &capacity + } + if inspected.HostConfig.Memory > 0 { + limit := uint64(inspected.HostConfig.Memory) + sample.MemoryLimitBytes = &limit + } + return sample, nil +} diff --git a/services/agents-api/internal/sandbox/docker/resources_test.go b/services/agents-api/internal/sandbox/docker/resources_test.go new file mode 100644 index 000000000..ea4db1096 --- /dev/null +++ b/services/agents-api/internal/sandbox/docker/resources_test.go @@ -0,0 +1,158 @@ +package docker + +import ( + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" + "github.com/moby/moby/api/types/container" + "github.com/moby/moby/client" +) + +func TestObserveVerifiesOwnershipThenReadsOneShotStats(t *testing.T) { + installationID := uuid.NewString() + target := runtimeobs.Target{ + TenantID: uuid.NewString(), SessionID: uuid.NewString(), EnvironmentID: uuid.NewString(), Mode: runtimeobs.ModeManaged, + Instance: runtimeobs.Instance{AllocationID: uuid.NewString(), ProviderKey: installationID, DeviceID: uuid.NewString()}, + } + observed := time.Now().UTC().Truncate(time.Microsecond) + started := observed.Add(-time.Minute) + statsRead := false + omitMeasurements := false + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + switch { + case r.Method == http.MethodGet && strings.HasSuffix(r.URL.Path, "/json"): + _ = json.NewEncoder(w).Encode(map[string]any{ + "Id": "container-id", + "State": map[string]any{"Status": "running", "Running": true, "StartedAt": started.Format(time.RFC3339Nano)}, + "HostConfig": map[string]any{"NanoCpus": 2_000_000_000, "Memory": 2048}, + "Config": map[string]any{"Labels": map[string]string{ + labelPrefix + "installation": installationID, + labelPrefix + "tenant": target.TenantID, + labelPrefix + "environment": target.EnvironmentID, + labelPrefix + "allocation": target.Instance.AllocationID, + }}, + }) + case r.Method == http.MethodGet && strings.HasSuffix(r.URL.Path, "/containers/container-id/stats"): + if r.URL.Query().Get("stream") != "false" || r.URL.Query().Get("one-shot") != "true" { + t.Errorf("stats request was not one-shot: %s", r.URL.RawQuery) + } + statsRead = true + if omitMeasurements { + _ = json.NewEncoder(w).Encode(map[string]any{"read": observed}) + return + } + _ = json.NewEncoder(w).Encode(container.StatsResponse{ + Read: observed, + CPUStats: container.CPUStats{CPUUsage: container.CPUUsage{TotalUsage: 1_500_000_000}}, + MemoryStats: container.MemoryStats{Usage: 1024}, + }) + default: + http.NotFound(w, r) + } + })) + defer server.Close() + c, err := client.New(client.WithHost(server.URL)) + if err != nil { + t.Fatal(err) + } + defer c.Close() + p, err := New(c, Config{InstallationID: installationID, Image: "fixture@sha256:" + strings.Repeat("a", 64), Network: "bridge", Seccomp: `{}`}) + if err != nil { + t.Fatal(err) + } + sample, err := p.Observe(t.Context(), target) + if err != nil || !statsRead || sample.CPUUsageSecondsTotal == nil || *sample.CPUUsageSecondsTotal != 1.5 || sample.MemoryUsageBytes == nil || *sample.MemoryUsageBytes != 1024 { + t.Fatalf("bad one-shot observation: %+v %v stats=%v", sample, err, statsRead) + } + omitMeasurements = true + missing, err := p.Observe(t.Context(), target) + if err != nil || missing.CPUUsageSecondsTotal != nil || missing.MemoryUsageBytes != nil || missing.CPUCapacityCores == nil || missing.MemoryLimitBytes == nil { + t.Fatalf("missing Docker measurements became zero: %+v %v", missing, err) + } + foreign := target + foreign.Instance.ProviderKey = uuid.NewString() + statsRead = false + if _, err := p.Observe(t.Context(), foreign); err == nil || statsRead { + t.Fatal("foreign provider identity reached Docker stats") + } +} + +func TestSampleFromDockerPreservesObservedZeroAndConfiguredCapacity(t *testing.T) { + observed := time.Date(2026, 9, 22, 1, 2, 3, 0, time.UTC) + started := observed.Add(-5 * time.Minute) + zero := uint64(0) + sample, err := sampleFromDocker(container.InspectResponse{ + State: &container.State{Running: true, StartedAt: started.Format(time.RFC3339Nano)}, + HostConfig: &container.HostConfig{Resources: container.Resources{NanoCPUs: 2_000_000_000, Memory: 2 * 1024 * 1024 * 1024}}, + }, testDockerStats(observed, &zero, &zero)) + if err != nil { + t.Fatal(err) + } + if sample.ObservedAt != observed || sample.StartedAt == nil || !sample.StartedAt.Equal(started) { + t.Fatalf("lost Docker observation time: %+v", sample) + } + if sample.CPUUsageSecondsTotal == nil || *sample.CPUUsageSecondsTotal != 0 || sample.MemoryUsageBytes == nil || *sample.MemoryUsageBytes != 0 { + t.Fatalf("observed zero became unavailable: %+v", sample) + } + if sample.CPUCapacityCores == nil || *sample.CPUCapacityCores != 2 || sample.MemoryLimitBytes == nil || *sample.MemoryLimitBytes != 2*1024*1024*1024 { + t.Fatalf("lost configured capacity: %+v", sample) + } +} + +func TestSampleFromDockerNormalizesCumulativeCPU(t *testing.T) { + observed := time.Now().UTC() + started := observed.Add(-time.Hour) + cpu, memory := uint64(2_500_000_000), uint64(4096) + sample, err := sampleFromDocker(container.InspectResponse{ + State: &container.State{Running: true, StartedAt: started.Format(time.RFC3339Nano)}, + HostConfig: &container.HostConfig{}, + }, testDockerStats(observed, &cpu, &memory)) + if err != nil { + t.Fatal(err) + } + if sample.CPUUsageSecondsTotal == nil || *sample.CPUUsageSecondsTotal != 2.5 || sample.MemoryUsageBytes == nil || *sample.MemoryUsageBytes != 4096 { + t.Fatalf("bad Docker normalization: %+v", sample) + } + if sample.CPUCapacityCores != nil || sample.MemoryLimitBytes != nil { + t.Fatalf("invented unconfigured capacity: %+v", sample) + } +} + +func TestSampleFromDockerRejectsIncompleteState(t *testing.T) { + observed := time.Now().UTC() + for _, inspected := range []container.InspectResponse{ + {}, + {State: &container.State{Running: false}, HostConfig: &container.HostConfig{}}, + {State: &container.State{Running: true, StartedAt: "invalid"}, HostConfig: &container.HostConfig{}}, + } { + if _, err := sampleFromDocker(inspected, dockerStatsResponse{Read: observed}); err == nil { + t.Fatal("accepted incomplete Docker state") + } + } +} + +func testDockerStats(observed time.Time, cpu, memory *uint64) dockerStatsResponse { + stats := dockerStatsResponse{Read: observed} + if cpu != nil { + stats.CPUStats = &struct { + CPUUsage *struct { + TotalUsage *uint64 `json:"total_usage"` + } `json:"cpu_usage"` + }{CPUUsage: &struct { + TotalUsage *uint64 `json:"total_usage"` + }{TotalUsage: cpu}} + } + if memory != nil { + stats.MemoryStats = &struct { + Usage *uint64 `json:"usage"` + }{Usage: memory} + } + return stats +} From 2a0715f0ea876e85d7d54127cbae097254ce23d0 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 15:39:51 +0800 Subject: [PATCH 02/51] Document runtime observability API and dashboard design --- .../agents-api/runtime-observability-api.md | 281 +++++++++++++ .../runtime-observability-design.md | 388 ++++++++++++++++++ contracts/agents-api/runtime-observability.md | 4 + 3 files changed, 673 insertions(+) create mode 100644 contracts/agents-api/runtime-observability-api.md create mode 100644 contracts/agents-api/runtime-observability-design.md diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md new file mode 100644 index 000000000..4e607a9d7 --- /dev/null +++ b/contracts/agents-api/runtime-observability-api.md @@ -0,0 +1,281 @@ +# Runtime observation API proposal + +Status: review proposal. These routes are not implemented and are not yet present +in `openapi.yaml`. + +This is an Agents Core extension, not an upstream OpenAI Agents resource. The +implementation must record that status in the coverage ledger and generated +OpenAPI contract. + +## Routes + +### List current Runtime observations + +```http +GET /v1/agents/runtime-observations?after={target_id}&limit=20&order=desc +OpenAI-Beta: agents=v1 +Authorization: Bearer ... +``` + +| Field | Rules | +| --- | --- | +| `after` | Observation ID from the previous page. Optional, supplied once. | +| `limit` | Integer 1–100, default 20. | +| `order` | `asc` or `desc`, default `desc`. | + +The list contains one current Runtime context for every Session visible to the +authenticated tenant, including explicit `none`, unsupported `self_hosted`, and +released managed contexts. Ordering uses the same Session creation-time and ID +keyset as the Session list. An observation ID is the Session UUID, so pagination +does not change when the underlying Runtime incarnation changes. Pages are not an +atomic telemetry snapshot; every row has its own `resolved_at`, and a successful +provider sample has its own `observed_at`. A client completes the entire page chain +before publishing a new Dashboard snapshot. + +```json +{ + "object": "list", + "data": [ + { + "id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", + "object": "agent.runtime_observation", + "session_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", + "environment_id": "env_...", + "mode": "openai_hosted", + "provider_type": "docker", + "instance": { + "kind": "managed_allocation", + "allocation_id": "alloc_...", + "device_id": "device_...", + "connection_generation": null + }, + "status": "observed", + "reason": null, + "allocation_created_at": 1789951200, + "resolved_at": 1789953021, + "observed_at": 1789953020, + "started_at": 1789951220, + "cpu": { + "usage_seconds_total": 482.75, + "capacity_cores": 2.0, + "usage_cores": 1.42, + "utilization_ratio": 0.71 + }, + "memory": { + "usage_bytes": 805306368, + "limit_bytes": 2147483648 + } + } + ], + "has_more": false, + "first_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", + "last_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3" +} +``` + +### Retrieve one Session's current Runtime observation + +```http +GET /v1/agents/sessions/{session_id}/runtime-observation +OpenAI-Beta: agents=v1 +Authorization: Bearer ... +``` + +This returns the same object shape as a list item. It never starts a Turn, creates +an Environment, provisions compute, renews a lease, or changes lifecycle state. + +A valid `environment:none` Session returns `200` with status `unsupported`; the +Session exists but has no attributable Runtime instance. A missing or foreign +Session returns the existing indistinguishable not-found error. + +## Resource schema + +### `RuntimeObservation` + +| Field | Type | Required | Semantics | +| --- | --- | --- | --- | +| `id` | string | yes | Session UUID; stable identity of this current-observation resource and its list cursor. | +| `object` | literal | yes | `agent.runtime_observation`. | +| `session_id` | string | yes | Authorized Core Session. | +| `environment_id` | string or null | yes | Null only for mode `none`. | +| `mode` | enum | yes | `none`, `self_hosted`, `openai_hosted`. | +| `provider_type` | string or null | yes | Forward-compatible safe source kind such as `docker`; null when no provider applies. Clients must not treat an unknown nonempty value as an error. | +| `instance` | object | yes | Provider-neutral current incarnation identity; explicit `kind=none` when no compute applies. | +| `status` | enum | yes | `observed`, `unsupported`, `unavailable`. | +| `reason` | enum or null | yes | Safe reason when status is not `observed`. | +| `allocation_created_at` | integer or null | yes | Unix seconds for managed allocation age. | +| `resolved_at` | integer | yes | Unix seconds when Core resolved identity and status for this row. | +| `observed_at` | integer or null | yes | Provider sample time; null without a sample. | +| `started_at` | integer or null | yes | Current compute incarnation start time. | +| `cpu` | object or null | yes | Null when no CPU fields were observed. | +| `memory` | object or null | yes | Null when no memory fields were observed. | + +### `RuntimeInstance` + +```json +{ + "kind": "managed_allocation", + "allocation_id": "alloc_...", + "device_id": "device_...", + "connection_generation": null +} +``` + +`kind` is `managed_allocation`, `self_hosted_connection`, or `none`. For a managed +context, `allocation_id` is the incarnation key; for self-hosted, the current +`connection_generation` is the incarnation key. Fields that do not apply are +explicit nulls. Provider-native container IDs, pod names, host paths, credentials, +and raw labels are not public fields. + +### `RuntimeCPUObservation` + +```json +{ + "usage_seconds_total": 482.75, + "capacity_cores": 2.0, + "usage_cores": 1.42, + "utilization_ratio": 0.71 +} +``` + +All fields are `number | null`. Values are finite and nonnegative; +`capacity_cores`, when present, is greater than zero. Numeric zero is observed +zero. Null is unavailable. `usage_cores` is the cumulative CPU delta divided by +the observation-time delta for two ordered samples of the same incarnation. +`utilization_ratio` is `usage_cores / capacity_cores`. It is not clamped: a value +above 1 is retained as provider/accounting evidence and is not interpreted as a +lifecycle signal. Both derived fields are null after a cache restart or whenever +either source sample is absent or invalid. The API never derives CPU rate from a +single sample. + +### `RuntimeMemoryObservation` + +```json +{ + "usage_bytes": 805306368, + "limit_bytes": 2147483648 +} +``` + +Both fields are `integer | null`. Values are nonnegative and safe JSON integers. +Zero usage is observed zero. A missing or unlimited provider limit is null. + +## Status and reason matrix + +| Status | Allowed reason | +| --- | --- | +| `observed` | null | +| `unsupported` | `runtime_mode_not_observable` | +| `unavailable` | `allocation_pending`, `runtime_not_running`, `source_not_configured`, `sample_timeout`, `sample_unavailable` | + +Ownership mismatch, malformed durable identity, corrupt provider evidence, and +authorization failure are not downgraded to unavailable rows. + +## Error responses + +Use the existing Agents API error envelope. + +| HTTP | Code | When | +| --- | --- | --- | +| 400 | `unsupported_parameter` | Unknown or duplicate query fields. | +| 400 | `invalid_request` | Empty or invalid limits, order, or malformed cursor. | +| 401 | `authentication_error` | Missing or invalid API authentication. | +| 404 | `not_found` | Missing or foreign Session/cursor, indistinguishably. | +| 429 | `rate_limit_exceeded` | Runtime sampling read budget exceeded. | +| 500 | `internal_error` | Integrity, ownership, or invalid provider evidence. | +| 503 | `execution_unavailable` | Required Runtime observation service is not configured. | + +Errors never include provider raw responses or credentials. + +## Freshness and caching + +- Return `Cache-Control: no-store`. +- An internal cache may coalesce reads for at most five seconds. +- `observed_at` is authoritative for freshness; HTTP response time is not. +- Clients mark samples stale according to their own explicit threshold. +- `ETag` is not proposed because observations change independently. + +## Client contract + +`packages/agents-client` should expose: + +```ts +type RuntimeObservationStatus = "observed" | "unsupported" | "unavailable"; +type RuntimeObservationReason = + | "runtime_mode_not_observable" + | "allocation_pending" + | "runtime_not_running" + | "source_not_configured" + | "sample_timeout" + | "sample_unavailable"; + +interface RuntimeObservation { + id: string; + object: "agent.runtime_observation"; + session_id: string; + environment_id: string | null; + mode: "none" | "self_hosted" | "openai_hosted"; + provider_type: string | null; + instance: { + kind: "managed_allocation" | "self_hosted_connection" | "none"; + allocation_id: string | null; + device_id: string | null; + connection_generation: string | null; + }; + status: RuntimeObservationStatus; + reason: RuntimeObservationReason | null; + allocation_created_at: number | null; + resolved_at: number; + observed_at: number | null; + started_at: number | null; + cpu: { + usage_seconds_total: number | null; + capacity_cores: number | null; + usage_cores: number | null; + utilization_ratio: number | null; + } | null; + memory: { + usage_bytes: number | null; + limit_bytes: number | null; + } | null; +} + +interface RuntimeObservationPage { + object: "list"; + data: RuntimeObservation[]; + has_more: boolean; + first_id: string | null; + last_id: string | null; +} + +interface RuntimeObservationClient { + list(options?: { + after?: string; + limit?: number; + order?: "asc" | "desc"; + }): Promise; + + retrieveForSession(sessionId: string): Promise; +} +``` + +The client validates every required field, enum, nullability rule, timestamp, and +finite number. Unknown additive fields are ignored. Malformed data rejects the +whole page; Web does not publish a partial snapshot. + +Web also applies a configured whole-refresh budget. If `has_more` remains true +when that budget is exhausted, it retains the prior complete snapshot and marks +the refresh incomplete; it does not publish partial values as global totals. +After both Runtime-observation and Session traversals complete, Web also requires +their Session ID sets to be identical. A mismatch caused by concurrent creation or +deletion makes the candidate incomplete and prevents publication. + +## Deliberately excluded + +- Token usage: use existing Session/Turn Usage. +- Billing and cost: product/backend concern. +- Historical series: optional later capability with a separate contract. +- Container logs and command output. +- Provider credentials or native configuration. +- Start, stop, pause, resume, restart, renew, or delete operations. +- Idle classification and automatic shutdown. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md new file mode 100644 index 000000000..70ca2fe63 --- /dev/null +++ b/contracts/agents-api/runtime-observability-design.md @@ -0,0 +1,388 @@ +# Runtime observability and Dashboard design + +Status: review proposal. Phase 1 provider abstraction and Docker sampling are +implemented; the public API, Web integration, history backend, additional +providers, and lifecycle automation described below are not implemented. + +## 1. Problem statement + +Operators need one Dashboard that answers four separate questions without +confusing their sources of truth: + +1. Which Runtime instances currently belong to which tenant, Session, and + Environment? +2. What compute is allocated and what is it consuming now? +3. How long has allocation, compute, and model work been active? +4. How many model tokens have been reported for the corresponding Sessions? + +The design must work across managed Docker now and later managed Kubernetes, +E2B, and authenticated self-hosted Runtime deployments. Metrics are operational +evidence. They must not become execution or lifecycle authority. + +## 2. Goals + +- Resolve every sample through durable Core identity before provider access. +- Keep provider-specific collection behind one source interface. +- Preserve observed zero, unavailable measurements, and unsupported modes as + different states. +- Provide a bounded read-only API suitable for Core Web and other operators. +- Let Web combine Runtime observations with existing Session and Turn usage + without copying execution truth into the browser. +- Keep current snapshots independent from an optional history backend. +- Define a safe path to future idle shutdown without implementing it implicitly. + +## 3. Non-goals + +- Redefining the pinned OpenAI Agents resources. +- Adding product users, organizations, billing, or authorization tables to Core. +- Treating a Session, daemon socket, container, pod, native harness Session, or + Turn as the same identity. +- Estimating missing CPU, memory, token, or duration values. +- Using telemetry, heartbeat age, low CPU, or Prometheus state to stop compute. +- Adding lifecycle actions to the first Dashboard release. +- Storing time-series samples in PostgreSQL. + +## 4. Source-of-truth model + +| Concern | Authority | Notes | +| --- | --- | --- | +| Tenant and Session ownership | Core database | Every public read is tenant-scoped. | +| Environment placement | Session configuration and Environment row | `none`, `self_hosted`, or `openai_hosted`. | +| Observation resource identity | Session ID | One current observation resource exists per tenant-owned Session. | +| Managed Runtime identity | `runtime_allocations` | Allocation and provider key identify the compute incarnation. | +| Self-hosted Runtime identity | Environment connection generation | Future telemetry must be generation-fenced. | +| Container/pod resource values | Selected provider source | Read-only, point-in-time evidence. | +| Turn state and busy duration | Core Turns | Never inferred from CPU. | +| Token usage | Existing Session/Turn usage | Missing native usage remains unknown. | +| Historical resource series | Optional telemetry backend | Not execution or lifecycle authority. | +| Idle shutdown decision | Future durable Core control state | Separate design and migration. | + +## 5. Identity chain + +```text +managed +tenant_id -> session_id -> environment_id -> runtime_allocation_id + -> provider_key -> provider-owned container/pod/instance + +self-hosted (future) +tenant_id -> session_id -> environment_id + -> device_id + connection_generation -> authenticated Runtime report + +none +tenant_id -> session_id + -> no Session-owned Runtime instance +``` + +Provider-native identifiers are never accepted from browser input. The resolver +starts from the authorized tenant and Session, loads the committed Environment and +allocation, and only then selects the configured source by persisted provider key. +The provider independently verifies its labels or equivalent ownership metadata. + +## 6. Component architecture + +```mermaid +flowchart LR + Web[Core Web Dashboard] --> Client[packages/agents-client] + Client --> API[Agents API read handlers] + API --> Service[runtimeobs.Service] + Service --> Resolver[durable identity resolver] + Resolver --> DB[(Core PostgreSQL)] + Service --> Registry[provider source registry] + Registry --> Docker[Docker Inspect and one-shot Stats] + Registry -. future .-> K8s[Kubernetes Metrics API or cAdvisor] + Registry -. future .-> E2B[E2B metrics adapter] + Registry -. future .-> Self[authenticated daemon telemetry] + Service -. optional export .-> Telemetry[OTLP or Prometheus pipeline] + Telemetry -. future history reads .-> History[operator history adapter] +``` + +### 6.1 `runtimeobs` + +Owns provider-neutral identity, mode resolution, source selection, sample +validation, and observation status. It must not import provider SDKs or mutate +Runtime lifecycle. + +### 6.2 Provider sources + +Each source receives a fully resolved target and returns one normalized sample. +A source must verify target ownership, make only bounded read calls, preserve +missing fields, return cumulative CPU seconds, and never create, renew, restart, +pause, or stop compute. + +Docker uses Inspect followed by non-streaming one-shot Stats. Kubernetes should +retain pod UID, container identity, and restart boundaries. E2B must use an +API-supported instance identity rather than display names. Self-hosted metrics +require authenticated daemon messages fenced by the current connection generation. + +### 6.3 API composition + +The API resolves durable rows first, samples sources with bounded concurrency, +maps only ordinary absence/timeouts to safe unavailable reasons, and fails closed +on ownership or integrity errors. It returns current observations only. + +### 6.4 Web composition + +Web loads the complete paginated Runtime observation collection before publishing +a new Dashboard snapshot. The collection follows the same Session creation-time +and ID keyset as the Session list, so a Runtime incarnation change cannot invalidate +pagination. It separately uses existing Session/Turn reads for status and tokens +and joins only by exact Session identity. Before publication, the set of Session +IDs from both complete traversals must be identical. Concurrent Session creation +or deletion can make the sets differ because the APIs have no shared snapshot +token; Web then discards the candidate, marks the refresh incomplete, and keeps +the previous successful snapshot visibly stale. + +The browser enforces a configured refresh budget for total pages, targets, and +elapsed time. Exhausting that budget is an incomplete refresh: Web retains the +previous complete snapshot and does not relabel partial aggregates as tenant-wide. + +## 7. Normalized sample + +```go +type Sample struct { + ObservedAt time.Time + StartedAt *time.Time + + CPUUsageSecondsTotal *float64 + CPUCapacityCores *float64 + MemoryUsageBytes *uint64 + MemoryLimitBytes *uint64 +} +``` + +Pointer presence is semantic. `0` means observed zero; `nil` means unavailable. +CPU percentage is derived from the delta between two cumulative samples and their +observation times. A single sample cannot truthfully supply CPU percentage. + +The API projection may additionally expose `usage_cores` and `utilization_ratio` +only when the service has two ordered samples for the same Runtime incarnation. +The process-local observation cache keeps the previous cumulative value for this +calculation. Its loss makes the derived fields temporarily null; it never changes +the cumulative source measurement or lifecycle state. + +## 8. Duration semantics + +| UI label | Calculation | Meaning | +| --- | --- | --- | +| Allocation age | allocation `created_at` to `released_at` or now | Age of Core's allocation record. | +| Compute uptime | provider `started_at` to sample `observed_at` | Age of the current compute incarnation. | +| Busy duration | Turn `started_at` to `completed_at` or now | Time model work has been active. | +| Idle duration | future durable `idle_since` | Not available in the current design. | + +Container restart resets compute uptime but not allocation age. Dashboard labels +must not collapse these values into one generic Runtime duration. + +## 9. Collection behavior + +### 9.1 Current snapshot path + +- List one current target context for each tenant-owned Session in the same stable + Session creation-time and ID order used by the Session list. A released managed + allocation remains attributable but reports `runtime_not_running`; `none` and + unsupported `self_hosted` remain explicit rows rather than disappearing. +- Default page size 20, maximum 100. +- Sample at most eight providers concurrently. +- Default per-source budget two seconds and whole-request budget ten seconds. +- Do not retry a source call inside the HTTP request. +- An optional process-local singleflight/cache may coalesce identical reads for up + to five seconds and retain the previous cumulative sample for CPU-rate + calculation. It is an optimization only and may be lost on restart. +- Do not write samples to the Core database. + +### 9.2 Error classification + +| Condition | API result | +| --- | --- | +| Mode `none` or unsupported `self_hosted` | Row status `unsupported`. | +| Managed allocation not created yet | `unavailable`, reason `allocation_pending`. | +| Owned Runtime absent or stopped | `unavailable`, reason `runtime_not_running`. | +| Source not configured | `unavailable`, reason `source_not_configured`. | +| Source deadline | `unavailable`, reason `sample_timeout`. | +| Ownership mismatch or invalid durable identity | Fail the request and log a sanitized integrity error. | +| Database/authentication failure | Existing safe API error mapping. | + +Raw Docker, Kubernetes, E2B, daemon, host, credential, or network diagnostics are +never returned to the browser. + +## 10. Historical metrics + +Current API reads and history are separate capabilities. The initial API does not +provide charts over time. A later operator-configured adapter may query an +OTLP/Prometheus-compatible backend. Core must not make that backend mandatory for +Session execution or current snapshot reads. + +Recommended instruments are: + +- `agents.runtime.cpu.usage` cumulative seconds; +- `agents.runtime.cpu.capacity` cores; +- `agents.runtime.memory.usage` bytes; +- `agents.runtime.memory.limit` bytes; +- `agents.runtime.sample` success/unavailable count; and +- `agents.runtime.sample.duration` seconds. + +Provider type, Runtime mode, and coarse status are safe low-cardinality labels. +High-cardinality identities require tenant-scoped access and retention policies; +they are not global Prometheus labels by default. + +## 11. Dashboard information architecture + +### 11.1 Overview + +- Active managed Runtime count. +- Observed CPU usage and known configured capacity. +- Observed memory usage and known limits. +- Reported Session tokens, together with the reporting Session count. +- Data freshness and source coverage. + +Aggregates include only present measurements. Each total states its denominator, +for example, `6.4 / 12 cores across 6 of 8 active Runtimes`. Unknown is never added +as zero. + +### 11.2 Runtime table + +Each row shows Session, Agent/harness when already available from the Session +snapshot, mode, observation status, CPU, memory, compute uptime, Turn state, and +reported tokens. Rows navigate to the existing Session view. No stop, restart, +pause, or delete actions appear in the first release. + +### 11.3 Detail view + +The detail surface shows exact Session/Environment/allocation identity, provider +type, observation timestamps, allocation age, compute uptime, and safe unavailable +reason. It displays only the Core-owned identifiers explicitly present in the +public contract. Provider-native container IDs, pod names, instance names, host +paths, and raw labels are never displayed. + +### 11.4 States + +- **Loading:** no previous complete Runtime snapshot. +- **Fresh:** every page loaded and each row carries its own resolution time; a + provider sample also carries its independent observation time. +- **Stale:** refresh failed; previous complete snapshot retained. +- **Unavailable row:** identity is valid, measurement is temporarily absent. +- **Unsupported row:** mode is recognized but has no qualified source. +- **Integrity failure:** do not publish a partial replacement snapshot. + +### 11.5 Web implementation shape + +The Web change belongs in Core Web, not the Core service repository. It uses +`packages/agents-client` as the only Runtime-observation transport and keeps four +seams separate: + +1. A client/parser module validates one page and exposes list and Session-scoped + retrieval methods. +2. A refresh coordinator loads all observation pages plus the canonical Session + collection, applies page/target/time budgets, requires exact equality of their + Session ID sets, and atomically swaps only a complete joined snapshot. A set + mismatch is an incomplete refresh, not a partial success. +3. A feature-local state model retains `last_complete`, current refresh status, + local filters, and the selected time range. It aborts an overlapping refresh + and marks old data stale after a failed or incomplete refresh. +4. Presentational components render summary coverage, the Runtime table, and a + Session detail surface. Trend components are absent unless a later history + capability and contract are configured. + +The initial refresh cadence is an operator-configured value, not an API guarantee. +Web pauses periodic reads when hidden, refreshes when visibility returns, and adds +jitter so multiple browsers do not synchronize. Filtering is local to the last +complete snapshot and never changes tenant authorization or provider selection. + +## 12. Token usage boundary + +Runtime observations do not duplicate token usage. Web uses the existing canonical +Session Usage snapshot and joins it to Runtime rows by `session_id`. The Dashboard +shows both total and coverage, such as `1.84M reported by 7/8 Sessions`. Missing or +incomplete native usage remains unknown. + +Cost and billing stay outside this Core API. A product may join billing in its own +authorized backend, never by exposing product credentials to Core Web. + +## 13. Security and tenancy + +- Authenticate with the existing Agents API mechanism. +- Scope resolution to the authenticated tenant before provider access. +- Do not accept provider key, allocation ID, container ID, pod UID, or device ID + as an authority-bearing query parameter. +- Bound per-request list size, concurrency, response bytes, and source deadlines; + Web separately bounds a complete multi-page refresh. +- Sanitize logs through `internal/obs/log`. +- Never return credentials, environment variables, Docker raw JSON, daemon status + payloads, host paths, image registry credentials, or backend credentials. +- Rate-limit collection separately from ordinary Session reads. + +## 14. Data model impact + +Current snapshot and Dashboard work require no migration. Existing +`runtime_allocations`, `environments`, Sessions, Turns, and Usage are sufficient. +No time-series table is proposed. + +Automatic idle shutdown is a separate feature. It requires durable fields such as +`activity_revision`, `idle_since`, and `shutdown_requested_at` with fenced state +transitions. That migration cannot read a monitoring backend as authority. + +## 15. Delivery plan + +### Phase 1: provider-neutral foundation + +- `runtimeobs` identity, resolver, source, sample, and service. +- Managed Docker Inspect/Stats source. +- CPU, memory, and current compute start time. +- Explicit unsupported and unavailable states. + +### Phase 2: current snapshot API + +- Add extension types under `contracts/agents-api/v1`. +- Add collection and Session-scoped handlers. +- Add `packages/agents-client` methods and raw HTTP/client coverage. +- Add bounded concurrency, timeout, authorization, and error tests. +- Regenerate the public OpenAPI contract. + +### Phase 3: Web Dashboard + +- Add Runtime observations as a third independent Dashboard collection. +- Publish only complete traversals and retain the previous snapshot on failure. +- Join existing Session Usage and Turn status by exact Session ID. +- Add responsive, keyboard-accessible current-resource views. + +### Phase 4: optional history + +- Add telemetry exporter and qualified operator backend. +- Define a separate history query adapter and retention/security policy. +- Add trend charts only when this capability is advertised. + +### Phase 5: additional sources + +- Kubernetes, E2B, and generation-fenced self-hosted telemetry. +- Each source requires independent mechanism and deployment acceptance. + +### Phase 6: idle policy + +- Separate durable activity and shutdown state machine. +- No automatic action until race, fencing, recovery, and operator-control + acceptance is complete. + +## 16. Acceptance criteria + +- Every observation proves tenant, Session, Environment, and Runtime-instance + association. +- Managed Docker emits correct present/absent semantics and never mutates compute. +- Unsupported modes never look like zero usage. +- Collection calls are bounded and one ordinary unavailable source invents no data. +- Ownership/integrity mismatch fails closed. +- Web never publishes a partial page traversal as a current snapshot. +- Web rejects cross-collection Session membership skew, including concurrent + Session create/delete cases, before publishing tenant-wide aggregates. +- Token totals report coverage and do not estimate missing usage. +- No lifecycle action is reachable from the first Dashboard. +- No new database table is required for current snapshots or history export. + +## 17. Review decisions required + +1. Accept the proposed API as a documented Core extension rather than an upstream + OpenAI resource. +2. Accept current snapshots without atomic cross-row time semantics; every row + exposes its own `observed_at`. +3. Confirm that history is optional and external, not a PostgreSQL sample table. +4. Confirm that the first Web release has no lifecycle controls. +5. Choose whether `self_hosted` remains visibly unsupported until authenticated + generation-fenced telemetry is qualified. diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 61d64ccc5..658a136ed 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -71,3 +71,7 @@ token aggregation, Kubernetes/E2B source, or automatic lifecycle action. The internal source interface is intended to admit those providers without changing Session attribution or the existing sandbox lifecycle interface. +The review proposal for later API and Web phases is split into the +[full design](runtime-observability-design.md) and the +[proposed public extension](runtime-observability-api.md). Neither document marks +those later phases as implemented. From 45ff7f7caa4453ee9cae0650f133a558bb6cf014 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 18:07:27 +0800 Subject: [PATCH 03/51] feat(agents-api): expose runtime observations --- contracts/agents-api/README.md | 10 + contracts/agents-api/openapi.yaml | 283 ++++++++++++++++++ .../agents-api/runtime-observability-api.md | 84 +++--- .../runtime-observability-design.md | 33 +- .../agents-api/v1/runtime_observations.go | 46 +++ packages/agents-client/src/client.test.ts | 153 ++++++++++ packages/agents-client/src/client.ts | 197 ++++++++++++ .../agents-client/src/protocol-types.test.ts | 21 ++ packages/agents-client/src/types.ts | 115 +++++++ services/agents-api/cmd/server/main.go | 21 +- services/agents-api/internal/api/handler.go | 31 +- .../agents-api/internal/api/handler_test.go | 15 +- .../internal/api/runtime_observations.go | 221 ++++++++++++++ .../internal/api/runtime_observations_test.go | 224 ++++++++++++++ .../internal/runtimeobs/identity.go | 4 + .../internal/runtimeobs/resolver.go | 19 +- .../internal/runtimeobs/resolver_test.go | 59 +++- .../agents-api/internal/runtimeobs/sample.go | 15 +- .../agents-api/internal/runtimeobs/service.go | 54 +++- .../internal/runtimeobs/service_test.go | 165 +++++++++- .../internal/sandbox/docker/resources.go | 6 +- .../internal/sandbox/docker/resources_test.go | 10 +- 22 files changed, 1674 insertions(+), 112 deletions(-) create mode 100644 contracts/agents-api/v1/runtime_observations.go create mode 100644 services/agents-api/internal/api/runtime_observations.go create mode 100644 services/agents-api/internal/api/runtime_observations_test.go diff --git a/contracts/agents-api/README.md b/contracts/agents-api/README.md index c331ddb01..5049296c6 100644 --- a/contracts/agents-api/README.md +++ b/contracts/agents-api/README.md @@ -98,6 +98,16 @@ paths start at `/vaults`, not `/agents/vaults`. | vaults | create, retrieve, list, delete | Create/retrieve/list/delete with independent tenant persistence, stored status filtering, atomic Credential cascade and frozen Session attachments; archive semantics and full hosted lifecycle parity remain missing | | vaults.credentials | create, retrieve, update, list, delete | Static-bearer create/retrieve/list/token replacement/deletion with scoped encrypted storage; Session attachment and exact-URL HTTPS MCP binding; OAuth, archive semantics and full hosted lifecycle parity remain missing | +## Core extension inventory + +The operations below are implemented public Core extensions. They are excluded +from the 42-operation upstream inventory and must not be counted as OpenAI Agents +compatibility. + +| Extension | Operations | Current coverage | +| --- | --- | --- | +| Runtime observations | `GET /v1/agents/runtime-observations`; `GET /v1/agents/sessions/{session_id}/runtime-observation` | Current, read-only, tenant-scoped Session contexts with stable Session-keyset pagination, bounded concurrent sampling, Docker metrics, explicit unsupported/unavailable states, strict `packages/agents-client` projection, and no lifecycle mutation. Kubernetes, E2B, self-hosted telemetry, history, CPU-rate derivation, and automatic idle policy remain unimplemented. See [Runtime observation API](runtime-observability-api.md). | + For each resource, verify the referenced request/response unions and observable behavior, not just the route. Non-text initial input, configuration options, text/image content, function results, environment variants, full Item/SSE diff --git a/contracts/agents-api/openapi.yaml b/contracts/agents-api/openapi.yaml index 3578bf4ff..90a206045 100644 --- a/contracts/agents-api/openapi.yaml +++ b/contracts/agents-api/openapi.yaml @@ -868,6 +868,180 @@ definitions: required: - type type: object + v1.RuntimeCPUObservation: + properties: + capacity_cores: + minimum: 5e-324 + type: number + x-nullable: true + usage_cores: + minimum: 0 + type: number + x-nullable: true + usage_seconds_total: + minimum: 0 + type: number + x-nullable: true + utilization_ratio: + minimum: 0 + type: number + x-nullable: true + required: + - capacity_cores + - usage_cores + - usage_seconds_total + - utilization_ratio + type: object + v1.RuntimeInstance: + properties: + allocation_id: + format: uuid + type: string + x-nullable: true + connection_generation: + format: uuid + type: string + x-nullable: true + device_id: + format: uuid + type: string + x-nullable: true + kind: + enum: + - managed_allocation + - self_hosted_connection + - none + type: string + required: + - allocation_id + - connection_generation + - device_id + - kind + type: object + v1.RuntimeMemoryObservation: + properties: + limit_bytes: + minimum: 1 + type: integer + x-nullable: true + usage_bytes: + minimum: 0 + type: integer + x-nullable: true + required: + - limit_bytes + - usage_bytes + type: object + v1.RuntimeObservation: + properties: + allocation_created_at: + minimum: 0 + type: integer + x-nullable: true + cpu: + allOf: + - $ref: '#/definitions/v1.RuntimeCPUObservation' + x-nullable: true + environment_id: + format: uuid + type: string + x-nullable: true + id: + format: uuid + type: string + instance: + $ref: '#/definitions/v1.RuntimeInstance' + memory: + allOf: + - $ref: '#/definitions/v1.RuntimeMemoryObservation' + x-nullable: true + mode: + enum: + - none + - self_hosted + - openai_hosted + type: string + object: + enum: + - agent.runtime_observation + type: string + observed_at: + minimum: 0 + type: integer + x-nullable: true + provider_type: + type: string + x-nullable: true + reason: + enum: + - runtime_mode_not_observable + - allocation_pending + - runtime_not_running + - source_not_configured + - sample_timeout + - sample_unavailable + type: string + x-nullable: true + resolved_at: + minimum: 0 + type: integer + session_id: + format: uuid + type: string + started_at: + minimum: 0 + type: integer + x-nullable: true + status: + enum: + - observed + - unsupported + - unavailable + type: string + required: + - allocation_created_at + - cpu + - environment_id + - id + - instance + - memory + - mode + - object + - observed_at + - provider_type + - reason + - resolved_at + - session_id + - started_at + - status + type: object + v1.RuntimeObservationList: + properties: + data: + items: + $ref: '#/definitions/v1.RuntimeObservation' + type: array + first_id: + format: uuid + type: string + x-nullable: true + has_more: + type: boolean + last_id: + format: uuid + type: string + x-nullable: true + object: + enum: + - list + type: string + required: + - data + - first_id + - has_more + - last_id + - object + type: object v1.SavedAgent: properties: created_at: @@ -2581,6 +2755,68 @@ paths: summary: Update an Environment Template tags: - Environment Templates + /agents/runtime-observations: + get: + description: Core extension listing one current Runtime context per tenant-owned + Session in Session creation order. Each row has an independent resolved_at + and optional provider observed_at; the page is not an atomic telemetry snapshot. + parameters: + - description: agents=v1 + in: header + name: OpenAI-Beta + required: true + type: string + - description: Last observation ID from the previous page + in: query + name: after + type: string + - default: 20 + description: Page size + in: query + maximum: 100 + minimum: 1 + name: limit + type: integer + - default: desc + description: Session creation order + enum: + - asc + - desc + in: query + name: order + type: string + produces: + - application/json + responses: + "200": + description: OK + schema: + $ref: '#/definitions/v1.RuntimeObservationList' + "400": + description: Bad Request + schema: + $ref: '#/definitions/v1.ErrorResponse' + "401": + description: Unauthorized + schema: + $ref: '#/definitions/v1.ErrorResponse' + "404": + description: Not Found + schema: + $ref: '#/definitions/v1.ErrorResponse' + "500": + description: Internal Server Error + schema: + $ref: '#/definitions/v1.ErrorResponse' + "503": + description: Service Unavailable + schema: + $ref: '#/definitions/v1.ErrorResponse' + security: + - BearerAuth: [] + summary: List current Runtime observations + tags: + - Runtime observations /agents/sessions: get: description: Cursor and results are scoped to the authenticated execution tenant. @@ -3348,6 +3584,53 @@ paths: summary: List persisted execution Items tags: - Items + /agents/sessions/{session_id}/runtime-observation: + get: + description: Core extension returning one tenant-scoped, read-only current Runtime + observation. It never provisions, renews, restarts, pauses or stops compute. + parameters: + - description: agents=v1 + in: header + name: OpenAI-Beta + required: true + type: string + - description: Session ID + in: path + name: session_id + required: true + type: string + produces: + - application/json + responses: + "200": + description: OK + schema: + $ref: '#/definitions/v1.RuntimeObservation' + "400": + description: Bad Request + schema: + $ref: '#/definitions/v1.ErrorResponse' + "401": + description: Unauthorized + schema: + $ref: '#/definitions/v1.ErrorResponse' + "404": + description: Not Found + schema: + $ref: '#/definitions/v1.ErrorResponse' + "500": + description: Internal Server Error + schema: + $ref: '#/definitions/v1.ErrorResponse' + "503": + description: Service Unavailable + schema: + $ref: '#/definitions/v1.ErrorResponse' + security: + - BearerAuth: [] + summary: Retrieve a Session Runtime observation + tags: + - Runtime observations /agents/sessions/{session_id}/subagents: get: description: Includes nested and closed Subagents. Cursors belong to the same diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 4e607a9d7..34394ffd3 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -1,7 +1,9 @@ # Runtime observation API proposal -Status: review proposal. These routes are not implemented and are not yet present -in `openapi.yaml`. +Status: Phase 2 implemented. The current-snapshot routes, strict +`packages/agents-client` projection, and generated `openapi.yaml` contract are +implemented. Core Web integration, historical queries, and lifecycle controls +remain outside this phase. This is an Agents Core extension, not an upstream OpenAI Agents resource. The implementation must record that status in the coverage ledger and generated @@ -40,13 +42,13 @@ before publishing a new Dashboard snapshot. "id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", "object": "agent.runtime_observation", "session_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", - "environment_id": "env_...", + "environment_id": "6c02fb71-5fa8-4298-93e8-57c6625a3fc2", "mode": "openai_hosted", "provider_type": "docker", "instance": { "kind": "managed_allocation", - "allocation_id": "alloc_...", - "device_id": "device_...", + "allocation_id": "d23ab94e-e40b-45bd-93a2-444f1f74642b", + "device_id": "2e434f4f-76aa-4e54-a707-4757036d90ef", "connection_generation": null }, "status": "observed", @@ -58,8 +60,8 @@ before publishing a new Dashboard snapshot. "cpu": { "usage_seconds_total": 482.75, "capacity_cores": 2.0, - "usage_cores": 1.42, - "utilization_ratio": 0.71 + "usage_cores": null, + "utilization_ratio": null }, "memory": { "usage_bytes": 805306368, @@ -181,7 +183,6 @@ Use the existing Agents API error envelope. | 400 | `invalid_request` | Empty or invalid limits, order, or malformed cursor. | | 401 | `authentication_error` | Missing or invalid API authentication. | | 404 | `not_found` | Missing or foreign Session/cursor, indistinguishably. | -| 429 | `rate_limit_exceeded` | Runtime sampling read budget exceeded. | | 500 | `internal_error` | Integrity, ownership, or invalid provider evidence. | | 503 | `execution_unavailable` | Required Runtime observation service is not configured. | @@ -190,14 +191,15 @@ Errors never include provider raw responses or credentials. ## Freshness and caching - Return `Cache-Control: no-store`. -- An internal cache may coalesce reads for at most five seconds. +- The Phase 2 implementation performs bounded direct reads and has no observation + cache. A later internal cache may coalesce reads for at most five seconds. - `observed_at` is authoritative for freshness; HTTP response time is not. - Clients mark samples stale according to their own explicit threshold. - `ETag` is not proposed because observations change independently. ## Client contract -`packages/agents-client` should expose: +`packages/agents-client` exposes: ```ts type RuntimeObservationStatus = "observed" | "unsupported" | "unavailable"; @@ -209,38 +211,13 @@ type RuntimeObservationReason = | "sample_timeout" | "sample_unavailable"; -interface RuntimeObservation { - id: string; - object: "agent.runtime_observation"; - session_id: string; - environment_id: string | null; - mode: "none" | "self_hosted" | "openai_hosted"; - provider_type: string | null; - instance: { - kind: "managed_allocation" | "self_hosted_connection" | "none"; - allocation_id: string | null; - device_id: string | null; - connection_generation: string | null; - }; - status: RuntimeObservationStatus; - reason: RuntimeObservationReason | null; - allocation_created_at: number | null; - resolved_at: number; - observed_at: number | null; - started_at: number | null; - cpu: { - usage_seconds_total: number | null; - capacity_cores: number | null; - usage_cores: number | null; - utilization_ratio: number | null; - } | null; - memory: { - usage_bytes: number | null; - limit_bytes: number | null; - } | null; -} +type RuntimeObservation = + | RuntimeObservedObservation + | RuntimeUnavailableObservation + | RuntimeNoneObservation + | RuntimeSelfHostedObservation; -interface RuntimeObservationPage { +interface RuntimeObservationList { object: "list"; data: RuntimeObservation[]; has_more: boolean; @@ -248,20 +225,33 @@ interface RuntimeObservationPage { last_id: string | null; } -interface RuntimeObservationClient { - list(options?: { +interface AgentCore { + listRuntimeObservations(options?: { after?: string; limit?: number; order?: "asc" | "desc"; - }): Promise; + }): Promise; - retrieveForSession(sessionId: string): Promise; + retrieveRuntimeObservation(sessionId: string): Promise; } ``` +These exported variants discriminate on `status` and `mode`; their instance, +reason, timestamps, CPU, and memory fields narrow accordingly. The exact variant +definitions live in `packages/agents-client/src/types.ts` and mirror the status +and reason matrix above. + The client validates every required field, enum, nullability rule, timestamp, and -finite number. Unknown additive fields are ignored. Malformed data rejects the -whole page; Web does not publish a partial snapshot. +finite number. The current pinned contract rejects unknown additive fields so an +unreviewed server expansion cannot silently cross the browser boundary. Malformed +data rejects the whole page; Web does not publish a partial snapshot. + +The generated OpenAPI 2 schema records field-level required/nullability rules, +UUID formats, reason enums, and numeric minima. OpenAPI 2 +cannot encode the complete cross-field discriminated union. The matrix above is +normative for wire consumers; the server projection and strict TypeScript +projector enforce it, and the exported TypeScript type prevents invalid +status/mode combinations in typed consumers. Web also applies a configured whole-refresh budget. If `has_more` remains true when that budget is exhausted, it retains the prior complete snapshot and marks diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 70ca2fe63..dd9dfa16d 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -1,8 +1,8 @@ # Runtime observability and Dashboard design -Status: review proposal. Phase 1 provider abstraction and Docker sampling are -implemented; the public API, Web integration, history backend, additional -providers, and lifecycle automation described below are not implemented. +Status: Phase 1 provider abstraction/Docker sampling and Phase 2 current-snapshot +API/client contract are implemented. Core Web integration, history backend, +additional providers, and lifecycle automation described below are not implemented. ## 1. Problem statement @@ -156,9 +156,11 @@ observation times. A single sample cannot truthfully supply CPU percentage. The API projection may additionally expose `usage_cores` and `utilization_ratio` only when the service has two ordered samples for the same Runtime incarnation. -The process-local observation cache keeps the previous cumulative value for this -calculation. Its loss makes the derived fields temporarily null; it never changes -the cumulative source measurement or lifecycle state. +A future process-local observation cache may keep the previous cumulative value +for this calculation. Phase 2 intentionally leaves both derived fields null +because it has only one provider sample per request. Cache loss must make the +derived fields temporarily null; it must never change the cumulative source +measurement or lifecycle state. ## 8. Duration semantics @@ -324,6 +326,8 @@ transitions. That migration cannot read a monitoring backend as authority. ### Phase 1: provider-neutral foundation +Implemented in the Docker observability foundation. + - `runtimeobs` identity, resolver, source, sample, and service. - Managed Docker Inspect/Stats source. - CPU, memory, and current compute start time. @@ -331,6 +335,10 @@ transitions. That migration cannot read a monitoring backend as authority. ### Phase 2: current snapshot API +Implemented by the Runtime Observation extension routes and +`packages/agents-client`. The generated OpenAPI contract records the extension; +this does not add an upstream OpenAI operation. + - Add extension types under `contracts/agents-api/v1`. - Add collection and Session-scoped handlers. - Add `packages/agents-client` methods and raw HTTP/client coverage. @@ -376,13 +384,12 @@ transitions. That migration cannot read a monitoring backend as authority. - No lifecycle action is reachable from the first Dashboard. - No new database table is required for current snapshots or history export. -## 17. Review decisions required +## 17. Recorded design decisions -1. Accept the proposed API as a documented Core extension rather than an upstream - OpenAI resource. -2. Accept current snapshots without atomic cross-row time semantics; every row +1. The API is a documented Core extension rather than an upstream OpenAI resource. +2. Current snapshots have no atomic cross-row time semantics; every row exposes its own `observed_at`. -3. Confirm that history is optional and external, not a PostgreSQL sample table. -4. Confirm that the first Web release has no lifecycle controls. -5. Choose whether `self_hosted` remains visibly unsupported until authenticated +3. History is optional and external, not a PostgreSQL sample table. +4. The first Web release has no lifecycle controls. +5. `self_hosted` remains visibly unsupported until authenticated, generation-fenced telemetry is qualified. diff --git a/contracts/agents-api/v1/runtime_observations.go b/contracts/agents-api/v1/runtime_observations.go new file mode 100644 index 000000000..d82d8d514 --- /dev/null +++ b/contracts/agents-api/v1/runtime_observations.go @@ -0,0 +1,46 @@ +package v1 + +type RuntimeObservation struct { + ID string `json:"id" binding:"required" format:"uuid"` + Object string `json:"object" enums:"agent.runtime_observation" binding:"required"` + SessionID string `json:"session_id" binding:"required" format:"uuid"` + EnvironmentID *string `json:"environment_id" extensions:"x-nullable" binding:"required" format:"uuid"` + Mode string `json:"mode" enums:"none,self_hosted,openai_hosted" binding:"required"` + ProviderType *string `json:"provider_type" extensions:"x-nullable" binding:"required" pattern:"^[a-z][a-z0-9_]{0,31}$"` + Instance RuntimeInstance `json:"instance" binding:"required"` + Status string `json:"status" enums:"observed,unsupported,unavailable" binding:"required"` + Reason *string `json:"reason" extensions:"x-nullable" binding:"required" enums:"runtime_mode_not_observable,allocation_pending,runtime_not_running,source_not_configured,sample_timeout,sample_unavailable"` + AllocationCreatedAt *int64 `json:"allocation_created_at" extensions:"x-nullable" binding:"required" minimum:"0"` + ResolvedAt int64 `json:"resolved_at" binding:"required" minimum:"0"` + ObservedAt *int64 `json:"observed_at" extensions:"x-nullable" binding:"required" minimum:"0"` + StartedAt *int64 `json:"started_at" extensions:"x-nullable" binding:"required" minimum:"0"` + CPU *RuntimeCPUObservation `json:"cpu" extensions:"x-nullable" binding:"required"` + Memory *RuntimeMemoryObservation `json:"memory" extensions:"x-nullable" binding:"required"` +} + +type RuntimeInstance struct { + Kind string `json:"kind" enums:"managed_allocation,self_hosted_connection,none" binding:"required"` + AllocationID *string `json:"allocation_id" extensions:"x-nullable" binding:"required" format:"uuid"` + DeviceID *string `json:"device_id" extensions:"x-nullable" binding:"required" format:"uuid"` + ConnectionGeneration *string `json:"connection_generation" extensions:"x-nullable" binding:"required" format:"uuid"` +} + +type RuntimeCPUObservation struct { + UsageSecondsTotal *float64 `json:"usage_seconds_total" extensions:"x-nullable" binding:"required" minimum:"0"` + CapacityCores *float64 `json:"capacity_cores" extensions:"x-nullable" binding:"required" minimum:"5e-324"` + UsageCores *float64 `json:"usage_cores" extensions:"x-nullable" binding:"required" minimum:"0"` + UtilizationRatio *float64 `json:"utilization_ratio" extensions:"x-nullable" binding:"required" minimum:"0"` +} + +type RuntimeMemoryObservation struct { + UsageBytes *uint64 `json:"usage_bytes" extensions:"x-nullable" binding:"required" minimum:"0"` + LimitBytes *uint64 `json:"limit_bytes" extensions:"x-nullable" binding:"required" minimum:"1"` +} + +type RuntimeObservationList struct { + Object string `json:"object" enums:"list" binding:"required"` + Data []RuntimeObservation `json:"data" binding:"required"` + HasMore bool `json:"has_more" binding:"required"` + FirstID *string `json:"first_id" extensions:"x-nullable" binding:"required" format:"uuid"` + LastID *string `json:"last_id" extensions:"x-nullable" binding:"required" format:"uuid"` +} diff --git a/packages/agents-client/src/client.test.ts b/packages/agents-client/src/client.test.ts index 13c0a237d..facd5e497 100644 --- a/packages/agents-client/src/client.test.ts +++ b/packages/agents-client/src/client.test.ts @@ -103,6 +103,42 @@ function messageItem(overrides: Record = {}): Record = {}): Record { + return { + id: runtimeSessionId, + object: "agent.runtime_observation", + session_id: runtimeSessionId, + environment_id: runtimeEnvironmentId, + mode: "openai_hosted", + provider_type: "docker", + instance: { + kind: "managed_allocation", + allocation_id: runtimeAllocationId, + device_id: runtimeDeviceId, + connection_generation: null, + }, + status: "observed", + reason: null, + allocation_created_at: 10, + resolved_at: 30, + observed_at: 20, + started_at: 10, + cpu: { + usage_seconds_total: 0, + capacity_cores: 2, + usage_cores: null, + utilization_ratio: null, + }, + memory: { usage_bytes: 0, limit_bytes: 1024 }, + ...overrides, + }; +} + describe("OpenAIAgentsClient", () => { afterEach(() => vi.unstubAllGlobals()); @@ -2385,4 +2421,121 @@ describe("OpenAIAgentsClient", () => { ).rejects.toThrow("createSession only supports the JSON response"); expect(calls).toHaveLength(0); }); + + it("retrieves a Runtime observation, preserves observed zeroes, and encodes the Session ID", async () => { + const calls: FetchCall[] = []; + const client = new OpenAIAgentsClient({ + baseUrl: "https://core.example/v1", + fetch: recordingFetch(jsonResponse(runtimeObservation()), calls), + }); + + await expect(client.retrieveRuntimeObservation(runtimeSessionId)).resolves.toMatchObject({ + id: runtimeSessionId, + cpu: { usage_seconds_total: 0 }, + memory: { usage_bytes: 0 }, + }); + expect(String(calls[0]?.input)).toBe( + `https://core.example/v1/agents/sessions/${runtimeSessionId}/runtime-observation`, + ); + }); + + it("lists Runtime observations with stable pagination metadata and query serialization", async () => { + const calls: FetchCall[] = []; + const body = { + object: "list", + data: [runtimeObservation()], + has_more: true, + first_id: runtimeSessionId, + last_id: runtimeSessionId, + }; + const client = new OpenAIAgentsClient({ + baseUrl: "https://core.example/v1/", + fetch: recordingFetch(jsonResponse(body), calls), + }); + + await expect(client.listRuntimeObservations({ + after: runtimeSessionId, limit: 1, order: "asc", + })).resolves.toMatchObject(body); + expect(String(calls[0]?.input)).toBe( + `https://core.example/v1/agents/runtime-observations?after=${runtimeSessionId}&limit=1&order=asc`, + ); + }); + + it("accepts an unsupported none-mode Runtime observation with explicit nulls", async () => { + const value = runtimeObservation({ + environment_id: null, + mode: "none", + provider_type: null, + instance: { kind: "none", allocation_id: null, device_id: null, connection_generation: null }, + status: "unsupported", + reason: "runtime_mode_not_observable", + allocation_created_at: null, + observed_at: null, + started_at: null, + cpu: null, + memory: null, + }); + const client = new OpenAIAgentsClient({ fetch: recordingFetch(jsonResponse(value), []) }); + await expect(client.retrieveRuntimeObservation(runtimeSessionId)).resolves.toMatchObject(value); + }); + + it.each([ + ["unknown field", () => ({ ...runtimeObservation(), provider_native_id: "hidden" })], + ["foreign Session", () => ({ ...runtimeObservation(), session_id: "55555555-5555-4555-8555-555555555555" })], + ["invalid status/reason", () => ({ ...runtimeObservation(), status: "observed", reason: "sample_timeout" })], + ["invalid mode/instance", () => ({ ...runtimeObservation(), mode: "none" })], + ["negative CPU", () => ({ ...runtimeObservation(), cpu: { + usage_seconds_total: -1, capacity_cores: 2, usage_cores: null, utilization_ratio: null, + } })], + ["non-numeric CPU", () => ({ ...runtimeObservation(), cpu: { + usage_seconds_total: "NaN", capacity_cores: 2, usage_cores: null, utilization_ratio: null, + } })], + ["zero CPU capacity", () => ({ ...runtimeObservation(), cpu: { + usage_seconds_total: 1, capacity_cores: 0, usage_cores: null, utilization_ratio: null, + } })], + ["unsafe memory", () => ({ ...runtimeObservation(), memory: { + usage_bytes: Number.MAX_SAFE_INTEGER + 1, limit_bytes: 1024, + } })], + ["zero memory limit", () => ({ ...runtimeObservation(), memory: { + usage_bytes: 1, limit_bytes: 0, + } })], + ])("rejects a Runtime observation with %s", async (_label, build) => { + const client = new OpenAIAgentsClient({ fetch: recordingFetch(jsonResponse(build()), []) }); + await expect(client.retrieveRuntimeObservation(runtimeSessionId)).rejects.toMatchObject({ + status: 502, + code: "invalid_runtime_observation", + }); + }); + + it.each([ + ["mismatched first_id", { + object: "list", data: [runtimeObservation()], has_more: false, + first_id: runtimeEnvironmentId, last_id: runtimeSessionId, + }], + ["duplicate IDs", { + object: "list", data: [runtimeObservation(), runtimeObservation()], has_more: false, + first_id: runtimeSessionId, last_id: runtimeSessionId, + }], + ["empty continuation", { + object: "list", data: [], has_more: true, first_id: null, last_id: null, + }], + ])("rejects a Runtime observation list with %s", async (_label, body) => { + const client = new OpenAIAgentsClient({ fetch: recordingFetch(jsonResponse(body), []) }); + await expect(client.listRuntimeObservations()).rejects.toMatchObject({ + status: 502, + code: "invalid_runtime_observation", + }); + }); + + it.each([ + { after: "not-a-uuid" }, + { limit: 0 }, + { limit: 101 }, + { order: "sideways" }, + ])("rejects invalid Runtime observation pagination before fetch", async (options) => { + const calls: FetchCall[] = []; + const client = new OpenAIAgentsClient({ fetch: recordingFetch(jsonResponse({}), calls) }); + await expect(client.listRuntimeObservations(options as never)).rejects.toThrow(TypeError); + expect(calls).toHaveLength(0); + }); }); diff --git a/packages/agents-client/src/client.ts b/packages/agents-client/src/client.ts index c2f32630b..dc4393f57 100644 --- a/packages/agents-client/src/client.ts +++ b/packages/agents-client/src/client.ts @@ -43,6 +43,8 @@ import type { StreamError, UpdateAgentInput, ReplaceVaultCredentialTokenInput, + RuntimeObservation, + RuntimeObservationList, Vault, VaultCredential, VaultCredentialDeleted, @@ -234,6 +236,18 @@ const knownItemTypes = new Set([ ]); const itemStatuses = new Set(["in_progress", "completed", "failed", "incomplete"]); const turnStatuses = new Set(["queued", "in_progress", "waiting", "completed", "failed", "cancelled"]); +const runtimeObservationFields = new Set([ + "id", "object", "session_id", "environment_id", "mode", "provider_type", "instance", "status", "reason", + "allocation_created_at", "resolved_at", "observed_at", "started_at", "cpu", "memory", +]); +const runtimeInstanceFields = new Set(["kind", "allocation_id", "device_id", "connection_generation"]); +const runtimeCPUFields = new Set(["usage_seconds_total", "capacity_cores", "usage_cores", "utilization_ratio"]); +const runtimeMemoryFields = new Set(["usage_bytes", "limit_bytes"]); +const runtimeObservationReasons = new Set([ + "runtime_mode_not_observable", "allocation_pending", "runtime_not_running", + "source_not_configured", "sample_timeout", "sample_unavailable", +]); +const runtimeProviderTypePattern = /^[a-z][a-z0-9_]{0,31}$/; function exactFields(value: Record, fields: Set): boolean { const keys = Object.keys(value); @@ -995,6 +1009,164 @@ function projectAgentSession( return session; } +function invalidRuntimeObservation(message = "Agent Core returned an invalid Runtime observation."): never { + throw new AgentCoreError(message, 502, "invalid_runtime_observation"); +} + +function nullableRuntimeNumber(value: unknown): number | null { + if (value === null) return null; + if (typeof value !== "number" || !Number.isFinite(value) || value < 0) { + return invalidRuntimeObservation(); + } + return value; +} + +function nullableRuntimeInteger(value: unknown): number | null { + const projected = nullableRuntimeNumber(value); + if (projected !== null && !Number.isSafeInteger(projected)) return invalidRuntimeObservation(); + return projected; +} + +function projectRuntimeObservation(value: unknown, expectedSessionId?: string): RuntimeObservation { + if (!isRecord(value) || !exactFields(value, runtimeObservationFields)) { + return invalidRuntimeObservation(); + } + const id = canonicalUuid(value.id); + const sessionId = canonicalUuid(value.session_id); + const environmentId = value.environment_id === null ? null : canonicalUuid(value.environment_id); + if ( + id === null || sessionId === null || id !== sessionId || + (expectedSessionId !== undefined && !sameUuid(sessionId, expectedSessionId)) || + value.object !== "agent.runtime_observation" || + (value.mode !== "none" && value.mode !== "self_hosted" && value.mode !== "openai_hosted") || + !(value.provider_type === null || ( + typeof value.provider_type === "string" && runtimeProviderTypePattern.test(value.provider_type) + )) || + !isRecord(value.instance) || !exactFields(value.instance, runtimeInstanceFields) || + (value.status !== "observed" && value.status !== "unsupported" && value.status !== "unavailable") || + !(value.reason === null || ( + typeof value.reason === "string" && runtimeObservationReasons.has(value.reason) + )) || + !isNonnegativeInteger(value.resolved_at) + ) return invalidRuntimeObservation(); + + const allocationId = value.instance.allocation_id === null ? null : canonicalUuid(value.instance.allocation_id); + const deviceId = value.instance.device_id === null ? null : canonicalUuid(value.instance.device_id); + const connectionGeneration = value.instance.connection_generation === null + ? null + : canonicalUuid(value.instance.connection_generation); + if ( + (value.instance.allocation_id !== null && allocationId === null) || + (value.instance.device_id !== null && deviceId === null) || + (value.instance.connection_generation !== null && connectionGeneration === null) + ) return invalidRuntimeObservation(); + + const allocationCreatedAt = nullableRuntimeInteger(value.allocation_created_at); + const observedAt = nullableRuntimeInteger(value.observed_at); + const startedAt = nullableRuntimeInteger(value.started_at); + const isNone = value.mode === "none"; + const isSelfHosted = value.mode === "self_hosted"; + const isManaged = value.mode === "openai_hosted"; + if ( + (isNone && ( + value.instance.kind !== "none" || environmentId !== null || value.provider_type !== null || + allocationId !== null || deviceId !== null || connectionGeneration !== null || allocationCreatedAt !== null + )) || + (isSelfHosted && ( + value.instance.kind !== "self_hosted_connection" || environmentId === null || + allocationId !== null || allocationCreatedAt !== null + )) || + (isManaged && ( + value.instance.kind !== "managed_allocation" || environmentId === null || connectionGeneration !== null || + (allocationId === null && (deviceId !== null || allocationCreatedAt !== null)) + )) + ) return invalidRuntimeObservation(); + + const observed = value.status === "observed"; + if ( + (observed && ( + !isManaged || allocationId === null || value.reason !== null || observedAt === null || + observedAt > value.resolved_at + )) || + (!observed && ( + observedAt !== null || startedAt !== null || value.cpu !== null || value.memory !== null + )) || + (value.status === "unsupported" && ( + (!isNone && !isSelfHosted) || value.reason !== "runtime_mode_not_observable" + )) || + (value.status === "unavailable" && ( + !isManaged || value.reason === null || value.reason === "runtime_mode_not_observable" + )) || + (startedAt !== null && observedAt !== null && startedAt > observedAt) || + (allocationCreatedAt !== null && allocationCreatedAt > value.resolved_at) + ) return invalidRuntimeObservation(); + + let cpu: RuntimeObservation["cpu"] = null; + if (value.cpu !== null) { + if (!observed || !isRecord(value.cpu) || !exactFields(value.cpu, runtimeCPUFields)) { + return invalidRuntimeObservation(); + } + cpu = { + usage_seconds_total: nullableRuntimeNumber(value.cpu.usage_seconds_total), + capacity_cores: nullableRuntimeNumber(value.cpu.capacity_cores), + usage_cores: nullableRuntimeNumber(value.cpu.usage_cores), + utilization_ratio: nullableRuntimeNumber(value.cpu.utilization_ratio), + }; + if ( + Object.values(cpu).every((entry) => entry === null) || + (cpu.capacity_cores !== null && cpu.capacity_cores === 0) + ) return invalidRuntimeObservation(); + } + + let memory: RuntimeObservation["memory"] = null; + if (value.memory !== null) { + if (!observed || !isRecord(value.memory) || !exactFields(value.memory, runtimeMemoryFields)) { + return invalidRuntimeObservation(); + } + memory = { + usage_bytes: nullableRuntimeInteger(value.memory.usage_bytes), + limit_bytes: nullableRuntimeInteger(value.memory.limit_bytes), + }; + if ( + (memory.usage_bytes === null && memory.limit_bytes === null) || + memory.limit_bytes === 0 + ) return invalidRuntimeObservation(); + } + + return { + id, object: "agent.runtime_observation", session_id: sessionId, environment_id: environmentId, + mode: value.mode, provider_type: value.provider_type, instance: { + kind: value.instance.kind as RuntimeObservation["instance"]["kind"], + allocation_id: allocationId, device_id: deviceId, connection_generation: connectionGeneration, + }, + status: value.status, reason: value.reason as RuntimeObservation["reason"], + allocation_created_at: allocationCreatedAt, resolved_at: value.resolved_at, + observed_at: observedAt, started_at: startedAt, cpu, memory, + } as RuntimeObservation; +} + +function projectRuntimeObservationList(value: unknown, options?: PageOptions): RuntimeObservationList { + if ( + !isRecord(value) || !exactFields(value, vaultListFields) || value.object !== "list" || + !Array.isArray(value.data) || typeof value.has_more !== "boolean" + ) return invalidRuntimeObservation("Agent Core returned an invalid Runtime observation list."); + const limit = options?.limit ?? 20; + if ( + !Number.isSafeInteger(limit) || limit < 1 || limit > 100 || + (options?.order !== undefined && options.order !== "asc" && options.order !== "desc") || + value.data.length > limit + ) return invalidRuntimeObservation("Agent Core returned an invalid Runtime observation list."); + const data = value.data.map((entry) => projectRuntimeObservation(entry)); + const firstId = data[0]?.id ?? null; + const lastId = data[data.length - 1]?.id ?? null; + if ( + new Set(data.map((entry) => entry.id)).size !== data.length || + value.first_id !== firstId || value.last_id !== lastId || + (value.has_more && data.length === 0) + ) return invalidRuntimeObservation("Agent Core returned an invalid Runtime observation list."); + return { object: "list", data, has_more: value.has_more, first_id: firstId, last_id: lastId }; +} + function projectStreamError(value: unknown): StreamError { if ( !isRecord(value) || !exactFields(value, streamErrorFields) || @@ -2016,6 +2188,31 @@ export class OpenAIAgentsClient implements AgentCore { return { ...page, data: page.data.map((session) => projectAgentSession(session)) }; } + async listRuntimeObservations(options?: PageOptions): Promise { + if ( + (options?.after !== undefined && canonicalUuid(options.after) === null) || + (options?.limit !== undefined && ( + !Number.isSafeInteger(options.limit) || options.limit < 1 || options.limit > 100 + )) || + (options?.order !== undefined && options.order !== "asc" && options.order !== "desc") + ) throw new TypeError("Runtime observation pagination options are invalid."); + const params = new URLSearchParams(); + addPageOptions(params, options); + const value = await this.request( + withQuery("/agents/runtime-observations", params), + { signal: options?.signal }, + ); + return projectRuntimeObservationList(value, options); + } + + async retrieveRuntimeObservation(sessionId: string, options?: ReadOptions): Promise { + const value = await this.request( + `/agents/sessions/${encodeURIComponent(sessionId)}/runtime-observation`, + { signal: options?.signal }, + ); + return projectRuntimeObservation(value, sessionId); + } + async createSession(input: CreateSessionInput, idempotencyKey = createIdempotencyKey()): Promise { if ((input as { stream?: boolean }).stream === true) { throw new TypeError("createSession only supports the JSON response; connect streamEvents after creation."); diff --git a/packages/agents-client/src/protocol-types.test.ts b/packages/agents-client/src/protocol-types.test.ts index 3c1384129..ddf80f36e 100644 --- a/packages/agents-client/src/protocol-types.test.ts +++ b/packages/agents-client/src/protocol-types.test.ts @@ -29,6 +29,8 @@ import type { OpenAIHostedAgentEnvironmentInput, OpenAIHostedAgentEnvironmentResource, RequiredAction, + RuntimeObservation, + RuntimeUnavailableReason, SelfHostedAgentEnvironment, SavedAgentToolInput, SourceFile, @@ -48,6 +50,25 @@ import type { VaultCredential, } from "./types"; +describe("Runtime Observation discriminated contract", () => { + it("narrows status, reason, mode, instance, and sample presence together", () => { + type Observed = Extract; + type Unavailable = Extract; + type NoneMode = Extract; + type SelfHosted = Extract; + + expectTypeOf().toEqualTypeOf<"openai_hosted">(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf(); + expectTypeOf().toEqualTypeOf<"none">(); + expectTypeOf().toEqualTypeOf<"self_hosted_connection">(); + }); +}); + describe("Parsar dadf64a7 basic managed Environment profile", () => { it("pins omitted/default, explicit-enabled, and explicit-disabled network input", () => { const inputs = hostedDadf64.inputs as Record; diff --git a/packages/agents-client/src/types.ts b/packages/agents-client/src/types.ts index 74a156353..8bc3b0f6b 100644 --- a/packages/agents-client/src/types.ts +++ b/packages/agents-client/src/types.ts @@ -682,6 +682,119 @@ export interface CreateSessionStreamOptions extends StreamOptions { onSession: (session: AgentSession) => void; } +export type RuntimeObservationStatus = "observed" | "unsupported" | "unavailable"; +export type RuntimeObservationReason = + | "runtime_mode_not_observable" + | "allocation_pending" + | "runtime_not_running" + | "source_not_configured" + | "sample_timeout" + | "sample_unavailable"; + +export type RuntimeUnavailableReason = Exclude; + +export interface RuntimeCPUObservation { + usage_seconds_total: number | null; + capacity_cores: number | null; + usage_cores: number | null; + utilization_ratio: number | null; +} + +export interface RuntimeMemoryObservation { + usage_bytes: number | null; + limit_bytes: number | null; +} + +interface RuntimeObservationBase { + id: string; + object: "agent.runtime_observation"; + session_id: string; + resolved_at: number; +} + +export interface RuntimeObservedObservation extends RuntimeObservationBase { + environment_id: string; + mode: "openai_hosted"; + provider_type: string | null; + instance: { + kind: "managed_allocation"; + allocation_id: string; + device_id: string | null; + connection_generation: null; + }; + status: "observed"; + reason: null; + allocation_created_at: number | null; + observed_at: number; + started_at: number | null; + cpu: RuntimeCPUObservation | null; + memory: RuntimeMemoryObservation | null; +} + +export interface RuntimeUnavailableObservation extends RuntimeObservationBase { + environment_id: string; + mode: "openai_hosted"; + provider_type: string | null; + instance: { + kind: "managed_allocation"; + allocation_id: string | null; + device_id: string | null; + connection_generation: null; + }; + status: "unavailable"; + reason: RuntimeUnavailableReason; + allocation_created_at: number | null; + observed_at: null; + started_at: null; + cpu: null; + memory: null; +} + +export interface RuntimeNoneObservation extends RuntimeObservationBase { + environment_id: null; + mode: "none"; + provider_type: null; + instance: { kind: "none"; allocation_id: null; device_id: null; connection_generation: null }; + status: "unsupported"; + reason: "runtime_mode_not_observable"; + allocation_created_at: null; + observed_at: null; + started_at: null; + cpu: null; + memory: null; +} + +export interface RuntimeSelfHostedObservation extends RuntimeObservationBase { + environment_id: string; + mode: "self_hosted"; + provider_type: string | null; + instance: { + kind: "self_hosted_connection"; + allocation_id: null; + device_id: string | null; + connection_generation: string | null; + }; + status: "unsupported"; + reason: "runtime_mode_not_observable"; + allocation_created_at: null; + observed_at: null; + started_at: null; + cpu: null; + memory: null; +} + +export type RuntimeObservation = + | RuntimeObservedObservation + | RuntimeUnavailableObservation + | RuntimeNoneObservation + | RuntimeSelfHostedObservation; + +export interface RuntimeObservationList extends ListPage { + object: "list"; + first_id: string | null; + last_id: string | null; +} + export interface AgentCore { listAgents(options?: PageOptions): Promise>; createAgent(input: CreateAgentInput): Promise; @@ -698,6 +811,8 @@ export interface AgentCore { replaceVaultCredentialToken(vaultId: string, credentialId: string, input: ReplaceVaultCredentialTokenInput): Promise; deleteVaultCredential(vaultId: string, credentialId: string): Promise; listSessions(options?: PageOptions & { agentId?: string }): Promise>; + listRuntimeObservations(options?: PageOptions): Promise; + retrieveRuntimeObservation(sessionId: string, options?: ReadOptions): Promise; createSession(input: CreateSessionInput, idempotencyKey?: string): Promise; createSessionStream( input: Omit, diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 757c2be80..d7c2daf88 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -28,6 +28,7 @@ import ( "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/execution" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtime" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeenrollment" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" "github.com/jackc/pgx/v5/pgxpool" ) @@ -89,9 +90,27 @@ func run() error { if err := executionStore.EnsureProjectScopes(ready, auth.ProjectScopes()); err != nil { return err } + observationSources := map[string]runtimeobs.Source{} + if managed != nil { + for key, provider := range managed.Providers { + source, ok := provider.(runtimeobs.Source) + if !ok { + continue + } + observationSources[key] = source + } + } + resolver, err := runtimeobs.NewResolver(executionStore) + if err != nil { + return err + } + observationService, err := runtimeobs.NewService(resolver, observationSources) + if err != nil { + return err + } var workerDone chan error var worker *execution.Worker - options := []api.Option{api.WithSubagents(executionStore), api.WithSkills(executionStore), api.WithSourceFiles(executionStore), api.WithSessionArtifacts(executionStore)} + options := []api.Option{api.WithSubagents(executionStore), api.WithSkills(executionStore), api.WithSourceFiles(executionStore), api.WithSessionArtifacts(executionStore), api.WithRuntimeObservations(observationService)} var daemonHandler http.Handler var registry *gateway.Registry if wsURL := os.Getenv("AGENTS_API_DAEMON_WS_URL"); wsURL != "" { diff --git a/services/agents-api/internal/api/handler.go b/services/agents-api/internal/api/handler.go index 4fa78bdbb..a154843ef 100644 --- a/services/agents-api/internal/api/handler.go +++ b/services/agents-api/internal/api/handler.go @@ -35,20 +35,21 @@ type ResourceStore interface { } type Handler struct { - policy execution.Policy - store ResourceStore - auth *Authenticator - harnesses map[string]bool - engine string - inputs InputSubmitter - executorURL string - hostedEnvironments bool - directoryReader EnvironmentDirectoryReader - fileWriter EnvironmentFileWriter - skills SkillStore - sourceFiles SourceFileStore - artifacts SessionArtifactStore - subagents SubagentStore + policy execution.Policy + store ResourceStore + auth *Authenticator + harnesses map[string]bool + engine string + inputs InputSubmitter + executorURL string + hostedEnvironments bool + directoryReader EnvironmentDirectoryReader + fileWriter EnvironmentFileWriter + skills SkillStore + sourceFiles SourceFileStore + artifacts SessionArtifactStore + subagents SubagentStore + runtimeObservations RuntimeObservationService } func NewHandler(s ResourceStore, auth *Authenticator, engine string, options ...Option) (http.Handler, error) { @@ -100,6 +101,8 @@ func NewHandler(s ResourceStore, auth *Authenticator, engine string, options ... r.Post("/agents/sessions", h.createSession) r.Get("/agents/sessions", h.listSessions) r.Get("/agents/sessions/{session_id}", h.getSession) + r.Get("/agents/sessions/{session_id}/runtime-observation", h.getRuntimeObservation) + r.Get("/agents/runtime-observations", h.listRuntimeObservations) r.Post("/agents/sessions/{session_id}", h.updateSession) r.Delete("/agents/sessions/{session_id}", h.deleteSession) r.Post("/agents/sessions/{session_id}/events", h.createEvents) diff --git a/services/agents-api/internal/api/handler_test.go b/services/agents-api/internal/api/handler_test.go index dbf08bcf7..b1c3befe2 100644 --- a/services/agents-api/internal/api/handler_test.go +++ b/services/agents-api/internal/api/handler_test.go @@ -19,8 +19,19 @@ import ( type recordingStore struct { ResourceStore - tenant string - input store.CreateSessionInput + tenant string + input store.CreateSessionInput + sessions []store.Session + nextSessionCursor string + listTenant string + listAfter string + listLimit int + listAscending bool +} + +func (s *recordingStore) ListSessions(_ context.Context, tenant, after string, limit int, ascending bool, _ *string) (store.SessionPage, error) { + s.listTenant, s.listAfter, s.listLimit, s.listAscending = tenant, after, limit, ascending + return store.SessionPage{Sessions: append([]store.Session(nil), s.sessions...), NextCursor: s.nextSessionCursor}, nil } func (s *recordingStore) GetSession(ctx context.Context, tenant, id string) (store.Session, error) { diff --git a/services/agents-api/internal/api/runtime_observations.go b/services/agents-api/internal/api/runtime_observations.go new file mode 100644 index 000000000..54d061a2d --- /dev/null +++ b/services/agents-api/internal/api/runtime_observations.go @@ -0,0 +1,221 @@ +package api + +import ( + "context" + "errors" + "net/http" + "regexp" + "sync" + "time" + + v1 "github.com/MiniMax-AI-Dev/parsar/contracts/agents-api/v1" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/go-chi/chi/v5" +) + +const ( + runtimeObservationConcurrency = 8 + runtimeObservationSourceBudget = 2 * time.Second + runtimeObservationRequestBudget = 10 * time.Second +) + +var runtimeProviderTypePattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) + +type RuntimeObservationService interface { + ObserveSession(context.Context, string, string) (runtimeobs.Observation, error) +} + +func WithRuntimeObservations(service RuntimeObservationService) Option { + return func(h *Handler) { h.runtimeObservations = service } +} + +// @Summary Retrieve a Session Runtime observation +// @Description Core extension returning one tenant-scoped, read-only current Runtime observation. It never provisions, renews, restarts, pauses or stops compute. +// @Tags Runtime observations +// @Produce json +// @Security BearerAuth +// @Param OpenAI-Beta header string true "agents=v1" +// @Param session_id path string true "Session ID" +// @Success 200 {object} v1.RuntimeObservation +// @Failure 400,401,404,500,503 {object} v1.ErrorResponse +// @Router /agents/sessions/{session_id}/runtime-observation [get] +func (h *Handler) getRuntimeObservation(w http.ResponseWriter, r *http.Request) { + if len(r.URL.Query()) != 0 { + writeError(w, http.StatusBadRequest, "unsupported_parameter", "Runtime observation retrieval does not accept query parameters.") + return + } + if h.runtimeObservations == nil { + writeError(w, http.StatusServiceUnavailable, "execution_unavailable", "Runtime observation is not configured on this service.") + return + } + ctx, cancel := context.WithTimeout(r.Context(), runtimeObservationSourceBudget) + defer cancel() + observation, err := h.runtimeObservations.ObserveSession(ctx, tenantID(r), chi.URLParam(r, "session_id")) + if err != nil { + writeStoreError(w, r, err) + return + } + response, err := runtimeObservationResponse(observation) + if err != nil { + writeStoreError(w, r, err) + return + } + writeJSON(w, http.StatusOK, response) +} + +// @Summary List current Runtime observations +// @Description Core extension listing one current Runtime context per tenant-owned Session in Session creation order. Each row has an independent resolved_at and optional provider observed_at; the page is not an atomic telemetry snapshot. +// @Tags Runtime observations +// @Produce json +// @Security BearerAuth +// @Param OpenAI-Beta header string true "agents=v1" +// @Param after query string false "Last observation ID from the previous page" +// @Param limit query int false "Page size" minimum(1) maximum(100) default(20) +// @Param order query string false "Session creation order" Enums(asc,desc) default(desc) +// @Success 200 {object} v1.RuntimeObservationList +// @Failure 400,401,404,500,503 {object} v1.ErrorResponse +// @Router /agents/runtime-observations [get] +func (h *Handler) listRuntimeObservations(w http.ResponseWriter, r *http.Request) { + if h.runtimeObservations == nil { + writeError(w, http.StatusServiceUnavailable, "execution_unavailable", "Runtime observation is not configured on this service.") + return + } + options, ok := readPage(w, r) + if !ok { + return + } + ctx, cancel := context.WithTimeout(r.Context(), runtimeObservationRequestBudget) + defer cancel() + page, err := h.store.ListSessions(ctx, tenantID(r), options.after, options.limit, options.ascending, nil) + if err != nil { + writeStoreError(w, r, err) + return + } + observations := make([]runtimeobs.Observation, len(page.Sessions)) + semaphore := make(chan struct{}, runtimeObservationConcurrency) + work, stop := context.WithCancel(ctx) + defer stop() + var wait sync.WaitGroup + var once sync.Once + var firstErr error + for index, session := range page.Sessions { + wait.Add(1) + go func(index int, sessionID string) { + defer wait.Done() + select { + case semaphore <- struct{}{}: + defer func() { <-semaphore }() + case <-work.Done(): + return + } + sampleCtx, sampleCancel := context.WithTimeout(work, runtimeObservationSourceBudget) + defer sampleCancel() + value, err := h.runtimeObservations.ObserveSession(sampleCtx, tenantID(r), sessionID) + if err != nil { + once.Do(func() { firstErr = err; stop() }) + return + } + observations[index] = value + }(index, session.ID) + } + wait.Wait() + if firstErr != nil { + writeStoreError(w, r, firstErr) + return + } + if err := ctx.Err(); err != nil { + writeError(w, http.StatusServiceUnavailable, "execution_unavailable", "Runtime observation collection exceeded its request budget.") + return + } + response := v1.RuntimeObservationList{Object: "list", Data: make([]v1.RuntimeObservation, 0, len(observations)), HasMore: page.NextCursor != ""} + for _, observation := range observations { + item, err := runtimeObservationResponse(observation) + if err != nil { + writeStoreError(w, r, err) + return + } + response.Data = append(response.Data, item) + } + if len(response.Data) > 0 { + response.FirstID = &response.Data[0].ID + response.LastID = &response.Data[len(response.Data)-1].ID + } + writeJSON(w, http.StatusOK, response) +} + +func runtimeObservationResponse(observation runtimeobs.Observation) (v1.RuntimeObservation, error) { + if observation.Target.SessionID == "" || observation.ResolvedAt.IsZero() || observation.ResolvedAt.Unix() < 0 { + return v1.RuntimeObservation{}, errors.New("invalid Runtime observation identity") + } + if !observation.Target.Instance.AllocationCreatedAt.IsZero() && + (observation.Target.Instance.AllocationCreatedAt.Unix() < 0 || observation.Target.Instance.AllocationCreatedAt.After(observation.ResolvedAt)) { + return v1.RuntimeObservation{}, errors.New("invalid Runtime allocation creation time") + } + if observation.Sample != nil { + if observation.Sample.ObservedAt.IsZero() || observation.Sample.ObservedAt.Unix() < 0 || observation.Sample.ObservedAt.After(observation.ResolvedAt) { + return v1.RuntimeObservation{}, errors.New("invalid Runtime sample time") + } + if observation.Sample.StartedAt != nil && + (observation.Sample.StartedAt.IsZero() || observation.Sample.StartedAt.Unix() < 0 || observation.Sample.StartedAt.After(observation.Sample.ObservedAt)) { + return v1.RuntimeObservation{}, errors.New("invalid Runtime start time") + } + } + result := v1.RuntimeObservation{ + ID: observation.Target.SessionID, Object: "agent.runtime_observation", SessionID: observation.Target.SessionID, + Mode: string(observation.Target.Mode), Status: string(observation.Status), ResolvedAt: observation.ResolvedAt.Unix(), + } + if observation.Target.EnvironmentID != "" { + result.EnvironmentID = &observation.Target.EnvironmentID + } + if observation.ProviderType != "" { + if !runtimeProviderTypePattern.MatchString(observation.ProviderType) { + return v1.RuntimeObservation{}, errors.New("invalid Runtime observation provider type") + } + result.ProviderType = &observation.ProviderType + } + if observation.Reason != "" { + result.Reason = &observation.Reason + } + switch observation.Target.Mode { + case runtimeobs.ModeManaged: + result.Instance.Kind = "managed_allocation" + if observation.Target.Instance.AllocationID != "" { + result.Instance.AllocationID = &observation.Target.Instance.AllocationID + } + if observation.Target.Instance.DeviceID != "" { + result.Instance.DeviceID = &observation.Target.Instance.DeviceID + } + if !observation.Target.Instance.AllocationCreatedAt.IsZero() { + created := observation.Target.Instance.AllocationCreatedAt.Unix() + result.AllocationCreatedAt = &created + } + case runtimeobs.ModeSelfHosted: + result.Instance.Kind = "self_hosted_connection" + if observation.Target.Instance.DeviceID != "" { + result.Instance.DeviceID = &observation.Target.Instance.DeviceID + } + if observation.Target.Instance.ConnectionGeneration != "" { + result.Instance.ConnectionGeneration = &observation.Target.Instance.ConnectionGeneration + } + case runtimeobs.ModeNone: + result.Instance.Kind = "none" + default: + return v1.RuntimeObservation{}, errors.New("invalid Runtime observation mode") + } + if observation.Sample == nil { + return result, nil + } + observedAt := observation.Sample.ObservedAt.Unix() + result.ObservedAt = &observedAt + if observation.Sample.StartedAt != nil { + startedAt := observation.Sample.StartedAt.Unix() + result.StartedAt = &startedAt + } + if observation.Sample.CPUUsageSecondsTotal != nil || observation.Sample.CPUCapacityCores != nil { + result.CPU = &v1.RuntimeCPUObservation{UsageSecondsTotal: observation.Sample.CPUUsageSecondsTotal, CapacityCores: observation.Sample.CPUCapacityCores} + } + if observation.Sample.MemoryUsageBytes != nil || observation.Sample.MemoryLimitBytes != nil { + result.Memory = &v1.RuntimeMemoryObservation{UsageBytes: observation.Sample.MemoryUsageBytes, LimitBytes: observation.Sample.MemoryLimitBytes} + } + return result, nil +} diff --git a/services/agents-api/internal/api/runtime_observations_test.go b/services/agents-api/internal/api/runtime_observations_test.go new file mode 100644 index 000000000..29768be94 --- /dev/null +++ b/services/agents-api/internal/api/runtime_observations_test.go @@ -0,0 +1,224 @@ +package api + +import ( + "context" + "encoding/json" + "errors" + "net/http" + "net/http/httptest" + "sync/atomic" + "testing" + "time" + + v1 "github.com/MiniMax-AI-Dev/parsar/contracts/agents-api/v1" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" + "github.com/google/uuid" +) + +type runtimeObservationFixture struct { + values map[string]runtimeobs.Observation +} + +func (f runtimeObservationFixture) ObserveSession(_ context.Context, tenant, session string) (runtimeobs.Observation, error) { + value := f.values[session] + value.Target.TenantID = tenant + return value, nil +} + +type runtimeObservationServiceFunc func(context.Context, string, string) (runtimeobs.Observation, error) + +func (f runtimeObservationServiceFunc) ObserveSession(ctx context.Context, tenant, session string) (runtimeobs.Observation, error) { + return f(ctx, tenant, session) +} + +func runtimeObservationRequest(handler http.Handler, path string) *httptest.ResponseRecorder { + request := httptest.NewRequest(http.MethodGet, path, nil) + request.Header.Set("Authorization", "Bearer test-api-key") + request.Header.Set("OpenAI-Beta", "agents=v1") + response := httptest.NewRecorder() + handler.ServeHTTP(response, request) + return response +} + +func TestRuntimeObservationRoutesUseSessionIdentityAndExactNullability(t *testing.T) { + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + sessionID := uuid.NewString() + service := runtimeObservationFixture{values: map[string]runtimeobs.Observation{ + sessionID: {Target: runtimeobs.Target{SessionID: sessionID, Mode: runtimeobs.ModeNone}, Status: runtimeobs.StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: now}, + }} + handler, saved, _ := testHandler(t, WithRuntimeObservations(service)) + saved.sessions = []store.Session{{ID: sessionID, CreatedAt: now}} + + for _, path := range []string{"/v1/agents/sessions/" + sessionID + "/runtime-observation", "/v1/agents/runtime-observations?limit=1"} { + response := runtimeObservationRequest(handler, path) + if response.Code != http.StatusOK { + t.Fatalf("%s returned %d: %s", path, response.Code, response.Body) + } + if path[len(path)-7:] == "limit=1" { + var page v1.RuntimeObservationList + if json.Unmarshal(response.Body.Bytes(), &page) != nil || len(page.Data) != 1 || page.FirstID == nil || *page.FirstID != sessionID || page.LastID == nil || *page.LastID != sessionID { + t.Fatalf("invalid observation page: %s", response.Body) + } + continue + } + var value v1.RuntimeObservation + if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.ID != sessionID || value.SessionID != sessionID || value.Instance.Kind != "none" || value.EnvironmentID != nil || value.ProviderType != nil || value.ObservedAt != nil || value.CPU != nil || value.Memory != nil || value.Reason == nil || *value.Reason != "runtime_mode_not_observable" { + t.Fatalf("invalid unsupported observation: %s", response.Body) + } + } +} + +func TestRuntimeObservationRoutesRequireConfiguredServiceAndRejectQueries(t *testing.T) { + handler, _, _ := testHandler(t) + missing := runtimeObservationRequest(handler, "/v1/agents/runtime-observations") + if missing.Code != http.StatusServiceUnavailable { + t.Fatalf("unconfigured service returned %d: %s", missing.Code, missing.Body) + } + + sessionID := uuid.NewString() + now := time.Now().UTC() + service := runtimeObservationFixture{values: map[string]runtimeobs.Observation{ + sessionID: {Target: runtimeobs.Target{SessionID: sessionID, Mode: runtimeobs.ModeNone}, Status: runtimeobs.StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: now}, + }} + handler, _, _ = testHandler(t, WithRuntimeObservations(service)) + invalid := runtimeObservationRequest(handler, "/v1/agents/sessions/"+sessionID+"/runtime-observation?provider=docker") + if invalid.Code != http.StatusBadRequest { + t.Fatalf("unsupported query returned %d: %s", invalid.Code, invalid.Body) + } +} + +func TestRuntimeObservationResponsePreservesObservedZero(t *testing.T) { + zeroCPU := float64(0) + zeroMemory := uint64(0) + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + sessionID, environmentID := uuid.NewString(), uuid.NewString() + value, err := runtimeObservationResponse(runtimeobs.Observation{ + Target: runtimeobs.Target{SessionID: sessionID, EnvironmentID: environmentID, Mode: runtimeobs.ModeManaged, Instance: runtimeobs.Instance{AllocationID: uuid.NewString(), DeviceID: uuid.NewString(), AllocationCreatedAt: now.Add(-time.Hour)}}, + Status: runtimeobs.StatusObserved, ProviderType: "docker", ResolvedAt: now, + Sample: &runtimeobs.Sample{ObservedAt: now, CPUUsageSecondsTotal: &zeroCPU, MemoryUsageBytes: &zeroMemory}, + }) + if err != nil || value.CPU == nil || value.CPU.UsageSecondsTotal == nil || *value.CPU.UsageSecondsTotal != 0 || value.Memory == nil || value.Memory.UsageBytes == nil || *value.Memory.UsageBytes != 0 { + t.Fatalf("observed zero was lost: %+v %v", value, err) + } +} + +func TestRuntimeObservationResponseRejectsTimesOutsidePublicContract(t *testing.T) { + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + preEpoch := time.Unix(-1, 0).UTC() + base := runtimeobs.Observation{ + Target: runtimeobs.Target{ + SessionID: uuid.NewString(), EnvironmentID: uuid.NewString(), Mode: runtimeobs.ModeManaged, + Instance: runtimeobs.Instance{AllocationID: uuid.NewString(), DeviceID: uuid.NewString(), AllocationCreatedAt: now.Add(-time.Hour)}, + }, + Status: runtimeobs.StatusObserved, ResolvedAt: now, + Sample: &runtimeobs.Sample{ObservedAt: now, StartedAt: timePointer(now.Add(-time.Minute))}, + } + for _, mutate := range []func(*runtimeobs.Observation){ + func(value *runtimeobs.Observation) { value.Target.Instance.AllocationCreatedAt = now.Add(time.Second) }, + func(value *runtimeobs.Observation) { value.Target.Instance.AllocationCreatedAt = preEpoch }, + func(value *runtimeobs.Observation) { value.Sample.ObservedAt = preEpoch }, + func(value *runtimeobs.Observation) { value.Sample.StartedAt = &preEpoch }, + } { + observation := base + sample := *base.Sample + observation.Sample = &sample + mutate(&observation) + if _, err := runtimeObservationResponse(observation); err == nil { + t.Fatalf("invalid Runtime time accepted: %+v", observation) + } + } +} + +func timePointer(value time.Time) *time.Time { return &value } + +func TestRuntimeObservationListPreservesStoreOrderAndTenantPagination(t *testing.T) { + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + sessionIDs := []string{uuid.NewString(), uuid.NewString(), uuid.NewString()} + var expectedTenant string + service := runtimeObservationServiceFunc(func(_ context.Context, tenant, session string) (runtimeobs.Observation, error) { + if tenant != expectedTenant { + return runtimeobs.Observation{}, errors.New("unexpected tenant") + } + if session == sessionIDs[0] { + time.Sleep(20 * time.Millisecond) + } + return runtimeobs.Observation{ + Target: runtimeobs.Target{SessionID: session, Mode: runtimeobs.ModeNone}, + Status: runtimeobs.StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: now, + }, nil + }) + handler, saved, tenant := testHandler(t, WithRuntimeObservations(service)) + expectedTenant = tenant + for _, id := range sessionIDs { + saved.sessions = append(saved.sessions, store.Session{ID: id, CreatedAt: now}) + } + saved.nextSessionCursor = "next" + + response := runtimeObservationRequest(handler, "/v1/agents/runtime-observations?after=cursor&limit=3&order=asc") + if response.Code != http.StatusOK { + t.Fatalf("list returned %d: %s", response.Code, response.Body) + } + var page v1.RuntimeObservationList + if err := json.Unmarshal(response.Body.Bytes(), &page); err != nil { + t.Fatal(err) + } + if len(page.Data) != len(sessionIDs) || !page.HasMore || saved.listTenant != tenant || saved.listAfter != "cursor" || saved.listLimit != 3 || !saved.listAscending { + t.Fatalf("pagination binding was not preserved: page=%+v store=%+v", page, saved) + } + for index, item := range page.Data { + if item.ID != sessionIDs[index] { + t.Fatalf("concurrent collection reordered page: %+v", page.Data) + } + } +} + +func TestRuntimeObservationListBoundsCollectionConcurrency(t *testing.T) { + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + var active, maximum atomic.Int32 + service := runtimeObservationServiceFunc(func(_ context.Context, _, session string) (runtimeobs.Observation, error) { + current := active.Add(1) + defer active.Add(-1) + for current > maximum.Load() && !maximum.CompareAndSwap(maximum.Load(), current) { + } + time.Sleep(15 * time.Millisecond) + return runtimeobs.Observation{ + Target: runtimeobs.Target{SessionID: session, Mode: runtimeobs.ModeNone}, + Status: runtimeobs.StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: now, + }, nil + }) + handler, saved, _ := testHandler(t, WithRuntimeObservations(service)) + for range 20 { + saved.sessions = append(saved.sessions, store.Session{ID: uuid.NewString(), CreatedAt: now}) + } + response := runtimeObservationRequest(handler, "/v1/agents/runtime-observations?limit=20") + if response.Code != http.StatusOK { + t.Fatalf("list returned %d: %s", response.Code, response.Body) + } + if got := maximum.Load(); got == 0 || got > runtimeObservationConcurrency { + t.Fatalf("collection concurrency = %d, want 1..%d", got, runtimeObservationConcurrency) + } +} + +func TestRuntimeObservationListRejectsWholePageOnIntegrityFailure(t *testing.T) { + now := time.Date(2026, 9, 22, 8, 0, 0, 0, time.UTC) + validID, invalidID := uuid.NewString(), uuid.NewString() + service := runtimeObservationFixture{values: map[string]runtimeobs.Observation{ + validID: { + Target: runtimeobs.Target{SessionID: validID, Mode: runtimeobs.ModeNone}, + Status: runtimeobs.StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: now, + }, + invalidID: {Target: runtimeobs.Target{Mode: runtimeobs.ModeNone}, Status: runtimeobs.StatusUnsupported, ResolvedAt: now}, + }} + handler, saved, _ := testHandler(t, WithRuntimeObservations(service)) + saved.sessions = []store.Session{{ID: validID, CreatedAt: now}, {ID: invalidID, CreatedAt: now}} + + response := runtimeObservationRequest(handler, "/v1/agents/runtime-observations?limit=2") + if response.Code != http.StatusInternalServerError { + t.Fatalf("integrity failure returned %d: %s", response.Code, response.Body) + } + var envelope v1.ErrorResponse + if err := json.Unmarshal(response.Body.Bytes(), &envelope); err != nil || envelope.Error.Code == "" { + t.Fatalf("integrity failure leaked a partial page: %s", response.Body) + } +} diff --git a/services/agents-api/internal/runtimeobs/identity.go b/services/agents-api/internal/runtimeobs/identity.go index 7e33f4bb9..17e38f7d4 100644 --- a/services/agents-api/internal/runtimeobs/identity.go +++ b/services/agents-api/internal/runtimeobs/identity.go @@ -1,5 +1,7 @@ package runtimeobs +import "time" + // Instance is one provider-owned Runtime incarnation. AllocationID is present // for managed compute. DeviceID and ConnectionGeneration are reserved for a // future authenticated self-hosted telemetry source. @@ -8,6 +10,8 @@ type Instance struct { ProviderKey string DeviceID string ConnectionGeneration string + AllocationState string + AllocationCreatedAt time.Time } // Target binds telemetry to durable Core identity. A Session is not itself a diff --git a/services/agents-api/internal/runtimeobs/resolver.go b/services/agents-api/internal/runtimeobs/resolver.go index 7b1047c0c..58a5f28df 100644 --- a/services/agents-api/internal/runtimeobs/resolver.go +++ b/services/agents-api/internal/runtimeobs/resolver.go @@ -9,7 +9,10 @@ import ( "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" ) -var ErrUnavailable = errors.New("Runtime observation unavailable") +var ( + ErrUnavailable = errors.New("Runtime observation unavailable") + ErrNotRunning = errors.New("Runtime is not running") +) type sessionStore interface { GetSession(context.Context, string, string) (store.Session, error) @@ -49,12 +52,18 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Tar if session.Environment == nil { return Target{}, errors.New("self-hosted Session is missing its Environment") } + if session.Environment.TenantID != session.TenantID || session.Environment.SessionID != session.ID { + return Target{}, errors.New("self-hosted Environment does not match resolved ownership") + } target.EnvironmentID = session.Environment.ID return target, nil case ModeManaged: if session.Environment == nil { return Target{}, errors.New("managed Session is missing its Environment") } + if session.Environment.TenantID != session.TenantID || session.Environment.SessionID != session.ID { + return Target{}, errors.New("managed Environment does not match resolved ownership") + } target.EnvironmentID = session.Environment.ID allocation, err := r.store.GetRuntimeAllocation(ctx, tenantID, target.EnvironmentID) if errors.Is(err, store.ErrNotFound) { @@ -63,7 +72,13 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Tar if err != nil { return Target{}, fmt.Errorf("resolve Runtime allocation: %w", err) } - target.Instance = Instance{AllocationID: allocation.ID, ProviderKey: allocation.ProviderKey, DeviceID: allocation.DeviceID} + if allocation.TenantID != tenantID || allocation.SessionID != session.ID || allocation.EnvironmentID != target.EnvironmentID { + return Target{}, errors.New("Runtime allocation does not match resolved ownership") + } + target.Instance = Instance{ + AllocationID: allocation.ID, ProviderKey: allocation.ProviderKey, DeviceID: allocation.DeviceID, + AllocationState: allocation.State, AllocationCreatedAt: allocation.CreatedAt, + } return target, nil default: return Target{}, errors.New("invalid stored Runtime environment type") diff --git a/services/agents-api/internal/runtimeobs/resolver_test.go b/services/agents-api/internal/runtimeobs/resolver_test.go index 738c79f2c..7edb4ba48 100644 --- a/services/agents-api/internal/runtimeobs/resolver_test.go +++ b/services/agents-api/internal/runtimeobs/resolver_test.go @@ -24,8 +24,11 @@ func (s resolverStore) GetRuntimeAllocation(context.Context, string, string) (st func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { r, err := NewResolver(resolverStore{ - session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment"}}, - allocation: store.RuntimeAllocation{ID: "allocation", ProviderKey: "provider", DeviceID: "device"}, + session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, + allocation: store.RuntimeAllocation{ + ID: "allocation", TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", + ProviderKey: "provider", DeviceID: "device", + }, }) if err != nil { t.Fatal(err) @@ -45,7 +48,7 @@ func TestResolverKeepsUnsupportedModesDistinct(t *testing.T) { environment *store.Environment }{ {mode: "none"}, - {mode: "self_hosted", environment: &store.Environment{ID: "environment"}}, + {mode: "self_hosted", environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, } { r, err := NewResolver(resolverStore{session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"` + tc.mode + `"}}`), Environment: tc.environment}}) if err != nil { @@ -60,7 +63,7 @@ func TestResolverKeepsUnsupportedModesDistinct(t *testing.T) { func TestResolverReportsManagedAllocationAsUnavailable(t *testing.T) { r, err := NewResolver(resolverStore{ - session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment"}}, + session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, allocationErr: store.ErrNotFound, }) if err != nil { @@ -71,3 +74,51 @@ func TestResolverReportsManagedAllocationAsUnavailable(t *testing.T) { t.Fatalf("allocation absence was not preserved: %+v %v", target, err) } } + +func TestResolverRejectsMismatchedEnvironmentOwnership(t *testing.T) { + for _, mode := range []string{"self_hosted", "openai_hosted"} { + for _, environment := range []store.Environment{ + {ID: "environment", TenantID: "other", SessionID: "session"}, + {ID: "environment", TenantID: "tenant", SessionID: "other"}, + } { + resolver, err := NewResolver(resolverStore{session: store.Session{ + ID: "session", TenantID: "tenant", + Configuration: []byte(`{"environment":{"type":"` + mode + `"}}`), Environment: &environment, + }}) + if err != nil { + t.Fatal(err) + } + if _, err := resolver.Resolve(t.Context(), "tenant", "session"); err == nil { + t.Fatalf("mismatched %s Environment accepted: %+v", mode, environment) + } + } + } +} + +func TestResolverRejectsMismatchedAllocationOwnership(t *testing.T) { + base := store.RuntimeAllocation{ + ID: "allocation", TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", + ProviderKey: "provider", DeviceID: "device", + } + for _, mutate := range []func(*store.RuntimeAllocation){ + func(value *store.RuntimeAllocation) { value.TenantID = "other" }, + func(value *store.RuntimeAllocation) { value.SessionID = "other" }, + func(value *store.RuntimeAllocation) { value.EnvironmentID = "other" }, + } { + allocation := base + mutate(&allocation) + resolver, err := NewResolver(resolverStore{ + session: store.Session{ + ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), + Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}, + }, + allocation: allocation, + }) + if err != nil { + t.Fatal(err) + } + if _, err := resolver.Resolve(t.Context(), "tenant", "session"); err == nil { + t.Fatalf("mismatched allocation accepted: %+v", allocation) + } + } +} diff --git a/services/agents-api/internal/runtimeobs/sample.go b/services/agents-api/internal/runtimeobs/sample.go index bf653c174..5b36e81f1 100644 --- a/services/agents-api/internal/runtimeobs/sample.go +++ b/services/agents-api/internal/runtimeobs/sample.go @@ -4,6 +4,7 @@ package runtimeobs import ( "errors" + "math" "time" ) @@ -36,19 +37,23 @@ type Sample struct { } func (s Sample) validate(now time.Time) error { - if s.ObservedAt.IsZero() || s.ObservedAt.After(now) { + if s.ObservedAt.IsZero() || s.ObservedAt.Unix() < 0 || s.ObservedAt.After(now) { return errors.New("invalid Runtime observation time") } - if s.StartedAt != nil && (s.StartedAt.IsZero() || s.StartedAt.After(s.ObservedAt)) { + if s.StartedAt != nil && (s.StartedAt.IsZero() || s.StartedAt.Unix() < 0 || s.StartedAt.After(s.ObservedAt)) { return errors.New("invalid Runtime start time") } - if s.CPUUsageSecondsTotal != nil && *s.CPUUsageSecondsTotal < 0 { + if s.CPUUsageSecondsTotal != nil && (*s.CPUUsageSecondsTotal < 0 || math.IsNaN(*s.CPUUsageSecondsTotal) || math.IsInf(*s.CPUUsageSecondsTotal, 0)) { return errors.New("invalid Runtime CPU usage") } - if s.CPUCapacityCores != nil && *s.CPUCapacityCores <= 0 { + if s.CPUCapacityCores != nil && (*s.CPUCapacityCores <= 0 || math.IsNaN(*s.CPUCapacityCores) || math.IsInf(*s.CPUCapacityCores, 0)) { return errors.New("invalid Runtime CPU capacity") } - if s.MemoryLimitBytes != nil && *s.MemoryLimitBytes == 0 { + const maxSafeJSONInteger = uint64(1<<53 - 1) + if s.MemoryUsageBytes != nil && *s.MemoryUsageBytes > maxSafeJSONInteger { + return errors.New("Runtime memory usage exceeds the public JSON integer range") + } + if s.MemoryLimitBytes != nil && (*s.MemoryLimitBytes == 0 || *s.MemoryLimitBytes > maxSafeJSONInteger) { return errors.New("invalid Runtime memory limit") } return nil diff --git a/services/agents-api/internal/runtimeobs/service.go b/services/agents-api/internal/runtimeobs/service.go index e05f7c678..e56d17348 100644 --- a/services/agents-api/internal/runtimeobs/service.go +++ b/services/agents-api/internal/runtimeobs/service.go @@ -8,10 +8,12 @@ import ( ) type Observation struct { - Target Target - Status Status - Sample *Sample - Reason string + Target Target + Status Status + Sample *Sample + Reason string + ProviderType string + ResolvedAt time.Time } type Service struct { @@ -36,25 +38,59 @@ func NewService(resolver TargetResolver, sources map[string]Source) (*Service, e func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string) (Observation, error) { target, err := s.resolver.Resolve(ctx, tenantID, sessionID) + resolvedAt := s.now() if errors.Is(err, ErrUnavailable) { - return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_allocation_unavailable"}, nil + if target.TenantID != tenantID || target.SessionID != sessionID || target.Mode != ModeManaged || target.EnvironmentID == "" { + return Observation{}, errors.New("Runtime observation resolver returned invalid pending allocation identity") + } + return Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, nil } if err != nil { return Observation{}, err } + if target.TenantID != tenantID || target.SessionID != sessionID { + return Observation{}, errors.New("Runtime observation resolver returned mismatched ownership") + } + if (target.Mode == ModeNone && target.EnvironmentID != "") || + ((target.Mode == ModeSelfHosted || target.Mode == ModeManaged) && target.EnvironmentID == "") { + return Observation{}, errors.New("Runtime observation resolver returned mismatched Environment identity") + } if target.Mode == ModeNone || target.Mode == ModeSelfHosted { - return Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable"}, nil + return Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: resolvedAt}, nil } if target.Mode != ModeManaged || target.Instance.AllocationID == "" || target.Instance.ProviderKey == "" { return Observation{}, errors.New("invalid managed Runtime observation target") } + if !target.Instance.AllocationCreatedAt.IsZero() && + (target.Instance.AllocationCreatedAt.Unix() < 0 || target.Instance.AllocationCreatedAt.After(resolvedAt)) { + return Observation{}, errors.New("invalid managed Runtime allocation creation time") + } + switch target.Instance.AllocationState { + case "creating": + return Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, nil + case "cleanup_pending", "released": + return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ResolvedAt: resolvedAt}, nil + case "running": + default: + return Observation{}, errors.New("invalid managed Runtime allocation state") + } source, ok := s.sources[target.Instance.ProviderKey] if !ok { - return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_source_unavailable"}, nil + return Observation{Target: target, Status: StatusUnavailable, Reason: "source_not_configured", ResolvedAt: resolvedAt}, nil + } + providerType := "" + if typed, ok := source.(interface{ ObservationProviderType() string }); ok { + providerType = typed.ObservationProviderType() } sample, err := source.Observe(ctx, target) + if errors.Is(err, context.DeadlineExceeded) { + return Observation{Target: target, Status: StatusUnavailable, Reason: "sample_timeout", ProviderType: providerType, ResolvedAt: s.now()}, nil + } + if errors.Is(err, ErrNotRunning) { + return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ProviderType: providerType, ResolvedAt: s.now()}, nil + } if errors.Is(err, ErrUnavailable) { - return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_sample_unavailable"}, nil + return Observation{Target: target, Status: StatusUnavailable, Reason: "sample_unavailable", ProviderType: providerType, ResolvedAt: s.now()}, nil } if err != nil { return Observation{}, fmt.Errorf("observe Runtime: %w", err) @@ -62,5 +98,5 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string if err := sample.validate(s.now()); err != nil { return Observation{}, err } - return Observation{Target: target, Status: StatusObserved, Sample: &sample}, nil + return Observation{Target: target, Status: StatusObserved, Sample: &sample, ProviderType: providerType, ResolvedAt: s.now()}, nil } diff --git a/services/agents-api/internal/runtimeobs/service_test.go b/services/agents-api/internal/runtimeobs/service_test.go index bff6bf33d..6cb76682f 100644 --- a/services/agents-api/internal/runtimeobs/service_test.go +++ b/services/agents-api/internal/runtimeobs/service_test.go @@ -3,6 +3,7 @@ package runtimeobs import ( "context" "errors" + "math" "testing" "time" ) @@ -12,8 +13,15 @@ type fixedResolver struct { err error } -func (r fixedResolver) Resolve(context.Context, string, string) (Target, error) { - return r.target, r.err +func (r fixedResolver) Resolve(_ context.Context, tenant, session string) (Target, error) { + target := r.target + if target.TenantID == "" { + target.TenantID = tenant + } + if target.SessionID == "" { + target.SessionID = session + } + return target, r.err } type fixedSource struct { @@ -27,28 +35,48 @@ func (s *fixedSource) Observe(context.Context, Target) (Sample, error) { return s.sample, s.err } +type typedSource struct { + *fixedSource + providerType string +} + +func (s typedSource) ObservationProviderType() string { return s.providerType } + +type blockingSource struct{} + +func (blockingSource) Observe(ctx context.Context, _ Target) (Sample, error) { + <-ctx.Done() + return Sample{}, ctx.Err() +} + +func (blockingSource) ObservationProviderType() string { return "docker" } + func TestServiceDoesNotCallSourcesForUnsupportedModes(t *testing.T) { for _, mode := range []Mode{ModeNone, ModeSelfHosted} { source := &fixedSource{} - service, err := NewService(fixedResolver{target: Target{Mode: mode}}, map[string]Source{"provider": source}) + target := Target{Mode: mode} + if mode == ModeSelfHosted { + target.EnvironmentID = "environment" + } + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": source}) if err != nil { t.Fatal(err) } observation, err := service.ObserveSession(t.Context(), "tenant", "session") - if err != nil || observation.Status != StatusUnsupported || source.calls != 0 { + if err != nil || observation.Status != StatusUnsupported || observation.Reason != "runtime_mode_not_observable" || source.calls != 0 { t.Fatalf("unsupported mode touched a source: %+v %v calls=%d", observation, err, source.calls) } } } func TestServicePreservesUnavailableAndObservedZero(t *testing.T) { - target := Target{Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider"}} + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} service, err := NewService(fixedResolver{target: target}, nil) if err != nil { t.Fatal(err) } observation, err := service.ObserveSession(t.Context(), "tenant", "session") - if err != nil || observation.Status != StatusUnavailable || observation.Sample != nil { + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "source_not_configured" || observation.Sample != nil { t.Fatalf("missing source was not unavailable: %+v %v", observation, err) } @@ -68,15 +96,20 @@ func TestServicePreservesUnavailableAndObservedZero(t *testing.T) { } func TestServiceMapsOnlyDeclaredUnavailability(t *testing.T) { - target := Target{Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider"}} + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} for _, tc := range []struct { - err error - wantError bool + err error + wantReason string + wantError bool }{ - {err: ErrUnavailable}, + {err: ErrUnavailable, wantReason: "sample_unavailable"}, + {err: ErrNotRunning, wantReason: "runtime_not_running"}, + {err: context.DeadlineExceeded, wantReason: "sample_timeout"}, {err: errors.New("Docker permission denied"), wantError: true}, } { - service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": &fixedSource{err: tc.err}}) + service, err := NewService(fixedResolver{target: target}, map[string]Source{ + "provider": typedSource{fixedSource: &fixedSource{err: tc.err}, providerType: "docker"}, + }) if err != nil { t.Fatal(err) } @@ -84,8 +117,116 @@ func TestServiceMapsOnlyDeclaredUnavailability(t *testing.T) { if (err != nil) != tc.wantError { t.Fatalf("wrong error classification: %+v %v", observation, err) } - if !tc.wantError && observation.Status != StatusUnavailable { + if !tc.wantError && (observation.Status != StatusUnavailable || observation.Reason != tc.wantReason || observation.ProviderType != "docker") { t.Fatalf("declared unavailability was not mapped: %+v", observation) } } } + +func TestServiceMapsAnActualSourceDeadlineWithoutLeakingIt(t *testing.T) { + target := Target{ + EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}, + } + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": blockingSource{}}) + if err != nil { + t.Fatal(err) + } + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Millisecond) + defer cancel() + observation, err := service.ObserveSession(ctx, "tenant", "session") + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "sample_timeout" || observation.ProviderType != "docker" { + t.Fatalf("source deadline was not safely classified: %+v %v", observation, err) + } +} + +func TestServiceClassifiesResolverAndTerminalAllocationUnavailability(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + service, err := NewService(fixedResolver{target: Target{SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged}, err: ErrUnavailable}, nil) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "allocation_pending" || !observation.ResolvedAt.Equal(now) { + t.Fatalf("pending allocation was not classified: %+v %v", observation, err) + } + + source := &fixedSource{} + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "creating"}} + service, err = NewService(fixedResolver{target: target}, map[string]Source{"provider": source}) + if err != nil { + t.Fatal(err) + } + observation, err = service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "allocation_pending" || source.calls != 0 { + t.Fatalf("creating allocation reached its provider: %+v %v calls=%d", observation, err, source.calls) + } + + target = Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "released"}} + service, err = NewService(fixedResolver{target: target}, map[string]Source{"provider": &fixedSource{}}) + if err != nil { + t.Fatal(err) + } + observation, err = service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "runtime_not_running" { + t.Fatalf("released allocation was not classified: %+v %v", observation, err) + } + + source = &fixedSource{} + target = Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{ + AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running", AllocationCreatedAt: time.Now().Add(time.Hour), + }} + service, err = NewService(fixedResolver{target: target}, map[string]Source{"provider": source}) + if err != nil { + t.Fatal(err) + } + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err == nil || source.calls != 0 { + t.Fatalf("future allocation creation reached its provider: %v calls=%d", err, source.calls) + } +} + +func TestServiceRejectsUnsafeProviderSamples(t *testing.T) { + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + preEpoch := time.Unix(-1, 0).UTC() + tooLarge := uint64(1 << 53) + for _, sample := range []Sample{ + {ObservedAt: preEpoch}, + {ObservedAt: now, StartedAt: &preEpoch}, + {ObservedAt: now, CPUUsageSecondsTotal: float64Pointer(-1)}, + {ObservedAt: now, CPUUsageSecondsTotal: float64Pointer(math.NaN())}, + {ObservedAt: now, CPUUsageSecondsTotal: float64Pointer(math.Inf(1))}, + {ObservedAt: now, CPUCapacityCores: float64Pointer(0)}, + {ObservedAt: now, MemoryUsageBytes: &tooLarge}, + {ObservedAt: now, MemoryLimitBytes: &tooLarge}, + } { + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": &fixedSource{sample: sample}}) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err == nil { + t.Fatalf("unsafe sample accepted: %+v", sample) + } + } +} + +func TestServiceRejectsMismatchedResolvedOwnership(t *testing.T) { + for _, target := range []Target{ + {TenantID: "other", SessionID: "session", Mode: ModeNone}, + {TenantID: "tenant", SessionID: "other", Mode: ModeNone}, + {TenantID: "tenant", SessionID: "session", EnvironmentID: "unexpected", Mode: ModeNone}, + {TenantID: "tenant", SessionID: "session", Mode: ModeSelfHosted}, + } { + service, err := NewService(fixedResolver{target: target}, nil) + if err != nil { + t.Fatal(err) + } + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err == nil { + t.Fatalf("mismatched ownership accepted: %+v", target) + } + } +} + +func float64Pointer(value float64) *float64 { return &value } diff --git a/services/agents-api/internal/sandbox/docker/resources.go b/services/agents-api/internal/sandbox/docker/resources.go index 2f444c961..f0f057de0 100644 --- a/services/agents-api/internal/sandbox/docker/resources.go +++ b/services/agents-api/internal/sandbox/docker/resources.go @@ -15,6 +15,8 @@ import ( var _ runtimeobs.Source = (*Provider)(nil) +func (*Provider) ObservationProviderType() string { return "docker" } + // Observe is read-only. Inspect verifies allocation ownership before Docker // statistics are requested; it never renews or changes the container. func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runtimeobs.Sample, error) { @@ -27,13 +29,13 @@ func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runti reference := sandbox.Reference{TenantID: target.TenantID, EnvironmentID: target.EnvironmentID, AllocationID: target.Instance.AllocationID} inspected, err := p.inspect(ctx, reference) if errors.Is(err, sandbox.ErrNotFound) { - return runtimeobs.Sample{}, runtimeobs.ErrUnavailable + return runtimeobs.Sample{}, runtimeobs.ErrNotRunning } if err != nil { return runtimeobs.Sample{}, err } if inspected.Container.State == nil || !inspected.Container.State.Running { - return runtimeobs.Sample{}, runtimeobs.ErrUnavailable + return runtimeobs.Sample{}, runtimeobs.ErrNotRunning } result, err := p.client.ContainerStats(ctx, inspected.Container.ID, client.ContainerStatsOptions{Stream: false, IncludePreviousSample: false}) if err != nil { diff --git a/services/agents-api/internal/sandbox/docker/resources_test.go b/services/agents-api/internal/sandbox/docker/resources_test.go index ea4db1096..8a21dbff6 100644 --- a/services/agents-api/internal/sandbox/docker/resources_test.go +++ b/services/agents-api/internal/sandbox/docker/resources_test.go @@ -2,6 +2,7 @@ package docker import ( "encoding/json" + "errors" "net/http" "net/http/httptest" "strings" @@ -24,13 +25,14 @@ func TestObserveVerifiesOwnershipThenReadsOneShotStats(t *testing.T) { started := observed.Add(-time.Minute) statsRead := false omitMeasurements := false + running := true server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") switch { case r.Method == http.MethodGet && strings.HasSuffix(r.URL.Path, "/json"): _ = json.NewEncoder(w).Encode(map[string]any{ "Id": "container-id", - "State": map[string]any{"Status": "running", "Running": true, "StartedAt": started.Format(time.RFC3339Nano)}, + "State": map[string]any{"Status": "running", "Running": running, "StartedAt": started.Format(time.RFC3339Nano)}, "HostConfig": map[string]any{"NanoCpus": 2_000_000_000, "Memory": 2048}, "Config": map[string]any{"Labels": map[string]string{ labelPrefix + "installation": installationID, @@ -76,6 +78,12 @@ func TestObserveVerifiesOwnershipThenReadsOneShotStats(t *testing.T) { if err != nil || missing.CPUUsageSecondsTotal != nil || missing.MemoryUsageBytes != nil || missing.CPUCapacityCores == nil || missing.MemoryLimitBytes == nil { t.Fatalf("missing Docker measurements became zero: %+v %v", missing, err) } + running = false + statsRead = false + if _, err := p.Observe(t.Context(), target); !errors.Is(err, runtimeobs.ErrNotRunning) || statsRead { + t.Fatalf("stopped Runtime was not classified before stats: %v stats=%v", err, statsRead) + } + running = true foreign := target foreign.Instance.ProviderKey = uuid.NewString() statsRead = false From 0a10d5b3ee8dbdb1c1a22a33eeda9e33021d423f Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 20:33:57 +0800 Subject: [PATCH 04/51] Add Runtime observability dashboard --- apps/web/e2e/agents-lifecycle.spec.ts | 17 +- apps/web/e2e/fixture-core.mjs | 9 + apps/web/src/App.tsx | 103 ++++++- .../src/features/dashboard/DashboardView.css | 289 ++++++++++++++++++ .../features/dashboard/DashboardView.test.tsx | 80 ++++- .../src/features/dashboard/DashboardView.tsx | 208 ++++++++++++- .../dashboard/dashboard-model.test.ts | 94 +++++- .../src/features/dashboard/dashboard-model.ts | 212 +++++++++++++ .../dashboard/runtime-snapshot.test.ts | 57 ++++ .../features/dashboard/runtime-snapshot.ts | 59 ++++ .../agents-api/runtime-observability-api.md | 8 +- .../runtime-observability-design.md | 19 +- docs/web/README.md | 11 +- docs/web/README.zh-CN.md | 12 +- 14 files changed, 1140 insertions(+), 38 deletions(-) create mode 100644 apps/web/src/features/dashboard/runtime-snapshot.test.ts create mode 100644 apps/web/src/features/dashboard/runtime-snapshot.ts diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 640a954ef..9fd4d700a 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2329,10 +2329,11 @@ test("presents Dashboard page-chain results and System boundaries without extra await expect.poll(async () => { const entries = await fixtureRequests(request); return [count(entries, "/v1/agents"), count(entries, "/v1/agents/sessions")]; - }).toEqual([count(before, "/v1/agents") + 1, count(before, "/v1/agents/sessions") + 1]); + }).toEqual([count(before, "/v1/agents") + 1, count(before, "/v1/agents/sessions") + 2]); const after = await fixtureRequests(request); expect(count(after, "/v1/agents")).toBe(count(before, "/v1/agents") + 1); - expect(count(after, "/v1/agents/sessions")).toBe(count(before, "/v1/agents/sessions") + 1); + expect(count(after, "/v1/agents/sessions")).toBe(count(before, "/v1/agents/sessions") + 2); + expect(count(after, "/v1/agents/runtime-observations")).toBe(count(before, "/v1/agents/runtime-observations") + 1); for (const path of detailPaths) expect(count(after, path)).toBe(count(before, path)); await attachScreenshot(page, testInfo, "desktop-dashboard-loaded-snapshot"); @@ -2387,11 +2388,12 @@ test("presents Dashboard page-chain results and System boundaries without extra return [count(entries, "/v1/agents"), count(entries, "/v1/agents/sessions")]; }).toEqual([ count(beforeSystemRefresh, "/v1/agents") + 1, - count(beforeSystemRefresh, "/v1/agents/sessions") + 1, + count(beforeSystemRefresh, "/v1/agents/sessions") + 2, ]); const afterSystemRefresh = await fixtureRequests(request); expect(count(afterSystemRefresh, "/v1/agents")).toBe(count(beforeSystemRefresh, "/v1/agents") + 1); - expect(count(afterSystemRefresh, "/v1/agents/sessions")).toBe(count(beforeSystemRefresh, "/v1/agents/sessions") + 1); + expect(count(afterSystemRefresh, "/v1/agents/sessions")).toBe(count(beforeSystemRefresh, "/v1/agents/sessions") + 2); + expect(count(afterSystemRefresh, "/v1/agents/runtime-observations")).toBe(count(beforeSystemRefresh, "/v1/agents/runtime-observations") + 1); for (const path of detailPaths) expect(count(afterSystemRefresh, path)).toBe(count(beforeSystemRefresh, path)); await attachScreenshot(page, testInfo, "desktop-system-contract-boundary"); @@ -2480,7 +2482,9 @@ test("publishes Dashboard counts only after every top-level Agent and Session pa await expect(dashboard.locator(".dashboard-summary > div").filter({ hasText: "Agents" })).toContainText("3"); await expect(dashboard.locator(".dashboard-summary > div").filter({ hasText: "Sessions" })).toContainText("2"); expect(agentAfters).toEqual([null, "agent_b"]); - expect(sessionAfters).toEqual([null, "session_snapshot"]); + expect(sessionAfters).toHaveLength(4); + expect(sessionAfters.filter((after) => after === null)).toHaveLength(2); + expect(sessionAfters.filter((after) => after === "session_snapshot")).toHaveLength(2); }); test("keeps the previous Dashboard result when pagination exceeds the safety limit", async ({ page, request }) => { @@ -2523,8 +2527,9 @@ test("keeps the previous Dashboard result when pagination exceeds the safety lim await refresh.click(); await expect.poll(() => reads).toBe(100); await expect(refresh).toBeEnabled(); - await expect(dashboard).toContainText("Using the last successful snapshot"); + await expect(dashboard).toContainText("Snapshot incomplete"); await expect(dashboard).toContainText("collection pagination exceeded the Web safety limit"); + await expect(dashboard).toContainText("Sessions changed while Runtime observations were loading"); await expect(loadedAgents).toContainText("3"); await dashboard.getByRole("button", { name: "Connection settings" }).click(); const connectionDialog = page.getByRole("dialog", { name: "Connect an Agent Core" }); diff --git a/apps/web/e2e/fixture-core.mjs b/apps/web/e2e/fixture-core.mjs index 5aebe44dd..640616184 100644 --- a/apps/web/e2e/fixture-core.mjs +++ b/apps/web/e2e/fixture-core.mjs @@ -995,6 +995,15 @@ const server = http.createServer(async (request, response) => { return sendJson(response, created, 201); } + // The legacy Web fixture uses human-readable Session IDs for interaction + // assertions. Runtime observation resources require canonical UUIDs, so this + // fixture advertises an empty, valid collection instead of inventing a false + // identity join. Positive Runtime rendering is covered by the typed component + // and coordinator tests with canonical identities. + if (request.method === "GET" && url.pathname === "/v1/agents/runtime-observations") { + return sendJson(response, page([])); + } + if (request.method === "GET" && url.pathname === "/v1/agents/sessions") { trackAbort(response, "sessionListReads"); const control = consumeControl("sessionList"); diff --git a/apps/web/src/App.tsx b/apps/web/src/App.tsx index e4a7d6f85..c0925e986 100644 --- a/apps/web/src/App.tsx +++ b/apps/web/src/App.tsx @@ -33,6 +33,12 @@ import { requestAgentUpdate, } from "./features/agents/agent-actions"; import { DashboardView } from "./features/dashboard/DashboardView"; +import { + loadRuntimeDashboardSnapshot, + RUNTIME_SNAPSHOT_REFRESH_MS, + RUNTIME_SNAPSHOT_TIMEOUT_MS, + type RuntimeDashboardSnapshot, +} from "./features/dashboard/runtime-snapshot"; import { SessionsView, type SessionDetailState, @@ -270,6 +276,10 @@ export function App() { const [sessionCollectionState, setSessionCollectionState] = useState("connecting"); const [sessionCollectionError, setSessionCollectionError] = useState(null); const [sessionCollectionHasSnapshot, setSessionCollectionHasSnapshot] = useState(false); + const [runtimeSnapshot, setRuntimeSnapshot] = useState(null); + const [runtimeCollectionState, setRuntimeCollectionState] = useState("connecting"); + const [runtimeCollectionError, setRuntimeCollectionError] = useState(null); + const [runtimeCollectionHasSnapshot, setRuntimeCollectionHasSnapshot] = useState(false); const [sessionAgentFilter, setSessionAgentFilter] = useState(null); const [filteredSessions, setFilteredSessions] = useState([]); const [filteredSessionCollectionState, setFilteredSessionCollectionState] = useState("connecting"); @@ -305,6 +315,8 @@ export function App() { const sessionCollectionRequestRef = useRef(0); const agentCollectionAbortRef = useRef(null); const sessionCollectionAbortRef = useRef(null); + const runtimeCollectionAbortRef = useRef(null); + const runtimeCollectionRequestRef = useRef(0); const filteredSessionCollectionAbortRef = useRef(null); const filteredSessionCollectionRequestRef = useRef(0); const sessionAgentFilterRef = useRef(sessionAgentFilter); @@ -516,6 +528,54 @@ export function App() { } }, [core, coreGeneration, notify]); + const refreshRuntimeSnapshot = useCallback(async () => { + if (coreGeneration !== connectionGenerationRef.current) return false; + runtimeCollectionAbortRef.current?.abort(); + const controller = new AbortController(); + runtimeCollectionAbortRef.current = controller; + const request = runtimeCollectionRequestRef.current + 1; + runtimeCollectionRequestRef.current = request; + let timedOut = false; + const timeout = window.setTimeout(() => { + timedOut = true; + controller.abort(); + }, RUNTIME_SNAPSHOT_TIMEOUT_MS); + setRuntimeCollectionState("connecting"); + setRuntimeCollectionError(null); + try { + const result = await settleCollection(() => loadRuntimeDashboardSnapshot( + core, + () => sessionCollectionRevisionRef.current, + controller.signal, + )); + if ( + coreGeneration !== connectionGenerationRef.current || + request !== runtimeCollectionRequestRef.current + ) return false; + if (result.status === "rejected") { + if (isAbort(result.reason) && !timedOut) return false; + const message = timedOut + ? "Runtime snapshot exceeded the 15 second Web refresh budget." + : errorMessage(result.reason); + setRuntimeCollectionState("failed"); + setRuntimeCollectionError(message); + return false; + } + if (result.value === null) { + setRuntimeCollectionState("failed"); + setRuntimeCollectionError("Session data changed while Runtime observations were loading. The previous complete snapshot was retained."); + return false; + } + setRuntimeSnapshot(result.value); + setRuntimeCollectionHasSnapshot(true); + setRuntimeCollectionState("ready"); + return true; + } finally { + window.clearTimeout(timeout); + if (runtimeCollectionAbortRef.current === controller) runtimeCollectionAbortRef.current = null; + } + }, [core, coreGeneration]); + const refreshFilteredSessions = useCallback(async (agentId: string) => { if ( coreGeneration !== connectionGenerationRef.current || @@ -855,9 +915,10 @@ export function App() { void refreshSessions(); void refreshVaults(); void refreshEnvironmentTemplates(); + void refreshRuntimeSnapshot(); const filter = sessionAgentFilterRef.current; if (filter) void refreshFilteredSessions(filter); - }, [refreshAgents, refreshEnvironmentTemplates, refreshFilteredSessions, refreshSessions, refreshVaults]); + }, [refreshAgents, refreshEnvironmentTemplates, refreshFilteredSessions, refreshRuntimeSnapshot, refreshSessions, refreshVaults]); const changeSessionAgentFilter = useCallback((agentId: string | null) => { if (sessionAgentFilterRef.current === agentId) return; @@ -894,6 +955,10 @@ export function App() { setSessions([]); setAgentCollectionHasSnapshot(false); setSessionCollectionHasSnapshot(false); + setRuntimeSnapshot(null); + setRuntimeCollectionState("connecting"); + setRuntimeCollectionError(null); + setRuntimeCollectionHasSnapshot(false); setItems([]); setTurns([]); setEnvironmentObservations(new Map()); @@ -908,7 +973,29 @@ export function App() { void refreshSessions(); void refreshVaults(); void refreshEnvironmentTemplates(); - }, [refreshAgents, refreshEnvironmentTemplates, refreshSessions, refreshVaults]); + void refreshRuntimeSnapshot(); + }, [refreshAgents, refreshEnvironmentTemplates, refreshRuntimeSnapshot, refreshSessions, refreshVaults]); + + useEffect(() => { + if (view !== "dashboard") return; + let timer: number | null = null; + const schedule = () => { + const jitter = Math.floor(Math.random() * 5_000); + timer = window.setTimeout(() => { + if (!document.hidden) void refreshRuntimeSnapshot(); + schedule(); + }, RUNTIME_SNAPSHOT_REFRESH_MS + jitter); + }; + const onVisibilityChange = () => { + if (!document.hidden) void refreshRuntimeSnapshot(); + }; + document.addEventListener("visibilitychange", onVisibilityChange); + schedule(); + return () => { + if (timer !== null) window.clearTimeout(timer); + document.removeEventListener("visibilitychange", onVisibilityChange); + }; + }, [refreshRuntimeSnapshot, view]); useEffect(() => { filteredSessionCollectionAbortRef.current?.abort(); @@ -1943,6 +2030,7 @@ export function App() { connectionGenerationRef.current += 1; agentCollectionRequestRef.current += 1; sessionCollectionRequestRef.current += 1; + runtimeCollectionRequestRef.current += 1; filteredSessionCollectionRequestRef.current += 1; vaultCollectionRequestRef.current += 1; agentCollectionRevisionRef.current = 0; @@ -1966,6 +2054,8 @@ export function App() { agentCollectionAbortRef.current = null; sessionCollectionAbortRef.current?.abort(); sessionCollectionAbortRef.current = null; + runtimeCollectionAbortRef.current?.abort(); + runtimeCollectionAbortRef.current = null; filteredSessionCollectionAbortRef.current?.abort(); filteredSessionCollectionAbortRef.current = null; vaultCollectionAbortRef.current?.abort(); @@ -1981,6 +2071,10 @@ export function App() { setSessionCollectionState("connecting"); setSessionCollectionError(null); setSessionCollectionHasSnapshot(false); + setRuntimeSnapshot(null); + setRuntimeCollectionState("connecting"); + setRuntimeCollectionError(null); + setRuntimeCollectionHasSnapshot(false); sessionAgentFilterRef.current = null; filteredSessionsRef.current = []; setSessionAgentFilter(null); @@ -2108,6 +2202,10 @@ export function App() { sessionCollectionState={sessionCollectionState} sessionCollectionError={sessionCollectionError} sessionCollectionHasSnapshot={sessionCollectionHasSnapshot} + runtimeSnapshot={runtimeSnapshot} + runtimeCollectionState={runtimeCollectionState} + runtimeCollectionError={runtimeCollectionError} + runtimeCollectionHasSnapshot={runtimeCollectionHasSnapshot} onRefresh={refreshDashboard} onCreateAgent={openAgentSetup} onStartSession={() => openSessionSetup()} @@ -2208,6 +2306,7 @@ export function App() { refreshing={ agentCollectionState === "connecting" || sessionCollectionState === "connecting" || + runtimeCollectionState === "connecting" || vaultCollectionState === "connecting" } onRefresh={refreshDashboard} diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 999874c09..b76218653 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -613,10 +613,289 @@ border-radius: 50%; } +.dashboard-runtime-panel { + margin-top: 16px; + overflow: hidden; +} + +.dashboard-runtime-freshness { + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 10px; + font-variant-numeric: tabular-nums; + line-height: 15px; + text-align: right; +} + +.dashboard-runtime-summary { + display: grid; + grid-template-columns: repeat(4, minmax(0, 1fr)); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-metric { + display: grid; + min-width: 0; + min-height: 106px; + padding: 16px; + grid-template-columns: 32px minmax(0, 1fr); + align-items: flex-start; + gap: 10px; +} + +.dashboard-runtime-metric + .dashboard-runtime-metric { + border-left: 1px solid var(--line); +} + +.dashboard-runtime-metric-icon { + display: inline-flex; + width: 32px; + height: 32px; + align-items: center; + justify-content: center; + color: var(--accent); + background: color-mix(in srgb, var(--accent) 8%, var(--surface)); + border: 1px solid color-mix(in srgb, var(--accent) 18%, var(--line)); + border-radius: 8px; +} + +.dashboard-runtime-metric > span:last-child { + display: grid; + min-width: 0; + gap: 2px; +} + +.dashboard-runtime-metric small, +.dashboard-runtime-metric > span:last-child > span { + color: var(--fg-muted); + font-size: 10px; + line-height: 14px; +} + +.dashboard-runtime-metric strong { + overflow: hidden; + font-family: var(--font-mono); + font-size: 17px; + font-variant-numeric: tabular-nums slashed-zero; + font-weight: 550; + line-height: 24px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-ledger { + min-width: 960px; +} + +.dashboard-runtime-panel:has(.dashboard-runtime-ledger) { + overflow-x: auto; +} + +.dashboard-runtime-header, +.dashboard-runtime-row { + display: grid; + grid-template-columns: minmax(190px, 1.45fr) minmax(140px, 1fr) minmax(130px, .9fr) minmax(150px, 1fr) minmax(140px, .9fr) minmax(110px, .75fr); + align-items: center; + column-gap: 14px; +} + +.dashboard-runtime-header { + min-height: 34px; + padding: 7px 14px; + color: var(--fg-muted); + font-size: 9px; + font-weight: 600; + letter-spacing: .04em; + line-height: 13px; + background: var(--surface-subtle); + border-bottom: 1px solid var(--line); + text-transform: uppercase; +} + +.dashboard-runtime-entry + .dashboard-runtime-entry { + border-top: 1px solid var(--line); +} + +.dashboard-runtime-row { + min-height: 70px; + padding: 11px 14px; + transition: background-color 150ms var(--ease-settle); +} + +.dashboard-runtime-row:hover, +.dashboard-runtime-row:focus-within { + background: var(--hover); +} + +.dashboard-runtime-row > span { + display: grid; + min-width: 0; + gap: 3px; +} + +.dashboard-runtime-row small { + overflow: hidden; + color: var(--fg-muted); + font-size: 9px; + line-height: 13px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-identity button { + width: fit-content; + max-width: 100%; + padding: 0; + overflow: hidden; + color: var(--fg); + font-size: 12px; + font-weight: 600; + line-height: 17px; + text-align: left; + text-overflow: ellipsis; + white-space: nowrap; + background: transparent; + border: 0; +} + +.dashboard-runtime-identity button:hover, +.dashboard-runtime-identity button:focus-visible { + color: var(--accent); +} + +.dashboard-runtime-status { + display: inline-flex; + width: fit-content; + align-items: center; + gap: 6px; + font-size: 10px; + font-weight: 550; + line-height: 15px; +} + +.dashboard-runtime-status > span { + width: 7px; + height: 7px; + background: var(--fg-muted); + border-radius: 50%; +} + +.dashboard-runtime-status-observed > span { + background: var(--success); + box-shadow: 0 0 0 3px color-mix(in srgb, var(--success) 11%, transparent); +} + +.dashboard-runtime-status-unavailable { + color: var(--warning); +} + +.dashboard-runtime-status-unavailable > span { + background: var(--warning); +} + +.dashboard-runtime-value strong { + overflow: hidden; + font-family: var(--font-mono); + font-size: 12px; + font-variant-numeric: tabular-nums slashed-zero; + font-weight: 550; + line-height: 17px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-bar { + display: block; + width: 100%; + max-width: 110px; + height: 3px; + margin-top: 2px; + overflow: hidden; + background: var(--surface-subtle); + border-radius: 999px; +} + +.dashboard-runtime-bar i { + display: block; + height: 100%; + background: var(--accent); + border-radius: inherit; +} + +.dashboard-runtime-detail { + padding: 0 14px 10px; + color: var(--fg-muted); + font-size: 10px; +} + +.dashboard-runtime-detail summary { + width: fit-content; + padding: 4px 0; + cursor: pointer; + user-select: none; +} + +.dashboard-runtime-detail[open] { + padding-top: 8px; + background: var(--surface-subtle); + border-top: 1px dashed var(--line); +} + +.dashboard-runtime-detail dl { + display: grid; + margin: 8px 0; + grid-template-columns: repeat(3, minmax(0, 1fr)); + gap: 8px 14px; +} + +.dashboard-runtime-detail dl > div { + min-width: 0; +} + +.dashboard-runtime-detail dt { + color: var(--fg-muted); + font-size: 9px; + text-transform: uppercase; +} + +.dashboard-runtime-detail dd { + margin: 2px 0 0; + overflow: hidden; + color: var(--fg); + font-family: var(--font-mono); + font-size: 10px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-detail button { + display: inline-flex; + align-items: center; + padding: 4px 7px; + gap: 4px; + color: var(--accent); + font-size: 10px; + background: transparent; + border: 1px solid color-mix(in srgb, var(--accent) 25%, var(--line)); + border-radius: 5px; +} + @media (max-width: 980px) { .dashboard-primary-grid { grid-template-columns: 1fr; } + + .dashboard-runtime-summary { + grid-template-columns: repeat(2, minmax(0, 1fr)); + } + + .dashboard-runtime-metric:nth-child(3) { + border-left: 0; + } + + .dashboard-runtime-metric:nth-child(n + 3) { + border-top: 1px solid var(--line); + } } @media (max-width: 760px) { @@ -701,6 +980,16 @@ align-items: flex-start; } + .dashboard-runtime-summary { + grid-template-columns: 1fr; + } + + .dashboard-runtime-metric + .dashboard-runtime-metric, + .dashboard-runtime-metric:nth-child(3) { + border-top: 1px solid var(--line); + border-left: 0; + } + .dashboard-panel header p { white-space: normal; } diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index f7c617ba4..5d70620aa 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -1,7 +1,7 @@ import { renderToStaticMarkup } from "react-dom/server"; import { describe, expect, it } from "vitest"; -import type { AgentSession, SavedAgent } from "@agents-core-web/agents-client"; +import type { AgentSession, RuntimeObservation, SavedAgent } from "@agents-core-web/agents-client"; import { DashboardView, type DashboardViewProps } from "./DashboardView"; @@ -74,6 +74,10 @@ function render(overrides: Partial = {}): string { sessionCollectionState="ready" sessionCollectionError={null} sessionCollectionHasSnapshot + runtimeSnapshot={{ sessions: [], observations: [], loadedAt: 1_700_000_000_000 }} + runtimeCollectionState="ready" + runtimeCollectionError={null} + runtimeCollectionHasSnapshot {...callbacks} {...overrides} />, @@ -121,6 +125,21 @@ describe("Dashboard loaded-result presentation", () => { expect(html).not.toContain("Sessions: Agent core request failed (502)"); }); + it("keeps a Runtime-only 503 scoped to the optional observation feature", () => { + const html = render({ + runtimeSnapshot: null, + runtimeCollectionState: "failed", + runtimeCollectionError: "Agent core request failed (503).", + runtimeCollectionHasSnapshot: false, + }); + + expect(html).toContain("Runtime: Agent core request failed (503)."); + expect(html).toContain("Runtime observations unavailable"); + expect(html).not.toContain("Agent Core backend is not ready"); + expect(html).not.toContain("Core backend is offline"); + expect(html).not.toContain("Open startup guide"); + }); + it("renders a compact actionable overview while preserving Environment qualifications", () => { const selfHosted: AgentSession["environment"] = { type: "self_hosted", @@ -174,6 +193,65 @@ describe("Dashboard loaded-result presentation", () => { expect(html).not.toContain("Execution ready"); }); + it("renders current Docker resources without inventing a CPU percentage or history", () => { + const hosted = session("11111111-1111-4111-8111-111111111111", { + metadata: { title: "Managed research" }, + environment: { + type: "openai_hosted", + id: "22222222-2222-4222-8222-222222222222", + capability_directories: [], + network: { access: "enabled", allowed_domains: [] }, + packages: { npm: [], python: [], system: [] }, + files: [], + plugins: [], + skills: [], + }, + usage: { + input_tokens: 30, + output_tokens: 12, + total_tokens: 42, + input_tokens_details: { cached_tokens: 7 }, + output_tokens_details: { reasoning_tokens: 3 }, + }, + }); + const observation: RuntimeObservation = { + id: hosted.id, + object: "agent.runtime_observation", + session_id: hosted.id, + environment_id: "22222222-2222-4222-8222-222222222222", + mode: "openai_hosted", + provider_type: "docker", + instance: { + kind: "managed_allocation", + allocation_id: "33333333-3333-4333-8333-333333333333", + device_id: null, + connection_generation: null, + }, + status: "observed", + reason: null, + allocation_created_at: 1_700_000_000, + resolved_at: 1_700_000_100, + observed_at: 1_700_000_090, + started_at: 1_700_000_010, + cpu: { usage_seconds_total: 73.5, capacity_cores: 2, usage_cores: null, utilization_ratio: null }, + memory: { usage_bytes: 536_870_912, limit_bytes: 2_147_483_648 }, + }; + const html = render({ + runtimeSnapshot: { sessions: [hosted], observations: [observation], loadedAt: 1_700_000_100_000 }, + }); + + expect(html).toContain("Runtime monitoring"); + expect(html).toContain("1/1 managed observed"); + expect(html).toContain("CPU time / capacity"); + expect(html).toContain("1m 13s / 2 cores"); + expect(html).toContain("512 MiB / 2.00 GiB"); + expect(html).toContain("Current Runtime observations"); + expect(html).toContain("Managed research"); + expect(html).toContain("Identity and sample details"); + expect(html).not.toContain("CPU %"); + expect(html).not.toContain("historical chart"); + }); + it("keeps partial Usage out of the primary overview", () => { const html = render({ sessions: [session("usage-unknown")] }); diff --git a/apps/web/src/features/dashboard/DashboardView.tsx b/apps/web/src/features/dashboard/DashboardView.tsx index c2cd9786d..1ab5f2d87 100644 --- a/apps/web/src/features/dashboard/DashboardView.tsx +++ b/apps/web/src/features/dashboard/DashboardView.tsx @@ -2,9 +2,13 @@ import { AlertTriangle, ArrowRight, Bot, + Cpu, + Gauge, + MemoryStick, MessageSquare, RefreshCw, Rows3, + Server, } from "lucide-react"; import { useMemo, type ReactNode } from "react"; @@ -14,12 +18,18 @@ import { StatusIcon, type StatusKind } from "../../components/StatusIcon"; import { backendFailureStatus } from "../../lib/core-readiness"; import { buildDashboardSnapshot, + buildRuntimeDashboardModel, dashboardEnvironmentLabel, dashboardStatusLabel, + formatDashboardBytes, + formatDashboardDuration, + formatDashboardTokens, formatDashboardTimestamp, + runtimeObservationStatusLabel, type DashboardCollectionState, type DashboardSessionRow, } from "./dashboard-model"; +import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; import "./DashboardView.css"; export interface DashboardViewProps { @@ -31,6 +41,10 @@ export interface DashboardViewProps { sessionCollectionState: DashboardCollectionState; sessionCollectionError: string | null; sessionCollectionHasSnapshot: boolean; + runtimeSnapshot: RuntimeDashboardSnapshot | null; + runtimeCollectionState: DashboardCollectionState; + runtimeCollectionError: string | null; + runtimeCollectionHasSnapshot: boolean; onRefresh: () => void; onCreateAgent: () => void; onStartSession: () => void; @@ -40,6 +54,109 @@ export interface DashboardViewProps { onOpenSession: (sessionId: string) => void; } +function RuntimeMetric({ + icon, + label, + value, + detail, +}: { + icon: ReactNode; + label: string; + value: string; + detail: string; +}) { + return ( +
+ + + {label} + {value} + {detail} + +
+ ); +} + +function percent(usage: number | null | undefined, limit: number | null | undefined): number | null { + if (typeof usage !== "number" || typeof limit !== "number" || limit <= 0) return null; + return Math.min(100, Math.max(0, usage / limit * 100)); +} + +function RuntimeTable({ + snapshot, + onOpenSession, +}: { + snapshot: RuntimeDashboardSnapshot; + onOpenSession: (sessionId: string) => void; +}) { + const model = buildRuntimeDashboardModel(snapshot.sessions, snapshot.observations); + return ( +
+
+ Session / Runtime + Observation + CPU + Memory + Uptime + Tokens +
+ {model.rows.map((row) => { + const observation = row.observation; + const memoryPercent = observation.status === "observed" + ? percent(observation.memory?.usage_bytes, observation.memory?.limit_bytes) + : null; + return ( +
+
+ + + {row.session.agentLabel} · {dashboardStatusLabel(row.session.status)} + + + + + {observation.mode === "openai_hosted" ? observation.provider_type ?? "Managed" : dashboardEnvironmentLabel(row.session.environmentProfile)} + + + {observation.status === "observed" ? formatDashboardDuration(observation.cpu?.usage_seconds_total ?? null) : "—"} + {observation.status === "observed" && typeof observation.cpu?.capacity_cores === "number" + ? `${observation.cpu.capacity_cores.toLocaleString("en-US")} cores capacity` + : "No current sample"} + + + {observation.status === "observed" ? formatDashboardBytes(observation.memory?.usage_bytes ?? null) : "—"} + {observation.status === "observed" ? `of ${formatDashboardBytes(observation.memory?.limit_bytes ?? null)}` : "No current sample"} + {memoryPercent !== null ? : null} + + + {formatDashboardDuration(row.computeUptimeSeconds)} + {row.allocationAgeSeconds === null ? "Allocation age unavailable" : `${formatDashboardDuration(row.allocationAgeSeconds)} allocation age`} + + + {formatDashboardTokens(row.session.totalTokens)} + {row.session.totalTokens === null ? "Usage not reported" : "Session reported"} + +
+
+ Identity and sample details +
+
Session
{observation.session_id}
+
Environment
{observation.environment_id ?? "Not applicable"}
+
Allocation
{observation.instance.allocation_id ?? "Not available"}
+
Provider
{observation.provider_type ?? "Not available"}
+
Resolved
{formatDashboardTimestamp(observation.resolved_at)}
+
Observed
{formatDashboardTimestamp(observation.observed_at)}
+
+ +
+
+ ); + })} +
+ ); +} + function collectionHasSnapshot(state: DashboardCollectionState, hasSnapshot: boolean): boolean { return state === "ready" || hasSnapshot; } @@ -232,6 +349,10 @@ export function DashboardView({ sessionCollectionState, sessionCollectionError, sessionCollectionHasSnapshot, + runtimeSnapshot, + runtimeCollectionState, + runtimeCollectionError, + runtimeCollectionHasSnapshot, onRefresh, onCreateAgent, onStartSession, @@ -241,21 +362,28 @@ export function DashboardView({ onOpenSession, }: DashboardViewProps) { const snapshot = useMemo(() => buildDashboardSnapshot(agents, sessions, 6, 5), [agents, sessions]); + const runtimeModel = useMemo(() => runtimeSnapshot + ? buildRuntimeDashboardModel(runtimeSnapshot.sessions, runtimeSnapshot.observations) + : null, [runtimeSnapshot]); const agentsAvailable = collectionHasSnapshot(agentCollectionState, agentCollectionHasSnapshot); const sessionsAvailable = collectionHasSnapshot(sessionCollectionState, sessionCollectionHasSnapshot); - const refreshing = agentCollectionState === "connecting" || sessionCollectionState === "connecting"; + const runtimeAvailable = runtimeCollectionState === "ready" || runtimeCollectionHasSnapshot; + const refreshing = agentCollectionState === "connecting" || sessionCollectionState === "connecting" || runtimeCollectionState === "connecting"; const hasStaleSnapshot = ( (agentCollectionState === "failed" && agentCollectionHasSnapshot) || - (sessionCollectionState === "failed" && sessionCollectionHasSnapshot) + (sessionCollectionState === "failed" && sessionCollectionHasSnapshot) || + (runtimeCollectionState === "failed" && runtimeCollectionHasSnapshot) ); - const hasUnavailableSource = !agentsAvailable || !sessionsAvailable; + const hasUnavailableSource = !agentsAvailable || !sessionsAvailable || !runtimeAvailable; const attentionCount = snapshot.statusCounts.requires_action + snapshot.statusCounts.failed; - const sourceErrors: Array = [ + const sourceErrors: Array = [ agentCollectionState === "failed" ? ["Agents", agentCollectionError] as const : null, sessionCollectionState === "failed" ? ["Sessions", sessionCollectionError] as const : null, - ].filter((entry): entry is readonly ["Agents" | "Sessions", string | null] => entry !== null); - const backendFailureStatuses = sourceErrors.map(([, error]) => backendFailureStatus(error)); - const backendUnavailable = sourceErrors.length > 0 && backendFailureStatuses.every(Boolean); + runtimeCollectionState === "failed" ? ["Runtime", runtimeCollectionError] as const : null, + ].filter((entry): entry is readonly ["Agents" | "Sessions" | "Runtime", string | null] => entry !== null); + const coreSourceErrors = sourceErrors.filter(([label]) => label !== "Runtime"); + const backendFailureStatuses = coreSourceErrors.map(([, error]) => backendFailureStatus(error)); + const backendUnavailable = coreSourceErrors.length > 0 && backendFailureStatuses.every(Boolean); const backendFailureDetail = Array.from(new Set(backendFailureStatuses.filter(Boolean))).map((status) => ( status === "network" ? "network failure" : `HTTP ${status}` )).join(" / "); @@ -304,6 +432,7 @@ export function DashboardView({
+
@@ -359,6 +488,71 @@ export function DashboardView({ +
+
+
+

Runtime monitoring

+

Current provider samples joined to an exact complete Session snapshot · no historical series

+
+ {runtimeModel ? ( + + {runtimeModel.summary.observedRuntimeCount}/{runtimeModel.summary.managedRuntimeCount} managed observed + {runtimeModel.summary.newestResolvedAt === null ? "" : ` · ${formatDashboardTimestamp(runtimeModel.summary.newestResolvedAt)}`} + + ) : null} +
+ {!runtimeAvailable || !runtimeSnapshot || !runtimeModel ? ( +

+

+ ) : ( + <> +
+ } + label="Active Runtimes" + value={runtimeModel.summary.observedRuntimeCount.toLocaleString("en-US")} + detail={`${runtimeModel.summary.managedRuntimeCount} managed · ${runtimeModel.summary.unavailableRuntimeCount} unavailable`} + /> + } + label="CPU time / capacity" + value={runtimeModel.summary.cpuUsageSecondsTotal === null && runtimeModel.summary.cpuCapacityCores === null + ? "No current sample" + : `${formatDashboardDuration(runtimeModel.summary.cpuUsageSecondsTotal)} / ${runtimeModel.summary.cpuCapacityCores?.toLocaleString("en-US") ?? "—"} cores`} + detail={`${runtimeModel.summary.cpuCoverageCount}/${runtimeModel.summary.observedRuntimeCount} observed Runtimes report CPU`} + /> + } + label="Memory" + value={runtimeModel.summary.memoryUsageBytes === null && runtimeModel.summary.memoryLimitBytes === null + ? "No current sample" + : `${formatDashboardBytes(runtimeModel.summary.memoryUsageBytes)} / ${formatDashboardBytes(runtimeModel.summary.memoryLimitBytes)}`} + detail={`${runtimeModel.summary.memoryCoverageCount}/${runtimeModel.summary.observedRuntimeCount} observed Runtimes report memory`} + /> + } + label="Reported tokens" + value={runtimeModel.summary.totalTokens === null ? "Not reported" : formatDashboardTokens(runtimeModel.summary.totalTokens)} + detail={`${runtimeModel.summary.tokenCoverageCount}/${runtimeModel.summary.sessionCount} Sessions report usage`} + /> +
+ {runtimeModel.rows.length ? ( + + ) : ( +

No Session-owned Runtime contexts in this snapshot.

+ )} + + )} +
+
diff --git a/apps/web/src/features/dashboard/dashboard-model.test.ts b/apps/web/src/features/dashboard/dashboard-model.test.ts index e638c883f..b6f35bd0b 100644 --- a/apps/web/src/features/dashboard/dashboard-model.test.ts +++ b/apps/web/src/features/dashboard/dashboard-model.test.ts @@ -1,13 +1,17 @@ import { describe, expect, it } from "vitest"; -import type { AgentSession, SavedAgent, TokenUsage } from "@agents-core-web/agents-client"; +import type { AgentSession, RuntimeObservation, SavedAgent, TokenUsage } from "@agents-core-web/agents-client"; import { buildDashboardSnapshot, + buildRuntimeDashboardModel, dashboardEnvironmentLabel, dashboardEnvironmentProfile, dashboardStatusLabel, formatDashboardTimestamp, + formatDashboardBytes, + formatDashboardDuration, + runtimeObservationStatusLabel, } from "./dashboard-model"; function agent(id: string, overrides: Partial = {}): SavedAgent { @@ -228,4 +232,92 @@ describe("Dashboard loaded-snapshot model", () => { model: "fixture/fallback", }); }); + + it("aggregates only present Runtime measurements and preserves coverage", () => { + const managed = session("11111111-1111-4111-8111-111111111111", { usage: usage(21) }); + const unsupported = session("22222222-2222-4222-8222-222222222222"); + const observations: RuntimeObservation[] = [{ + id: managed.id, + object: "agent.runtime_observation", + session_id: managed.id, + environment_id: "33333333-3333-4333-8333-333333333333", + mode: "openai_hosted", + provider_type: "docker", + instance: { kind: "managed_allocation", allocation_id: "44444444-4444-4444-8444-444444444444", device_id: null, connection_generation: null }, + status: "observed", + reason: null, + allocation_created_at: 100, + resolved_at: 220, + observed_at: 210, + started_at: 150, + cpu: { usage_seconds_total: 3.5, capacity_cores: 2, usage_cores: null, utilization_ratio: null }, + memory: { usage_bytes: 512, limit_bytes: 2048 }, + }, { + id: unsupported.id, + object: "agent.runtime_observation", + session_id: unsupported.id, + environment_id: null, + mode: "none", + provider_type: null, + instance: { kind: "none", allocation_id: null, device_id: null, connection_generation: null }, + status: "unsupported", + reason: "runtime_mode_not_observable", + allocation_created_at: null, + resolved_at: 225, + observed_at: null, + started_at: null, + cpu: null, + memory: null, + }]; + const model = buildRuntimeDashboardModel([managed, unsupported], observations); + + expect(model.summary).toMatchObject({ + sessionCount: 2, + managedRuntimeCount: 1, + observedRuntimeCount: 1, + cpuUsageSecondsTotal: 3.5, + cpuCapacityCores: 2, + cpuCoverageCount: 1, + memoryUsageBytes: 512, + memoryLimitBytes: 2048, + memoryCoverageCount: 1, + totalTokens: 21, + tokenCoverageCount: 1, + oldestResolvedAt: 220, + newestResolvedAt: 225, + }); + expect(model.rows[0]?.computeUptimeSeconds).toBe(60); + expect(model.rows[0]?.allocationAgeSeconds).toBe(120); + expect(runtimeObservationStatusLabel(observations[0]!)).toBe("Observed"); + expect(formatDashboardBytes(2048)).toBe("2.00 KiB"); + expect(formatDashboardDuration(90)).toBe("1m 30s"); + }); + + it("does not infer a released allocation lifetime from the current resolution time", () => { + const stopped = session("11111111-1111-4111-8111-111111111111"); + const observation: RuntimeObservation = { + id: stopped.id, + object: "agent.runtime_observation", + session_id: stopped.id, + environment_id: "33333333-3333-4333-8333-333333333333", + mode: "openai_hosted", + provider_type: "docker", + instance: { + kind: "managed_allocation", + allocation_id: "44444444-4444-4444-8444-444444444444", + device_id: null, + connection_generation: null, + }, + status: "unavailable", + reason: "runtime_not_running", + allocation_created_at: 100, + resolved_at: 10_000, + observed_at: null, + started_at: null, + cpu: null, + memory: null, + }; + + expect(buildRuntimeDashboardModel([stopped], [observation]).rows[0]?.allocationAgeSeconds).toBeNull(); + }); }); diff --git a/apps/web/src/features/dashboard/dashboard-model.ts b/apps/web/src/features/dashboard/dashboard-model.ts index 2239bce16..89e36e2a4 100644 --- a/apps/web/src/features/dashboard/dashboard-model.ts +++ b/apps/web/src/features/dashboard/dashboard-model.ts @@ -1,5 +1,6 @@ import type { AgentSession, + RuntimeObservation, SavedAgent, SessionStatus, TokenUsage, @@ -45,6 +46,35 @@ export interface DashboardSnapshot { recentSessions: DashboardSessionRow[]; } +export interface RuntimeDashboardRow { + session: DashboardSessionRow; + observation: RuntimeObservation; + computeUptimeSeconds: number | null; + allocationAgeSeconds: number | null; +} + +export interface RuntimeDashboardSummary { + sessionCount: number; + managedRuntimeCount: number; + observedRuntimeCount: number; + unavailableRuntimeCount: number; + cpuUsageSecondsTotal: number | null; + cpuCapacityCores: number | null; + cpuCoverageCount: number; + memoryUsageBytes: number | null; + memoryLimitBytes: number | null; + memoryCoverageCount: number; + totalTokens: number | null; + tokenCoverageCount: number; + oldestResolvedAt: number | null; + newestResolvedAt: number | null; +} + +export interface RuntimeDashboardModel { + summary: RuntimeDashboardSummary; + rows: RuntimeDashboardRow[]; +} + const sessionStatuses = new Set([ "idle", "in_progress", @@ -229,3 +259,185 @@ export function formatDashboardTimestamp(value: number | null): string { if (seconds === null) return "Unknown"; return `${new Date(seconds * 1_000).toISOString().slice(0, 16).replace("T", " ")} UTC`; } + +function safeFiniteNonNegative(value: unknown): number | null { + return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; +} + +function safeAdd(left: number, right: number): number | null { + const value = left + right; + return Number.isSafeInteger(left) && Number.isSafeInteger(right) && Number.isSafeInteger(value) ? value : null; +} + +function elapsedSeconds(start: number | null, end: number | null): number | null { + if (start === null || end === null || end < start) return null; + return end - start; +} + +export function buildRuntimeDashboardModel( + sessions: readonly AgentSession[], + observations: readonly RuntimeObservation[], +): RuntimeDashboardModel { + const sessionsById = new Map(sessions.map((session) => [session.id, session])); + const rows: RuntimeDashboardRow[] = []; + let managedRuntimeCount = 0; + let observedRuntimeCount = 0; + let unavailableRuntimeCount = 0; + let cpuUsageSecondsTotal = 0; + let cpuUsageKnown = false; + let cpuUsageSafe = true; + let cpuCapacityCores = 0; + let cpuCapacityKnown = false; + let cpuCapacitySafe = true; + let cpuCoverageCount = 0; + let memoryUsageBytes = 0; + let memoryUsageKnown = false; + let memoryUsageSafe = true; + let memoryLimitBytes = 0; + let memoryLimitKnown = false; + let memoryLimitSafe = true; + let memoryCoverageCount = 0; + let totalTokens = 0; + let tokensKnown = false; + let tokensSafe = true; + let tokenCoverageCount = 0; + let oldestResolvedAt: number | null = null; + let newestResolvedAt: number | null = null; + + for (const observation of observations) { + const session = sessionsById.get(observation.session_id); + if (!session) continue; + const sessionRow = toSessionRow(session); + const resolvedAt = canonicalTimestamp(observation.resolved_at); + if (resolvedAt !== null) { + oldestResolvedAt = oldestResolvedAt === null ? resolvedAt : Math.min(oldestResolvedAt, resolvedAt); + newestResolvedAt = newestResolvedAt === null ? resolvedAt : Math.max(newestResolvedAt, resolvedAt); + } + if (observation.mode === "openai_hosted") managedRuntimeCount += 1; + if (observation.status === "observed") { + observedRuntimeCount += 1; + const cpuUsage = safeFiniteNonNegative(observation.cpu?.usage_seconds_total); + const cpuCapacity = safeFiniteNonNegative(observation.cpu?.capacity_cores); + if (cpuUsage !== null) { + const next = cpuUsageSecondsTotal + cpuUsage; + if (Number.isFinite(next)) { + cpuUsageSecondsTotal = next; + cpuUsageKnown = true; + } else cpuUsageSafe = false; + } + if (cpuCapacity !== null) { + const next = cpuCapacityCores + cpuCapacity; + if (Number.isFinite(next)) { + cpuCapacityCores = next; + cpuCapacityKnown = true; + } else cpuCapacitySafe = false; + } + if (cpuUsage !== null || cpuCapacity !== null) cpuCoverageCount += 1; + + const memoryUsage = safeNonNegativeInteger(observation.memory?.usage_bytes); + const memoryLimit = safeNonNegativeInteger(observation.memory?.limit_bytes); + if (memoryUsage !== null) { + const next = safeAdd(memoryUsageBytes, memoryUsage); + if (next !== null) { + memoryUsageBytes = next; + memoryUsageKnown = true; + } else memoryUsageSafe = false; + } + if (memoryLimit !== null) { + const next = safeAdd(memoryLimitBytes, memoryLimit); + if (next !== null) { + memoryLimitBytes = next; + memoryLimitKnown = true; + } else memoryLimitSafe = false; + } + if (memoryUsage !== null || memoryLimit !== null) memoryCoverageCount += 1; + } else if (observation.status === "unavailable") { + unavailableRuntimeCount += 1; + } + + if (sessionRow.totalTokens !== null) { + const next = safeAdd(totalTokens, sessionRow.totalTokens); + if (next !== null) { + totalTokens = next; + tokensKnown = true; + } else tokensSafe = false; + tokenCoverageCount += 1; + } + rows.push({ + session: sessionRow, + observation, + computeUptimeSeconds: observation.status === "observed" + ? elapsedSeconds(canonicalTimestamp(observation.started_at), canonicalTimestamp(observation.observed_at)) + : null, + allocationAgeSeconds: observation.mode === "openai_hosted" && observation.reason !== "runtime_not_running" + ? elapsedSeconds(canonicalTimestamp(observation.allocation_created_at), resolvedAt) + : null, + }); + } + + const statusOrder = { observed: 0, unavailable: 1, unsupported: 2 } as const; + rows.sort((left, right) => ( + statusOrder[left.observation.status] - statusOrder[right.observation.status] + || compareSessionRows(left.session, right.session) + )); + return { + summary: { + sessionCount: rows.length, + managedRuntimeCount, + observedRuntimeCount, + unavailableRuntimeCount, + cpuUsageSecondsTotal: cpuUsageKnown && cpuUsageSafe ? cpuUsageSecondsTotal : null, + cpuCapacityCores: cpuCapacityKnown && cpuCapacitySafe ? cpuCapacityCores : null, + cpuCoverageCount, + memoryUsageBytes: memoryUsageKnown && memoryUsageSafe ? memoryUsageBytes : null, + memoryLimitBytes: memoryLimitKnown && memoryLimitSafe ? memoryLimitBytes : null, + memoryCoverageCount, + totalTokens: tokensKnown && tokensSafe ? totalTokens : null, + tokenCoverageCount, + oldestResolvedAt, + newestResolvedAt, + }, + rows, + }; +} + +export function formatDashboardBytes(value: number | null): string { + if (value === null) return "Unavailable"; + const units = ["B", "KiB", "MiB", "GiB", "TiB"]; + let amount = value; + let index = 0; + while (amount >= 1024 && index < units.length - 1) { + amount /= 1024; + index += 1; + } + const digits = amount >= 100 || index === 0 ? 0 : amount >= 10 ? 1 : 2; + return `${amount.toFixed(digits)} ${units[index]}`; +} + +export function formatDashboardDuration(value: number | null): string { + if (value === null || !Number.isFinite(value) || value < 0) return "Unavailable"; + const seconds = Math.floor(value); + if (seconds < 60) return `${seconds}s`; + const minutes = Math.floor(seconds / 60); + if (minutes < 60) return `${minutes}m ${seconds % 60}s`; + const hours = Math.floor(minutes / 60); + if (hours < 24) return `${hours}h ${minutes % 60}m`; + return `${Math.floor(hours / 24)}d ${hours % 24}h`; +} + +export function formatDashboardTokens(value: number | null): string { + if (value === null) return "Unavailable"; + return value.toLocaleString("en-US"); +} + +export function runtimeObservationStatusLabel(observation: RuntimeObservation): string { + if (observation.status === "observed") return "Observed"; + if (observation.status === "unsupported") return "Unsupported"; + switch (observation.reason) { + case "allocation_pending": return "Allocation pending"; + case "runtime_not_running": return "Not running"; + case "source_not_configured": return "Source unavailable"; + case "sample_timeout": return "Sample timeout"; + case "sample_unavailable": return "Sample unavailable"; + } +} diff --git a/apps/web/src/features/dashboard/runtime-snapshot.test.ts b/apps/web/src/features/dashboard/runtime-snapshot.test.ts new file mode 100644 index 000000000..4e852e59a --- /dev/null +++ b/apps/web/src/features/dashboard/runtime-snapshot.test.ts @@ -0,0 +1,57 @@ +import { describe, expect, it } from "vitest"; + +import type { + AgentCore, + AgentSession, + RuntimeObservation, +} from "@agents-core-web/agents-client"; + +import { + loadRuntimeDashboardSnapshot, + RuntimeSnapshotIncompleteError, +} from "./runtime-snapshot"; + +const session = { id: "11111111-1111-4111-8111-111111111111" } as AgentSession; +const observation = { + id: session.id, + session_id: session.id, +} as RuntimeObservation; + +function core( + sessions: AgentSession[], + observations: RuntimeObservation[], +): Pick { + return { + listSessions: async () => ({ object: "list", data: sessions, has_more: false, first_id: session.id, last_id: session.id }), + listRuntimeObservations: async () => ({ object: "list", data: observations, has_more: false, first_id: session.id, last_id: session.id }), + }; +} + +describe("Runtime Dashboard snapshot coordination", () => { + it("publishes only exact Session and observation identity sets", async () => { + const value = await loadRuntimeDashboardSnapshot(core([session], [observation]), () => 4); + expect(value?.sessions).toEqual([session]); + expect(value?.observations).toEqual([observation]); + + await expect(loadRuntimeDashboardSnapshot(core([session], []), () => 4)).rejects.toBeInstanceOf( + RuntimeSnapshotIncompleteError, + ); + }); + + it("discards a candidate when local Session state changes during collection", async () => { + let revision = 1; + const changing = core([session], [observation]); + changing.listRuntimeObservations = async () => { + revision += 1; + return { object: "list", data: [observation], has_more: false, first_id: session.id, last_id: session.id }; + }; + + await expect(loadRuntimeDashboardSnapshot(changing, () => revision)).resolves.toBeNull(); + }); + + it("fails closed when the target budget is exceeded", async () => { + await expect(loadRuntimeDashboardSnapshot(core([session], [observation]), () => 1, undefined, 0)).rejects.toThrow( + "target budget", + ); + }); +}); diff --git a/apps/web/src/features/dashboard/runtime-snapshot.ts b/apps/web/src/features/dashboard/runtime-snapshot.ts new file mode 100644 index 000000000..79d794406 --- /dev/null +++ b/apps/web/src/features/dashboard/runtime-snapshot.ts @@ -0,0 +1,59 @@ +import type { + AgentCore, + AgentSession, + RuntimeObservation, +} from "@agents-core-web/agents-client"; + +import { listAllCollectionPages } from "../../lib/collection-pagination"; + +export const RUNTIME_SNAPSHOT_TARGET_LIMIT = 10_000; +export const RUNTIME_SNAPSHOT_TIMEOUT_MS = 15_000; +export const RUNTIME_SNAPSHOT_REFRESH_MS = 30_000; + +export interface RuntimeDashboardSnapshot { + sessions: AgentSession[]; + observations: RuntimeObservation[]; + loadedAt: number; +} + +export class RuntimeSnapshotIncompleteError extends Error { + constructor(message: string) { + super(message); + this.name = "RuntimeSnapshotIncompleteError"; + } +} + +function identitySet(values: readonly { id: string }[]): Set { + return new Set(values.map((value) => value.id)); +} + +function setsEqual(left: ReadonlySet, right: ReadonlySet): boolean { + if (left.size !== right.size) return false; + for (const value of left) if (!right.has(value)) return false; + return true; +} + +export async function loadRuntimeDashboardSnapshot( + core: Pick, + readSessionRevision: () => number, + signal?: AbortSignal, + targetLimit = RUNTIME_SNAPSHOT_TARGET_LIMIT, +): Promise { + const revision = readSessionRevision(); + const [sessions, observations] = await Promise.all([ + listAllCollectionPages((options) => core.listSessions(options), signal), + listAllCollectionPages((options) => core.listRuntimeObservations(options), signal), + ]); + signal?.throwIfAborted(); + + if (revision !== readSessionRevision()) return null; + if (sessions.length > targetLimit || observations.length > targetLimit) { + throw new RuntimeSnapshotIncompleteError("Runtime snapshot exceeded the Web target budget."); + } + if (!setsEqual(identitySet(sessions), identitySet(observations))) { + throw new RuntimeSnapshotIncompleteError( + "Sessions changed while Runtime observations were loading. The previous complete snapshot was retained.", + ); + } + return { sessions, observations, loadedAt: Date.now() }; +} diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 34394ffd3..4d6345a51 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -1,9 +1,9 @@ -# Runtime observation API proposal +# Runtime observation API -Status: Phase 2 implemented. The current-snapshot routes, strict +Status: Phase 2 and initial Core Web consumption implemented. The current-snapshot routes, strict `packages/agents-client` projection, and generated `openapi.yaml` contract are -implemented. Core Web integration, historical queries, and lifecycle controls -remain outside this phase. +implemented and consumed by the Dashboard through complete Session/observation +identity joins. Historical queries and lifecycle controls remain outside this phase. This is an Agents Core extension, not an upstream OpenAI Agents resource. The implementation must record that status in the coverage ledger and generated diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index dd9dfa16d..d49bf3ed5 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -1,8 +1,9 @@ # Runtime observability and Dashboard design -Status: Phase 1 provider abstraction/Docker sampling and Phase 2 current-snapshot -API/client contract are implemented. Core Web integration, history backend, -additional providers, and lifecycle automation described below are not implemented. +Status: Phase 1 provider abstraction/Docker sampling, Phase 2 current-snapshot +API/client contract, and the initial Core Web current-snapshot Dashboard are +implemented. The history backend, additional providers, and lifecycle automation +described below are not implemented. ## 1. Problem statement @@ -267,7 +268,7 @@ paths, and raw labels are never displayed. ### 11.5 Web implementation shape -The Web change belongs in Core Web, not the Core service repository. It uses +The Web implementation belongs in Core Web, not the Core service layer. It uses `packages/agents-client` as the only Runtime-observation transport and keeps four seams separate: @@ -284,10 +285,12 @@ seams separate: Session detail surface. Trend components are absent unless a later history capability and contract are configured. -The initial refresh cadence is an operator-configured value, not an API guarantee. -Web pauses periodic reads when hidden, refreshes when visibility returns, and adds -jitter so multiple browsers do not synchronize. Filtering is local to the last -complete snapshot and never changes tenant authorization or provider selection. +The initial implementation uses a 30-second Web cadence plus up to five seconds +of jitter and a 15-second whole-refresh budget. These are Web configuration, not +API guarantees. Web pauses periodic reads when hidden, refreshes when visibility +returns, and adds jitter so multiple browsers do not synchronize. Filtering is +local to the last complete snapshot and never changes tenant authorization or +provider selection. ## 12. Token usage boundary diff --git a/docs/web/README.md b/docs/web/README.md index 3703cddeb..c94068ae1 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -11,7 +11,8 @@ credentials or execution into the browser. ## What you can do -- **Operate from one Dashboard** — see loaded Agents, active Sessions, work that needs +- **Operate from one Dashboard** — see loaded Agents, active Sessions, current Runtime + CPU/memory evidence, compute uptime, reported token coverage, work that needs attention, recent activity, and the two common create flows. - **Build reusable Agents** — start from a blank Agent or a practical template, then configure its model, instructions, text behavior, Functions, and HTTP MCP servers. @@ -29,8 +30,10 @@ credentials or execution into the browser. ### Dashboard Dashboard is the starting point. It summarizes the current Agent and Session results, -highlights Sessions that need attention, and links directly to Agent creation or a new -Session. +loads a complete tenant-scoped Runtime observation snapshot, shows current Docker +resource evidence and coverage without inventing missing values, highlights Sessions +that need attention, and links directly to Agent creation or a new Session. Historical +charts remain absent until an operator configures a separate history capability. ### Agents @@ -61,7 +64,7 @@ service is not mistaken for a ready model execution path. | Area | User experience | | --- | --- | -| Dashboard | Agent and Session overview, attention queue, recent activity, quick actions | +| Dashboard | Agent and Session overview, current Runtime CPU/memory/uptime and token coverage, attention queue, recent activity, quick actions | | Agents | Create, search, inspect, edit, delete, use templates, and start Sessions | | Sessions | Durable conversation history, Agent filtering, live events, cancellation, retry and continuation | | Trace | Turn history, usage when reported by Core, command output, Function and patch activity | diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index f444beddb..a1afb5249 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -10,8 +10,8 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 ## 可以做什么 -- **通过 Dashboard 统一管理**:查看 Agent、活跃 Session、需要关注的工作、最近活动 - 和常用创建入口。 +- **通过 Dashboard 统一管理**:查看 Agent、活跃 Session、Runtime 当前 CPU/内存 + 证据、计算运行时长、Token 覆盖率、需要关注的工作、最近活动和常用创建入口。 - **创建可复用 Agent**:从空白配置或实用模板开始,设置模型、指令、文本行为、 Function 和 HTTP MCP 服务。 - **运行持久化对话**:创建 Session、发送消息、查看实时事件、重新打开历史工作、 @@ -27,8 +27,10 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 ### Dashboard -Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果、需要关注的 Session, -并可直接进入创建 Agent 或启动 Session 的流程。 +Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载完整的租户级 Runtime +观测快照,并在不把缺失值伪装成 0 的前提下展示当前 Docker 资源和数据覆盖率。页面也 +展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。在运维方配置 +独立历史能力之前,页面不会伪造历史趋势图。 ### Agents @@ -56,7 +58,7 @@ System 展示当前 Core 对 Web 暴露的能力,并区分 API 访问、Vault | 区域 | 用户可以完成的工作 | | --- | --- | -| Dashboard | 查看 Agent/Session 概览、关注队列、最近活动和快捷入口 | +| Dashboard | 查看 Agent/Session 概览、Runtime 当前 CPU/内存/运行时长与 Token 覆盖率、关注队列、最近活动和快捷入口 | | Agents | 创建、搜索、查看、编辑、删除、使用模板并启动 Session | | Sessions | 持久化对话、按 Agent 筛选、实时事件、取消、重试和继续执行 | | Trace | 查看 Turn 历史、Core 报告的 Usage、命令输出、Function 和 Patch 活动 | From 745f51adff0eb68e823353bb50b87ff8251be665 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 21:49:03 +0800 Subject: [PATCH 05/51] Extend Runtime observability dashboard --- apps/web/package.json | 1 + .../src/features/dashboard/DashboardView.css | 467 +++++++++++++++--- .../features/dashboard/DashboardView.test.tsx | 19 +- .../src/features/dashboard/DashboardView.tsx | 148 +----- .../dashboard/RuntimeObservabilityContent.tsx | 349 +++++++++++++ .../dashboard/dashboard-model.test.ts | 48 ++ .../src/features/dashboard/dashboard-model.ts | 9 +- .../runtime-observability-design.md | 11 +- docs/web/README.md | 6 +- docs/web/README.zh-CN.md | 5 +- pnpm-lock.yaml | 22 + 11 files changed, 852 insertions(+), 233 deletions(-) create mode 100644 apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx diff --git a/apps/web/package.json b/apps/web/package.json index 98bf865e2..2a7f1da60 100644 --- a/apps/web/package.json +++ b/apps/web/package.json @@ -11,6 +11,7 @@ }, "dependencies": { "@agents-core-web/agents-client": "workspace:*", + "@tanstack/react-table": "^8.21.3", "lucide-react": "^1.22.0", "react": "^19.2.7", "react-dom": "^19.2.7", diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index b76218653..13778f0fb 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -683,74 +683,302 @@ white-space: nowrap; } -.dashboard-runtime-ledger { - min-width: 960px; +.dashboard-runtime-insights { + display: grid; + grid-template-columns: minmax(0, 1.15fr) minmax(0, .85fr); + border-bottom: 1px solid var(--line); } -.dashboard-runtime-panel:has(.dashboard-runtime-ledger) { - overflow-x: auto; +.dashboard-runtime-insight { + min-width: 0; + padding: 14px; } -.dashboard-runtime-header, -.dashboard-runtime-row { - display: grid; - grid-template-columns: minmax(190px, 1.45fr) minmax(140px, 1fr) minmax(130px, .9fr) minmax(150px, 1fr) minmax(140px, .9fr) minmax(110px, .75fr); - align-items: center; - column-gap: 14px; +.dashboard-runtime-insight + .dashboard-runtime-insight { + border-left: 1px solid var(--line); +} + +.dashboard-runtime-insight > header, +.dashboard-runtime-targets > header { + display: flex; + align-items: flex-start; + justify-content: space-between; + gap: 12px; +} + +.dashboard-runtime-insight h3, +.dashboard-runtime-targets h3, +.dashboard-runtime-insight p, +.dashboard-runtime-targets p { + margin: 0; +} + +.dashboard-runtime-insight h3, +.dashboard-runtime-targets h3 { + font-size: 12px; + font-weight: 600; + line-height: 17px; } -.dashboard-runtime-header { - min-height: 34px; - padding: 7px 14px; +.dashboard-runtime-insight header p, +.dashboard-runtime-targets header p, +.dashboard-runtime-insight header > span, +.dashboard-runtime-targets header > span { color: var(--fg-muted); font-size: 9px; - font-weight: 600; - letter-spacing: .04em; line-height: 13px; +} + +.dashboard-runtime-insight header > span, +.dashboard-runtime-targets header > span { + flex: 0 0 auto; + font-family: var(--font-mono); + font-variant-numeric: tabular-nums; +} + +.dashboard-runtime-health-grid { + display: grid; + margin-top: 12px; + grid-template-columns: repeat(2, minmax(0, 1fr)); + gap: 7px; +} + +.dashboard-runtime-health { + display: grid; + min-width: 0; + padding: 8px 9px 8px 19px; + position: relative; + gap: 1px; + color: var(--fg); + text-align: left; background: var(--surface-subtle); - border-bottom: 1px solid var(--line); - text-transform: uppercase; + border: 1px solid var(--line); + border-radius: 6px; } -.dashboard-runtime-entry + .dashboard-runtime-entry { - border-top: 1px solid var(--line); +.dashboard-runtime-health::before { + width: 6px; + height: 6px; + position: absolute; + top: 12px; + left: 8px; + content: ""; + background: var(--fg-muted); + border-radius: 50%; } -.dashboard-runtime-row { - min-height: 70px; - padding: 11px 14px; - transition: background-color 150ms var(--ease-settle); +.dashboard-runtime-health-observed::before { + background: var(--success); +} + +.dashboard-runtime-health-unavailable::before { + background: var(--warning); } -.dashboard-runtime-row:hover, -.dashboard-runtime-row:focus-within { +.dashboard-runtime-health:hover, +.dashboard-runtime-health:focus-visible { background: var(--hover); + border-color: color-mix(in srgb, var(--accent) 35%, var(--line)); +} + +.dashboard-runtime-health strong, +.dashboard-runtime-health span { + overflow: hidden; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-health strong { + font-size: 10px; + font-weight: 550; + line-height: 14px; +} + +.dashboard-runtime-health span, +.dashboard-runtime-health-more { + color: var(--fg-muted); + font-size: 9px; + line-height: 13px; +} + +.dashboard-runtime-health-more { + align-self: center; + padding-left: 4px; +} + +.dashboard-runtime-coverage-grid { + display: grid; + margin-top: 12px; + grid-template-columns: repeat(2, minmax(0, 1fr)); + gap: 7px; } -.dashboard-runtime-row > span { +.dashboard-runtime-coverage-grid > div { display: grid; min-width: 0; - gap: 3px; + padding: 9px 10px; + gap: 1px; + background: var(--surface-subtle); + border: 1px solid var(--line); + border-radius: 6px; } -.dashboard-runtime-row small { - overflow: hidden; +.dashboard-runtime-coverage-grid strong { + font-family: var(--font-mono); + font-size: 10px; + font-variant-numeric: tabular-nums; + font-weight: 600; + line-height: 14px; +} + +.dashboard-runtime-coverage-grid span { color: var(--fg-muted); font-size: 9px; line-height: 13px; - text-overflow: ellipsis; - white-space: nowrap; } -.dashboard-runtime-identity button { +.dashboard-runtime-coverage-warning strong { + color: var(--warning); +} + +.dashboard-runtime-coverage-unsupported strong { + color: var(--fg-muted); +} + +.dashboard-runtime-targets { + min-width: 0; +} + +.dashboard-runtime-targets > header { + min-height: 58px; + align-items: center; + padding: 11px 14px; + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-toolbar { + display: flex; + min-height: 48px; + align-items: center; + padding: 8px 14px; + gap: 8px; + background: var(--surface-subtle); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-toolbar label { + display: flex; + align-items: center; + gap: 6px; + color: var(--fg-muted); + font-size: 9px; +} + +.dashboard-runtime-toolbar input, +.dashboard-runtime-toolbar select { + height: 30px; + color: var(--fg); + font: inherit; + background: var(--surface); + border: 1px solid var(--line); + border-radius: 5px; +} + +.dashboard-runtime-search { + width: min(360px, 100%); + height: 30px; + margin-right: auto; + padding: 0 8px; + background: var(--surface); + border: 1px solid var(--line); + border-radius: 5px; +} + +.dashboard-runtime-search input { + width: 100%; + min-width: 0; + height: auto; + padding: 0; + background: transparent; + border: 0; + outline: 0; +} + +.dashboard-runtime-toolbar select { + min-width: 112px; + padding: 0 24px 0 8px; +} + +.dashboard-runtime-table-scroll { + overflow-x: auto; +} + +.dashboard-runtime-table { + width: 100%; + min-width: 1080px; + border-collapse: collapse; + font-size: 10px; + line-height: 14px; +} + +.dashboard-runtime-table th { + height: 34px; + padding: 0 10px; + color: var(--fg-muted); + font-size: 9px; + font-weight: 600; + text-align: left; + background: var(--surface-subtle); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-table th button { + display: inline-flex; + align-items: center; + padding: 0; + gap: 4px; + color: inherit; + font: inherit; + background: transparent; + border: 0; +} + +.dashboard-runtime-table td { + height: 68px; + padding: 9px 10px; + vertical-align: middle; + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-table tbody tr:hover { + background: var(--hover); +} + +.dashboard-runtime-table th:first-child, +.dashboard-runtime-table td:first-child { + width: 220px; + padding-left: 14px; +} + +.dashboard-runtime-target-identity, +.dashboard-runtime-table-value { + display: grid; + min-width: 0; + gap: 2px; +} + +.dashboard-runtime-target-identity { + position: relative; +} + +.dashboard-runtime-target-identity > button { width: fit-content; max-width: 100%; padding: 0; overflow: hidden; color: var(--fg); - font-size: 12px; + font: inherit; font-weight: 600; - line-height: 17px; text-align: left; text-overflow: ellipsis; white-space: nowrap; @@ -758,11 +986,63 @@ border: 0; } -.dashboard-runtime-identity button:hover, -.dashboard-runtime-identity button:focus-visible { +.dashboard-runtime-target-identity > button:hover, +.dashboard-runtime-target-identity > button:focus-visible { color: var(--accent); } +.dashboard-runtime-target-identity > small, +.dashboard-runtime-table-value small { + overflow: hidden; + color: var(--fg-muted); + font-size: 9px; + line-height: 13px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-target-identity details { + color: var(--fg-muted); + font-size: 9px; +} + +.dashboard-runtime-target-identity summary { + width: fit-content; + cursor: pointer; +} + +.dashboard-runtime-target-identity dl { + display: grid; + width: 320px; + margin: 6px 0 0; + padding: 8px; + position: absolute; + z-index: 2; + gap: 4px; + background: var(--surface); + border: 1px solid var(--line); + border-radius: 6px; + box-shadow: var(--shadow-control); +} + +.dashboard-runtime-target-identity dl > div { + display: grid; + grid-template-columns: 72px minmax(0, 1fr); +} + +.dashboard-runtime-target-identity dt, +.dashboard-runtime-target-identity dd { + margin: 0; + overflow: hidden; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-target-identity dd { + color: var(--fg); + font-family: var(--font-mono); +} + .dashboard-runtime-status { display: inline-flex; width: fit-content; @@ -793,7 +1073,11 @@ background: var(--warning); } -.dashboard-runtime-value strong { +.dashboard-runtime-status-unsupported { + color: var(--fg-muted); +} + +.dashboard-runtime-table-value strong { overflow: hidden; font-family: var(--font-mono); font-size: 12px; @@ -822,64 +1106,49 @@ border-radius: inherit; } -.dashboard-runtime-detail { - padding: 0 14px 10px; +.dashboard-runtime-no-results { + margin: 0; + padding: 20px 14px; color: var(--fg-muted); font-size: 10px; + text-align: center; + border-bottom: 1px solid var(--line); } -.dashboard-runtime-detail summary { - width: fit-content; - padding: 4px 0; - cursor: pointer; - user-select: none; -} - -.dashboard-runtime-detail[open] { - padding-top: 8px; - background: var(--surface-subtle); - border-top: 1px dashed var(--line); -} - -.dashboard-runtime-detail dl { - display: grid; - margin: 8px 0; - grid-template-columns: repeat(3, minmax(0, 1fr)); - gap: 8px 14px; -} - -.dashboard-runtime-detail dl > div { - min-width: 0; -} - -.dashboard-runtime-detail dt { +.dashboard-runtime-pagination { + display: flex; + min-height: 46px; + align-items: center; + justify-content: space-between; + padding: 8px 14px; color: var(--fg-muted); + font-family: var(--font-mono); font-size: 9px; - text-transform: uppercase; } -.dashboard-runtime-detail dd { - margin: 2px 0 0; - overflow: hidden; - color: var(--fg); - font-family: var(--font-mono); - font-size: 10px; - text-overflow: ellipsis; - white-space: nowrap; +.dashboard-runtime-pagination > div { + display: flex; + gap: 6px; } -.dashboard-runtime-detail button { +.dashboard-runtime-pagination button { display: inline-flex; + min-height: 28px; align-items: center; - padding: 4px 7px; + padding: 4px 8px; gap: 4px; - color: var(--accent); - font-size: 10px; - background: transparent; - border: 1px solid color-mix(in srgb, var(--accent) 25%, var(--line)); + color: var(--fg); + font: inherit; + background: var(--surface-subtle); + border: 1px solid var(--line); border-radius: 5px; } +.dashboard-runtime-pagination button:disabled { + cursor: not-allowed; + opacity: .45; +} + @media (max-width: 980px) { .dashboard-primary-grid { grid-template-columns: 1fr; @@ -896,6 +1165,15 @@ .dashboard-runtime-metric:nth-child(n + 3) { border-top: 1px solid var(--line); } + + .dashboard-runtime-insights { + grid-template-columns: 1fr; + } + + .dashboard-runtime-insight + .dashboard-runtime-insight { + border-top: 1px solid var(--line); + border-left: 0; + } } @media (max-width: 760px) { @@ -949,6 +1227,26 @@ .dashboard-metric:nth-child(n + 3) { border-top: 1px solid var(--line); } + + .dashboard-runtime-toolbar { + align-items: stretch; + flex-wrap: wrap; + } + + .dashboard-runtime-search { + width: 100%; + flex-basis: 100%; + margin-right: 0; + } + + .dashboard-runtime-toolbar label:not(.dashboard-runtime-search) { + flex: 1; + } + + .dashboard-runtime-toolbar select { + min-width: 0; + flex: 1; + } } @media (max-width: 520px) { @@ -990,6 +1288,17 @@ border-left: 0; } + .dashboard-runtime-health-grid, + .dashboard-runtime-coverage-grid { + grid-template-columns: 1fr; + } + + .dashboard-runtime-targets > header { + align-items: flex-start; + flex-direction: column; + gap: 4px; + } + .dashboard-panel header p { white-space: normal; } diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 5d70620aa..c0887eba7 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -245,10 +245,25 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("CPU time / capacity"); expect(html).toContain("1m 13s / 2 cores"); expect(html).toContain("512 MiB / 2.00 GiB"); - expect(html).toContain("Current Runtime observations"); + expect(html).toContain("Runtime health"); + expect(html).toContain("Observed · 2023-11-14 22:14 UTC"); + expect(html).not.toContain("Observed · 10s"); + expect(html).toContain("Usage coverage"); + expect(html).toContain("CPU 1/1"); + expect(html).toContain("Memory 1/1"); + expect(html).toContain("Tokens 1/1"); + expect(html).toContain("Unsupported 0"); + expect(html).toContain("Runtime targets"); + expect(html).toContain("Search Runtime targets"); + expect(html).toContain("All statuses"); + expect(html).toContain("All modes"); + expect(html).toContain(''); + expect(html).toContain("CPU time"); expect(html).toContain("Managed research"); - expect(html).toContain("Identity and sample details"); + expect(html).toContain("Identity"); + expect(html).toContain("Unknown remains unknown, never zero"); expect(html).not.toContain("CPU %"); + expect(html).not.toContain("CPU now"); expect(html).not.toContain("historical chart"); }); diff --git a/apps/web/src/features/dashboard/DashboardView.tsx b/apps/web/src/features/dashboard/DashboardView.tsx index 1ab5f2d87..4fd6cc68c 100644 --- a/apps/web/src/features/dashboard/DashboardView.tsx +++ b/apps/web/src/features/dashboard/DashboardView.tsx @@ -2,13 +2,9 @@ import { AlertTriangle, ArrowRight, Bot, - Cpu, - Gauge, - MemoryStick, MessageSquare, RefreshCw, Rows3, - Server, } from "lucide-react"; import { useMemo, type ReactNode } from "react"; @@ -21,14 +17,11 @@ import { buildRuntimeDashboardModel, dashboardEnvironmentLabel, dashboardStatusLabel, - formatDashboardBytes, - formatDashboardDuration, - formatDashboardTokens, formatDashboardTimestamp, - runtimeObservationStatusLabel, type DashboardCollectionState, type DashboardSessionRow, } from "./dashboard-model"; +import { RuntimeObservabilityContent } from "./RuntimeObservabilityContent"; import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; import "./DashboardView.css"; @@ -54,109 +47,6 @@ export interface DashboardViewProps { onOpenSession: (sessionId: string) => void; } -function RuntimeMetric({ - icon, - label, - value, - detail, -}: { - icon: ReactNode; - label: string; - value: string; - detail: string; -}) { - return ( -
- - - {label} - {value} - {detail} - -
- ); -} - -function percent(usage: number | null | undefined, limit: number | null | undefined): number | null { - if (typeof usage !== "number" || typeof limit !== "number" || limit <= 0) return null; - return Math.min(100, Math.max(0, usage / limit * 100)); -} - -function RuntimeTable({ - snapshot, - onOpenSession, -}: { - snapshot: RuntimeDashboardSnapshot; - onOpenSession: (sessionId: string) => void; -}) { - const model = buildRuntimeDashboardModel(snapshot.sessions, snapshot.observations); - return ( -
-
- Session / Runtime - Observation - CPU - Memory - Uptime - Tokens -
- {model.rows.map((row) => { - const observation = row.observation; - const memoryPercent = observation.status === "observed" - ? percent(observation.memory?.usage_bytes, observation.memory?.limit_bytes) - : null; - return ( -
-
- - - {row.session.agentLabel} · {dashboardStatusLabel(row.session.status)} - - - - - {observation.mode === "openai_hosted" ? observation.provider_type ?? "Managed" : dashboardEnvironmentLabel(row.session.environmentProfile)} - - - {observation.status === "observed" ? formatDashboardDuration(observation.cpu?.usage_seconds_total ?? null) : "—"} - {observation.status === "observed" && typeof observation.cpu?.capacity_cores === "number" - ? `${observation.cpu.capacity_cores.toLocaleString("en-US")} cores capacity` - : "No current sample"} - - - {observation.status === "observed" ? formatDashboardBytes(observation.memory?.usage_bytes ?? null) : "—"} - {observation.status === "observed" ? `of ${formatDashboardBytes(observation.memory?.limit_bytes ?? null)}` : "No current sample"} - {memoryPercent !== null ? : null} - - - {formatDashboardDuration(row.computeUptimeSeconds)} - {row.allocationAgeSeconds === null ? "Allocation age unavailable" : `${formatDashboardDuration(row.allocationAgeSeconds)} allocation age`} - - - {formatDashboardTokens(row.session.totalTokens)} - {row.session.totalTokens === null ? "Usage not reported" : "Session reported"} - -
-
- Identity and sample details -
-
Session
{observation.session_id}
-
Environment
{observation.environment_id ?? "Not applicable"}
-
Allocation
{observation.instance.allocation_id ?? "Not available"}
-
Provider
{observation.provider_type ?? "Not available"}
-
Resolved
{formatDashboardTimestamp(observation.resolved_at)}
-
Observed
{formatDashboardTimestamp(observation.observed_at)}
-
- -
-
- ); - })} -
- ); -} - function collectionHasSnapshot(state: DashboardCollectionState, hasSnapshot: boolean): boolean { return state === "ready" || hasSnapshot; } @@ -514,38 +404,12 @@ export function DashboardView({

) : ( <> -
- } - label="Active Runtimes" - value={runtimeModel.summary.observedRuntimeCount.toLocaleString("en-US")} - detail={`${runtimeModel.summary.managedRuntimeCount} managed · ${runtimeModel.summary.unavailableRuntimeCount} unavailable`} - /> - } - label="CPU time / capacity" - value={runtimeModel.summary.cpuUsageSecondsTotal === null && runtimeModel.summary.cpuCapacityCores === null - ? "No current sample" - : `${formatDashboardDuration(runtimeModel.summary.cpuUsageSecondsTotal)} / ${runtimeModel.summary.cpuCapacityCores?.toLocaleString("en-US") ?? "—"} cores`} - detail={`${runtimeModel.summary.cpuCoverageCount}/${runtimeModel.summary.observedRuntimeCount} observed Runtimes report CPU`} - /> - } - label="Memory" - value={runtimeModel.summary.memoryUsageBytes === null && runtimeModel.summary.memoryLimitBytes === null - ? "No current sample" - : `${formatDashboardBytes(runtimeModel.summary.memoryUsageBytes)} / ${formatDashboardBytes(runtimeModel.summary.memoryLimitBytes)}`} - detail={`${runtimeModel.summary.memoryCoverageCount}/${runtimeModel.summary.observedRuntimeCount} observed Runtimes report memory`} - /> - } - label="Reported tokens" - value={runtimeModel.summary.totalTokens === null ? "Not reported" : formatDashboardTokens(runtimeModel.summary.totalTokens)} - detail={`${runtimeModel.summary.tokenCoverageCount}/${runtimeModel.summary.sessionCount} Sessions report usage`} - /> -
{runtimeModel.rows.length ? ( - + ) : (

No Session-owned Runtime contexts in this snapshot.

)} diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx new file mode 100644 index 000000000..c580856cc --- /dev/null +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -0,0 +1,349 @@ +import { + ChevronDown, + ChevronLeft, + ChevronRight, + ChevronsUpDown, + ChevronUp, + Cpu, + Gauge, + MemoryStick, + Search, + Server, +} from "lucide-react"; +import { useMemo, useState, type ReactNode } from "react"; +import { + flexRender, + getCoreRowModel, + getFilteredRowModel, + getPaginationRowModel, + getSortedRowModel, + useReactTable, + type ColumnDef, + type FilterFn, + type SortingState, +} from "@tanstack/react-table"; + +import { + buildRuntimeDashboardModel, + dashboardEnvironmentLabel, + dashboardStatusLabel, + formatDashboardBytes, + formatDashboardDuration, + formatDashboardTimestamp, + formatDashboardTokens, + runtimeObservationStatusLabel, + type RuntimeDashboardRow, +} from "./dashboard-model"; +import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; + +const PAGE_SIZE = 10; + +function RuntimeMetric({ + icon, + label, + value, + detail, +}: { + icon: ReactNode; + label: string; + value: string; + detail: string; +}) { + return ( +
+ + + {label} + {value} + {detail} + +
+ ); +} + +function percent(usage: number | null | undefined, limit: number | null | undefined): number | null { + if (typeof usage !== "number" || typeof limit !== "number" || limit <= 0) return null; + return Math.min(100, Math.max(0, usage / limit * 100)); +} + +function coverage(known: number, total: number): string { + if (total === 0) return "No observed Runtimes"; + return `${Math.round(known / total * 100)}% coverage`; +} + +function runtimeModeLabel(row: RuntimeDashboardRow): string { + if (row.observation.mode === "openai_hosted") { + const provider = row.observation.provider_type; + return provider ? `Managed ${provider === "docker" ? "Docker" : provider}` : "Managed"; + } + return dashboardEnvironmentLabel(row.session.environmentProfile); +} + +function observationTimestamp(row: RuntimeDashboardRow, loadedAt: number): number | null { + const timestamp = row.observation.status === "observed" + ? row.observation.observed_at + : row.observation.resolved_at; + if (!Number.isSafeInteger(timestamp) || timestamp < 0 || timestamp > Math.floor(loadedAt / 1_000)) return null; + return timestamp; +} + +function healthDetail(row: RuntimeDashboardRow, loadedAt: number): string { + const status = runtimeObservationStatusLabel(row.observation); + const timestamp = observationTimestamp(row, loadedAt); + return timestamp === null ? status : `${status} · ${formatDashboardTimestamp(timestamp)}`; +} + +function SortHeader({ + label, + sorted, + onClick, +}: { + label: string; + sorted: false | "asc" | "desc"; + onClick: (event: unknown) => void; +}) { + const Icon = sorted === "asc" ? ChevronUp : sorted === "desc" ? ChevronDown : ChevronsUpDown; + return ( + + ); +} + +const runtimeGlobalFilter: FilterFn = (row, _columnId, value) => { + const query = String(value).trim().toLocaleLowerCase(); + if (!query) return true; + const item = row.original; + return [ + item.session.title, + item.session.agentLabel, + item.observation.session_id, + item.observation.environment_id, + item.observation.instance.allocation_id, + item.observation.provider_type, + item.observation.status, + item.observation.reason, + runtimeModeLabel(item), + ].some((candidate) => typeof candidate === "string" && candidate.toLocaleLowerCase().includes(query)); +}; + +function RuntimeTargets({ + rows, + onOpenSession, +}: { + rows: RuntimeDashboardRow[]; + onOpenSession: (sessionId: string) => void; +}) { + const [sorting, setSorting] = useState([]); + const [globalFilter, setGlobalFilter] = useState(""); + const [statusFilter, setStatusFilter] = useState("all"); + const [modeFilter, setModeFilter] = useState("all"); + const filteredRows = useMemo(() => rows.filter((row) => ( + (statusFilter === "all" || row.observation.status === statusFilter) && + (modeFilter === "all" || row.observation.mode === modeFilter) + )), [modeFilter, rows, statusFilter]); + const columns = useMemo[]>(() => [{ + id: "session", + accessorFn: (row) => row.session.title, + header: ({ column }) => undefined)} />, + cell: ({ row }) => { + const item = row.original; + return ( +
+ + {item.session.agentLabel} +
+ Identity +
+
Session
{item.observation.session_id}
+
Environment
{item.observation.environment_id ?? "Not applicable"}
+
Allocation
{item.observation.instance.allocation_id ?? "Not available"}
+
Resolved
{formatDashboardTimestamp(item.observation.resolved_at)}
+
+
+
+ ); + }, + }, { + id: "mode", + accessorFn: runtimeModeLabel, + header: ({ column }) => undefined)} />, + cell: ({ row }) => {runtimeModeLabel(row.original)}, + }, { + id: "status", + accessorFn: (row) => row.observation.status, + header: ({ column }) => undefined)} />, + cell: ({ row }) => ( + + + ), + }, { + id: "cpu", + accessorFn: (row) => row.observation.cpu?.usage_seconds_total ?? -1, + header: ({ column }) => undefined)} />, + cell: ({ row }) => { + const cpu = row.original.observation.status === "observed" ? row.original.observation.cpu : null; + return {formatDashboardDuration(cpu?.usage_seconds_total ?? null)}{typeof cpu?.capacity_cores === "number" ? `${cpu.capacity_cores.toLocaleString("en-US")} cores` : "Capacity unknown"}; + }, + }, { + id: "memory", + accessorFn: (row) => row.observation.memory?.usage_bytes ?? -1, + header: ({ column }) => undefined)} />, + cell: ({ row }) => { + const memory = row.original.observation.status === "observed" ? row.original.observation.memory : null; + const memoryPercent = percent(memory?.usage_bytes, memory?.limit_bytes); + return ( + + {formatDashboardBytes(memory?.usage_bytes ?? null)} + {memory?.limit_bytes == null ? "Limit unknown" : `of ${formatDashboardBytes(memory.limit_bytes)}`} + {memoryPercent !== null ? : null} + + ); + }, + }, { + id: "uptime", + accessorFn: (row) => row.computeUptimeSeconds ?? -1, + header: ({ column }) => undefined)} />, + cell: ({ row }) => {formatDashboardDuration(row.original.computeUptimeSeconds)}{row.original.allocationAgeSeconds === null ? "Allocation age unknown" : `${formatDashboardDuration(row.original.allocationAgeSeconds)} allocated`}, + }, { + id: "sessionStatus", + accessorFn: (row) => row.session.status, + header: ({ column }) => undefined)} />, + cell: ({ row }) => {dashboardStatusLabel(row.original.session.status)}, + }, { + id: "tokens", + accessorFn: (row) => row.session.totalTokens ?? -1, + header: ({ column }) => undefined)} />, + cell: ({ row }) => {formatDashboardTokens(row.original.session.totalTokens)}{row.original.session.totalTokens === null ? "Not reported" : "Session reported"}, + }], [onOpenSession]); + const table = useReactTable({ + data: filteredRows, + columns, + state: { sorting, globalFilter }, + onSortingChange: setSorting, + onGlobalFilterChange: setGlobalFilter, + globalFilterFn: runtimeGlobalFilter, + getCoreRowModel: getCoreRowModel(), + getFilteredRowModel: getFilteredRowModel(), + getSortedRowModel: getSortedRowModel(), + getPaginationRowModel: getPaginationRowModel(), + initialState: { pagination: { pageIndex: 0, pageSize: PAGE_SIZE } }, + }); + const visibleRows = table.getFilteredRowModel().rows.length; + + return ( +
+
+
+

Runtime targets

+

Read-only Session navigation · missing measurements remain unknown

+
+ {visibleRows.toLocaleString("en-US")} visible +
+
+ + + +
+
+
+ + {table.getHeaderGroups().map((headerGroup) => ( + + {headerGroup.headers.map((header) => )} + + ))} + + + {table.getRowModel().rows.map((row) => ( + + {row.getVisibleCells().map((cell) => )} + + ))} + +
{flexRender(header.column.columnDef.header, header.getContext())}
{flexRender(cell.column.columnDef.cell, cell.getContext())}
+ {visibleRows === 0 ?

No Runtime targets match these filters.

: null} +
+ {table.getPageCount() > 1 ? ( +
+ Page {table.getState().pagination.pageIndex + 1} of {table.getPageCount()} +
+ + +
+
+ ) : null} + + ); +} + +export function RuntimeObservabilityContent({ + snapshot, + stale, + onOpenSession, +}: { + snapshot: RuntimeDashboardSnapshot; + stale: boolean; + onOpenSession: (sessionId: string) => void; +}) { + const model = useMemo(() => buildRuntimeDashboardModel(snapshot.sessions, snapshot.observations), [snapshot]); + const summary = model.summary; + + return ( + <> +
+ } label="Active Runtimes" value={summary.observedRuntimeCount.toLocaleString("en-US")} detail={`${summary.managedRuntimeCount} managed · ${summary.unavailableRuntimeCount} unavailable`} /> + } label="CPU time / capacity" value={summary.cpuUsageSecondsTotal === null && summary.cpuCapacityCores === null ? "No current sample" : `${formatDashboardDuration(summary.cpuUsageSecondsTotal)} / ${summary.cpuCapacityCores?.toLocaleString("en-US") ?? "—"} cores`} detail={`${summary.cpuCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report CPU`} /> + } label="Memory" value={summary.memoryUsageBytes === null && summary.memoryLimitBytes === null ? "No current sample" : `${formatDashboardBytes(summary.memoryUsageBytes)} / ${formatDashboardBytes(summary.memoryLimitBytes)}`} detail={`${summary.memoryCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report memory`} /> + } label="Reported tokens" value={formatDashboardTokens(summary.totalTokens)} detail={`${summary.tokenCoverageCount}/${summary.sessionCount} Sessions report usage`} /> +
+ +
+
+

Runtime health

Status and observation freshness by Session

{summary.sessionCount} contexts
+
+ {model.rows.slice(0, 8).map((row) => ( + + ))} + {model.rows.length > 8 ? +{model.rows.length - 8} more : null} +
+
+
+

Usage coverage

Unknown remains unknown, never zero

{stale ? "retained snapshot" : "current snapshot"}
+
+
CPU {summary.cpuCoverageCount}/{summary.observedRuntimeCount}{coverage(summary.cpuCoverageCount, summary.observedRuntimeCount)}
+
Memory {summary.memoryCoverageCount}/{summary.observedRuntimeCount}{coverage(summary.memoryCoverageCount, summary.observedRuntimeCount)}
+
Tokens {summary.tokenCoverageCount}/{summary.sessionCount}{coverage(summary.tokenCoverageCount, summary.sessionCount)}
+
Unavailable {summary.unavailableRuntimeCount}current observations
+
Unsupported {summary.unsupportedRuntimeCount}self-hosted or none
+
+
+
+ + + + ); +} diff --git a/apps/web/src/features/dashboard/dashboard-model.test.ts b/apps/web/src/features/dashboard/dashboard-model.test.ts index b6f35bd0b..03a167088 100644 --- a/apps/web/src/features/dashboard/dashboard-model.test.ts +++ b/apps/web/src/features/dashboard/dashboard-model.test.ts @@ -275,6 +275,8 @@ describe("Dashboard loaded-snapshot model", () => { sessionCount: 2, managedRuntimeCount: 1, observedRuntimeCount: 1, + unavailableRuntimeCount: 0, + unsupportedRuntimeCount: 1, cpuUsageSecondsTotal: 3.5, cpuCapacityCores: 2, cpuCoverageCount: 1, @@ -320,4 +322,50 @@ describe("Dashboard loaded-snapshot model", () => { expect(buildRuntimeDashboardModel([stopped], [observation]).rows[0]?.allocationAgeSeconds).toBeNull(); }); + + it("does not count capacity-only or limit-only samples as usage coverage", () => { + const managed = session("11111111-1111-4111-8111-111111111111", { + environment: { + type: "openai_hosted", + id: "33333333-3333-4333-8333-333333333333", + capability_directories: [], + network: { access: "disabled", allowed_domains: [] }, + packages: { npm: [], python: [], system: [] }, + files: [], + plugins: [], + skills: [], + }, + }); + const observation: RuntimeObservation = { + id: managed.id, + object: "agent.runtime_observation", + session_id: managed.id, + environment_id: "33333333-3333-4333-8333-333333333333", + mode: "openai_hosted", + provider_type: "docker", + instance: { + kind: "managed_allocation", + allocation_id: "44444444-4444-4444-8444-444444444444", + device_id: null, + connection_generation: null, + }, + status: "observed", + reason: null, + allocation_created_at: null, + resolved_at: 220, + observed_at: 210, + started_at: null, + cpu: { usage_seconds_total: null, capacity_cores: 2, usage_cores: null, utilization_ratio: null }, + memory: { usage_bytes: null, limit_bytes: 2048 }, + }; + + expect(buildRuntimeDashboardModel([managed], [observation]).summary).toMatchObject({ + cpuUsageSecondsTotal: null, + cpuCapacityCores: 2, + cpuCoverageCount: 0, + memoryUsageBytes: null, + memoryLimitBytes: 2048, + memoryCoverageCount: 0, + }); + }); }); diff --git a/apps/web/src/features/dashboard/dashboard-model.ts b/apps/web/src/features/dashboard/dashboard-model.ts index 89e36e2a4..5c31b55f1 100644 --- a/apps/web/src/features/dashboard/dashboard-model.ts +++ b/apps/web/src/features/dashboard/dashboard-model.ts @@ -58,6 +58,7 @@ export interface RuntimeDashboardSummary { managedRuntimeCount: number; observedRuntimeCount: number; unavailableRuntimeCount: number; + unsupportedRuntimeCount: number; cpuUsageSecondsTotal: number | null; cpuCapacityCores: number | null; cpuCoverageCount: number; @@ -283,6 +284,7 @@ export function buildRuntimeDashboardModel( let managedRuntimeCount = 0; let observedRuntimeCount = 0; let unavailableRuntimeCount = 0; + let unsupportedRuntimeCount = 0; let cpuUsageSecondsTotal = 0; let cpuUsageKnown = false; let cpuUsageSafe = true; @@ -332,7 +334,7 @@ export function buildRuntimeDashboardModel( cpuCapacityKnown = true; } else cpuCapacitySafe = false; } - if (cpuUsage !== null || cpuCapacity !== null) cpuCoverageCount += 1; + if (cpuUsage !== null) cpuCoverageCount += 1; const memoryUsage = safeNonNegativeInteger(observation.memory?.usage_bytes); const memoryLimit = safeNonNegativeInteger(observation.memory?.limit_bytes); @@ -350,9 +352,11 @@ export function buildRuntimeDashboardModel( memoryLimitKnown = true; } else memoryLimitSafe = false; } - if (memoryUsage !== null || memoryLimit !== null) memoryCoverageCount += 1; + if (memoryUsage !== null) memoryCoverageCount += 1; } else if (observation.status === "unavailable") { unavailableRuntimeCount += 1; + } else { + unsupportedRuntimeCount += 1; } if (sessionRow.totalTokens !== null) { @@ -386,6 +390,7 @@ export function buildRuntimeDashboardModel( managedRuntimeCount, observedRuntimeCount, unavailableRuntimeCount, + unsupportedRuntimeCount, cpuUsageSecondsTotal: cpuUsageKnown && cpuUsageSafe ? cpuUsageSecondsTotal : null, cpuCapacityCores: cpuCapacityKnown && cpuCapacitySafe ? cpuCapacityCores : null, cpuCoverageCount, diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index d49bf3ed5..f3522662d 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -244,8 +244,10 @@ as zero. ### 11.2 Runtime table Each row shows Session, Agent/harness when already available from the Session -snapshot, mode, observation status, CPU, memory, compute uptime, Turn state, and -reported tokens. Rows navigate to the existing Session view. No stop, restart, +snapshot, mode, observation status, cumulative CPU time, memory, compute uptime, +Session state, and reported tokens. The semantic table supports local search, +status and mode filters, sortable columns, and bounded pagination over the last +complete snapshot. Rows navigate to the existing Session view. No stop, restart, pause, or delete actions appear in the first release. ### 11.3 Detail view @@ -281,8 +283,9 @@ seams separate: 3. A feature-local state model retains `last_complete`, current refresh status, local filters, and the selected time range. It aborts an overlapping refresh and marks old data stale after a failed or incomplete refresh. -4. Presentational components render summary coverage, the Runtime table, and a - Session detail surface. Trend components are absent unless a later history +4. Presentational components render summary metrics, a compact health matrix, + explicit usage coverage, the Runtime table, and an identity detail surface. + Trend components and chart dependencies are absent unless a later history capability and contract are configured. The initial implementation uses a 30-second Web cadence plus up to five seconds diff --git a/docs/web/README.md b/docs/web/README.md index c94068ae1..398ccacc1 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -32,8 +32,10 @@ credentials or execution into the browser. Dashboard is the starting point. It summarizes the current Agent and Session results, loads a complete tenant-scoped Runtime observation snapshot, shows current Docker resource evidence and coverage without inventing missing values, highlights Sessions -that need attention, and links directly to Agent creation or a new Session. Historical -charts remain absent until an operator configures a separate history capability. +that need attention, and links directly to Agent creation or a new Session. Runtime +health and coverage cards sit above a searchable, filterable, sortable, paginated +semantic table. Historical charts remain absent until an operator configures a +separate history capability. ### Agents diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index a1afb5249..77b804710 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -29,8 +29,9 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载完整的租户级 Runtime 观测快照,并在不把缺失值伪装成 0 的前提下展示当前 Docker 资源和数据覆盖率。页面也 -展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。在运维方配置 -独立历史能力之前,页面不会伪造历史趋势图。 +展示 Runtime health 与 coverage 卡片,以及支持搜索、状态/模式筛选、排序、分页的语义 +表格;同时展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。 +在运维方配置独立历史能力之前,页面不会伪造历史趋势图,也不会引入趋势图组件。 ### Agents diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index dd71da556..d46c38103 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -23,6 +23,9 @@ importers: '@agents-core-web/agents-client': specifier: workspace:* version: link:../../packages/agents-client + '@tanstack/react-table': + specifier: ^8.21.3 + version: 8.21.3(react-dom@19.3.0(react@19.3.0))(react@19.3.0) lucide-react: specifier: ^1.22.0 version: 1.47.0(react@19.3.0) @@ -493,6 +496,17 @@ packages: '@standard-schema/spec@1.1.0': resolution: {integrity: sha512-l2aFy5jALhniG5HgqrD6jXLi/rUWrKvqN/qJx6yoJsgKhblVd+iqqU4RCXavm/jPityDo5TCvKMnpjKnOriy0w==} + '@tanstack/react-table@8.21.3': + resolution: {integrity: sha512-5nNMTSETP4ykGegmVkhjcS8tTLW6Vl4axfEGQN3v0zdHYbK4UfoqfPChclTrJ4EoK9QynqAu9oUf8VEmrpZ5Ww==} + engines: {node: '>=12'} + peerDependencies: + react: '>=16.8' + react-dom: '>=16.8' + + '@tanstack/table-core@8.21.3': + resolution: {integrity: sha512-ldZXEhOBb8Is7xLs01fR3YEc3DERiz5silj8tnGkFZytt1abEvl/GhUmCE0PMLaMPTa3Jk4HbKmRlHmu+gCftg==} + engines: {node: '>=12'} + '@types/chai@5.2.3': resolution: {integrity: sha512-Mw558oeA9fFbv65/y4mHtXDs9bPnFMZAL/jxdPFUpOHHIXX91mcgEHbS5Lahr+pwZFR8A7GQleRWeI6cGFC2UA==} @@ -1763,6 +1777,14 @@ snapshots: '@standard-schema/spec@1.1.0': {} + '@tanstack/react-table@8.21.3(react-dom@19.3.0(react@19.3.0))(react@19.3.0)': + dependencies: + '@tanstack/table-core': 8.21.3 + react: 19.3.0 + react-dom: 19.3.0(react@19.3.0) + + '@tanstack/table-core@8.21.3': {} + '@types/chai@5.2.3': dependencies: '@types/deep-eql': 4.0.2 From 2b8dd07c7be88ac7fc077ddaedfd793150d404de Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 22:01:33 +0800 Subject: [PATCH 06/51] Adapt Runtime observations to deployment providers --- services/agents-api/cmd/server/main.go | 9 +++------ 1 file changed, 3 insertions(+), 6 deletions(-) diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 6c97bf45c..46603fadc 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -96,12 +96,9 @@ func run() error { } observationSources := map[string]runtimeobs.Source{} if managed != nil { - for key, provider := range managed.Providers { - source, ok := provider.(runtimeobs.Source) - if !ok { - continue - } - observationSources[key] = source + source, ok := managed.Provider.(runtimeobs.Source) + if ok { + observationSources[managed.InstallationID] = source } } resolver, err := runtimeobs.NewResolver(executionStore) From d74d6bdbc2f429fd3cc4b720270bcb1d413a1b61 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 22:03:51 +0800 Subject: [PATCH 07/51] Add microsandbox Runtime observations --- CONTRIBUTING.md | 6 + .../agents-api/runtime-observability-api.md | 2 +- .../runtime-observability-design.md | 22 ++- contracts/agents-api/runtime-observability.md | 23 ++- services/agents-api/cmd/server/main.go | 8 +- .../agents-api/deploy/microsandbox/README.md | 18 +++ .../internal/runtimeobs/identity.go | 10 +- .../internal/runtimeobs/resolver.go | 1 + .../internal/runtimeobs/resolver_test.go | 8 +- .../internal/sandbox/microsandbox/identity.go | 2 +- .../internal/sandbox/microsandbox/provider.go | 3 + .../sandbox/microsandbox/resources.go | 71 +++++++++ .../sandbox/microsandbox/resources_test.go | 136 ++++++++++++++++++ .../internal/sandbox/microsandbox/types.go | 13 +- .../tools/microsandbox-provider/README.md | 8 ++ .../tools/microsandbox-provider/backend.go | 29 ++++ .../tools/microsandbox-provider/main.go | 2 + .../microsandbox-provider/metrics_test.go | 23 +++ 18 files changed, 363 insertions(+), 22 deletions(-) create mode 100644 services/agents-api/internal/sandbox/microsandbox/resources.go create mode 100644 services/agents-api/internal/sandbox/microsandbox/resources_test.go create mode 100644 services/agents-api/tools/microsandbox-provider/metrics_test.go diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c1065c007..94ec4031b 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -243,6 +243,12 @@ a provider source. Observation never extends a lease or changes compute lifecycl Keep observed zero, unavailable data and unsupported Runtime modes distinct. Metrics may inform operators, but automatic suspension requires durable Core-owned activity state and must not use a monitoring backend as lifecycle authority. +Managed Docker observes one non-streaming Inspect/Stats sample. Managed microsandbox +observes the exact persisted compute generation through the existing one-shot helper +and pinned SDK Metrics call. Preserve cumulative CPU seconds, memory usage/limit and +compute uptime semantics across both. Do not use microsandbox's instantaneous CPU +percent, wake suspended compute, or expose provider-native identifiers to fill a +common field. In V1, our daemon fills the user-side executor role. Users deploy daemon, the selected harness, local tools and workspace together. Do not require Codex diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 4d6345a51..7c967cb18 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -101,7 +101,7 @@ Session returns the existing indistinguishable not-found error. | `session_id` | string | yes | Authorized Core Session. | | `environment_id` | string or null | yes | Null only for mode `none`. | | `mode` | enum | yes | `none`, `self_hosted`, `openai_hosted`. | -| `provider_type` | string or null | yes | Forward-compatible safe source kind such as `docker`; null when no provider applies. Clients must not treat an unknown nonempty value as an error. | +| `provider_type` | string or null | yes | Forward-compatible safe source kind such as `docker` or `microsandbox`; null when no provider applies. Clients must not treat an unknown nonempty value as an error. | | `instance` | object | yes | Provider-neutral current incarnation identity; explicit `kind=none` when no compute applies. | | `status` | enum | yes | `observed`, `unsupported`, `unavailable`. | | `reason` | enum or null | yes | Safe reason when status is not `observed`. | diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index d49bf3ed5..c6322104d 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -1,9 +1,10 @@ # Runtime observability and Dashboard design -Status: Phase 1 provider abstraction/Docker sampling, Phase 2 current-snapshot -API/client contract, and the initial Core Web current-snapshot Dashboard are -implemented. The history backend, additional providers, and lifecycle automation -described below are not implemented. +Status: provider abstraction with Docker and microsandbox sampling, the +current-snapshot API/client contract, and the initial Core Web current-snapshot +Dashboard are implemented. The history backend and other provider sources are not +implemented. Microsandbox idle suspension is a separate durable lifecycle feature; +it does not consume this telemetry as authority. ## 1. Problem statement @@ -90,6 +91,7 @@ flowchart LR Resolver --> DB[(Core PostgreSQL)] Service --> Registry[provider source registry] Registry --> Docker[Docker Inspect and one-shot Stats] + Registry --> Micro[microsandbox exact-compute one-shot Metrics] Registry -. future .-> K8s[Kubernetes Metrics API or cAdvisor] Registry -. future .-> E2B[E2B metrics adapter] Registry -. future .-> Self[authenticated daemon telemetry] @@ -115,6 +117,13 @@ retain pod UID, container identity, and restart boundaries. E2B must use an API-supported instance identity rather than display names. Self-hosted metrics require authenticated daemon messages fenced by the current connection generation. +Microsandbox uses the persisted allocation's opaque compute receipt to select the +exact current generation. The pure-Go Core adapter sends a read-only request to the +existing one-shot Linux helper; the helper verifies allocation labels, compute ID, +generation and snapshot provenance before calling `SandboxHandle.Metrics`. It maps +`VCPUTimeNs`, memory usage/limit and uptime into the common sample. Instantaneous +`CPUPercent`, host RSS, disk/network counters and overlay usage remain unprojected. + ### 6.3 API composition The API resolves durable rows first, samples sources with bounded concurrency, @@ -329,10 +338,11 @@ transitions. That migration cannot read a monitoring backend as authority. ### Phase 1: provider-neutral foundation -Implemented in the Docker observability foundation. +Implemented for Docker and microsandbox. - `runtimeobs` identity, resolver, source, sample, and service. - Managed Docker Inspect/Stats source. +- Managed microsandbox exact-compute point-in-time Metrics source. - CPU, memory, and current compute start time. - Explicit unsupported and unavailable states. @@ -377,6 +387,8 @@ this does not add an upstream OpenAI operation. - Every observation proves tenant, Session, Environment, and Runtime-instance association. - Managed Docker emits correct present/absent semantics and never mutates compute. +- Managed microsandbox verifies the exact current generation and never mutates, + resumes, pauses, snapshots, stops, or removes compute while observing it. - Unsupported modes never look like zero usage. - Collection calls are bounded and one ordinary unavailable source invents no data. - Ownership/integrity mismatch fails closed. diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 658a136ed..7d9fae351 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -19,8 +19,8 @@ equivalent ownership data. A Session, daemon connection, process, container, and native harness Session are different identities and must not be substituted for one another. -The first implementation supports managed Docker allocations. `self_hosted` and -`none` are recognized but explicitly unsupported. A future self-hosted source must +The current implementation supports managed Docker and microsandbox allocations. +`self_hosted` and `none` are recognized but explicitly unsupported. A future self-hosted source must use authenticated daemon telemetry fenced by the current connection generation. Core must not attribute shared host statistics to an `environment:none` Session. @@ -47,6 +47,17 @@ are read-only; observation must not renew, restart, create, or stop the containe The Docker `StartedAt` value defines current compute uptime and resets after a container restart. +Microsandbox reports cumulative vCPU time, current guest memory usage, its effective +memory limit, and compute uptime through the pinned SDK's point-in-time metrics +operation. Core invokes that SDK only through the existing one-shot Linux helper. +The helper first verifies the allocation labels and exact persisted compute +generation/ID, then reads metrics under the allocation lock. Restored generations +therefore reset compute uptime without resetting allocation age. Paused, stopped, +suspended, metrics-disabled, and no-current-sample states are unavailable, never +observed zero. The SDK also supplies instantaneous CPU percent, host RSS, disk, +network, and overlay values; those are intentionally outside this public sample +until their cross-provider semantics and API fields are designed. + ## Duration boundaries These durations answer different questions and must remain separate: @@ -66,10 +77,10 @@ not become the lifecycle authority. ## First-phase boundary -The first phase adds no migration, public endpoint, Web view, metrics backend, -token aggregation, Kubernetes/E2B source, or automatic lifecycle action. The -internal source interface is intended to admit those providers without changing -Session attribution or the existing sandbox lifecycle interface. +The current-snapshot implementation adds no migration, metrics backend, token +duplication, Kubernetes/E2B source, or telemetry-driven lifecycle action. The +internal source interface admits Docker and microsandbox without changing Session +attribution or the existing sandbox lifecycle interface. The review proposal for later API and Web phases is split into the [full design](runtime-observability-design.md) and the diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 91dbc861f..65e274056 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -92,12 +92,8 @@ func run() error { } observationSources := map[string]runtimeobs.Source{} if managed != nil { - for key, provider := range managed.Providers { - source, ok := provider.(runtimeobs.Source) - if !ok { - continue - } - observationSources[key] = source + if source, ok := managed.Provider.(runtimeobs.Source); ok { + observationSources[managed.InstallationID] = source } } resolver, err := runtimeobs.NewResolver(executionStore) diff --git a/services/agents-api/deploy/microsandbox/README.md b/services/agents-api/deploy/microsandbox/README.md index abaac2d0b..984f94953 100644 --- a/services/agents-api/deploy/microsandbox/README.md +++ b/services/agents-api/deploy/microsandbox/README.md @@ -143,6 +143,24 @@ The outward Core URL must be reachable from the guest. `localhost` inside the microVM refers to the guest. Model provider settings continue to use the existing private `AGENTS_API_EXECUTION_OPTIONS_FILE` contract. +## Runtime observations + +The configured microsandbox provider uses the same public Runtime observation rows +and Core Web Dashboard as Docker. Each request resolves Session, Environment, +allocation, installation and the persisted current compute generation before +invoking the helper. The helper performs one read-only `SandboxHandle.Metrics` +call; it does not wake suspended compute or extend retention. + +The current projection includes cumulative vCPU seconds, configured vCPU capacity, +guest memory usage/limit and compute uptime. Suspended or otherwise non-running +compute reports `unavailable` with `runtime_not_running`; metrics that are disabled +or have no current SDK sample report `sample_unavailable`. The SDK's instantaneous +CPU percentage, host RSS, disk, network and overlay measurements are not yet +exposed. Token usage continues to come from Session/Turn usage, not this provider. + +Observation is operational evidence only. The idle suspension state machine uses +its durable activity and compute-phase records and never consults Dashboard samples. + ## Park, restore and verification Core first reserves the idle allocation. The daemon rejects suspension while diff --git a/services/agents-api/internal/runtimeobs/identity.go b/services/agents-api/internal/runtimeobs/identity.go index 17e38f7d4..f0ccdd14b 100644 --- a/services/agents-api/internal/runtimeobs/identity.go +++ b/services/agents-api/internal/runtimeobs/identity.go @@ -1,6 +1,9 @@ package runtimeobs -import "time" +import ( + "encoding/json" + "time" +) // Instance is one provider-owned Runtime incarnation. AllocationID is present // for managed compute. DeviceID and ConnectionGeneration are reserved for a @@ -12,6 +15,11 @@ type Instance struct { ConnectionGeneration string AllocationState string AllocationCreatedAt time.Time + ComputePhase string + // ProviderState is the provider-owned, persisted compute receipt. It is + // internal-only and lets an observation source verify the exact current + // incarnation without accepting provider identifiers from the caller. + ProviderState json.RawMessage } // Target binds telemetry to durable Core identity. A Session is not itself a diff --git a/services/agents-api/internal/runtimeobs/resolver.go b/services/agents-api/internal/runtimeobs/resolver.go index 58a5f28df..d18fb6e99 100644 --- a/services/agents-api/internal/runtimeobs/resolver.go +++ b/services/agents-api/internal/runtimeobs/resolver.go @@ -78,6 +78,7 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Tar target.Instance = Instance{ AllocationID: allocation.ID, ProviderKey: allocation.ProviderKey, DeviceID: allocation.DeviceID, AllocationState: allocation.State, AllocationCreatedAt: allocation.CreatedAt, + ComputePhase: allocation.ComputePhase, ProviderState: append(json.RawMessage(nil), allocation.ComputeState...), } return target, nil default: diff --git a/services/agents-api/internal/runtimeobs/resolver_test.go b/services/agents-api/internal/runtimeobs/resolver_test.go index 7edb4ba48..e46e1bb6c 100644 --- a/services/agents-api/internal/runtimeobs/resolver_test.go +++ b/services/agents-api/internal/runtimeobs/resolver_test.go @@ -27,7 +27,7 @@ func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, allocation: store.RuntimeAllocation{ ID: "allocation", TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", - ProviderKey: "provider", DeviceID: "device", + ProviderKey: "provider", DeviceID: "device", ComputePhase: "running", ComputeState: []byte(`{"current":{"name":"sandbox"}}`), }, }) if err != nil { @@ -40,6 +40,12 @@ func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { if target.TenantID != "tenant" || target.SessionID != "session" || target.EnvironmentID != "environment" || target.Mode != ModeManaged || target.Instance.AllocationID != "allocation" || target.Instance.ProviderKey != "provider" || target.Instance.DeviceID != "device" { t.Fatalf("incorrect managed identity binding: %+v", target) } + if string(target.Instance.ProviderState) != `{"current":{"name":"sandbox"}}` { + t.Fatalf("provider state was not retained: %s", target.Instance.ProviderState) + } + if target.Instance.ComputePhase != "running" { + t.Fatalf("compute phase was not retained: %s", target.Instance.ComputePhase) + } } func TestResolverKeepsUnsupportedModesDistinct(t *testing.T) { diff --git a/services/agents-api/internal/sandbox/microsandbox/identity.go b/services/agents-api/internal/sandbox/microsandbox/identity.go index dd3e89b4b..7ef474c36 100644 --- a/services/agents-api/internal/sandbox/microsandbox/identity.go +++ b/services/agents-api/internal/sandbox/microsandbox/identity.go @@ -103,7 +103,7 @@ func ValidateRequest(q Request) error { if q.Bootstrap == nil || q.Bootstrap.Reference != q.Reference || ValidateBootstrap(*q.Bootstrap) != nil { return sandbox.ErrInvalid } - case "inspect", "kill", "resume_compute": + case "inspect", "kill", "resume_compute", "metrics": if q.Operation == "resume_compute" && q.Compute.ID == "" { return sandbox.ErrInvalid } diff --git a/services/agents-api/internal/sandbox/microsandbox/provider.go b/services/agents-api/internal/sandbox/microsandbox/provider.go index 3b74f72cd..feafd6c78 100644 --- a/services/agents-api/internal/sandbox/microsandbox/provider.go +++ b/services/agents-api/internal/sandbox/microsandbox/provider.go @@ -4,6 +4,7 @@ import ( "context" "errors" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" ) @@ -58,6 +59,8 @@ func (p *Provider) call(ctx context.Context, q Request) (Response, error) { return out, sandbox.ErrNotFound case "command_unconfirmed": return out, sandbox.ErrCommandUnconfirmed + case "metrics_unavailable": + return out, runtimeobs.ErrUnavailable default: return out, ErrUnconfirmed } diff --git a/services/agents-api/internal/sandbox/microsandbox/resources.go b/services/agents-api/internal/sandbox/microsandbox/resources.go new file mode 100644 index 000000000..2ab56d96f --- /dev/null +++ b/services/agents-api/internal/sandbox/microsandbox/resources.go @@ -0,0 +1,71 @@ +package microsandbox + +import ( + "context" + "encoding/json" + "errors" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" +) + +var _ runtimeobs.Source = (*Provider)(nil) + +func (*Provider) ObservationProviderType() string { return "microsandbox" } + +// Observe reads one point-in-time SDK metrics snapshot through the existing +// one-shot helper. The persisted compute receipt selects the exact generation; +// browser input and provider display names never select a sandbox. +func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runtimeobs.Sample, error) { + if target.Mode != runtimeobs.ModeManaged || target.Instance.AllocationID == "" { + return runtimeobs.Sample{}, sandbox.ErrInvalid + } + if target.Instance.ProviderKey != p.config.InstallationID { + return runtimeobs.Sample{}, sandbox.ErrOwnership + } + reference := sandbox.Reference{TenantID: target.TenantID, EnvironmentID: target.EnvironmentID, AllocationID: target.Instance.AllocationID} + if target.Instance.ComputePhase == "suspended" { + return runtimeobs.Sample{}, runtimeobs.ErrNotRunning + } + var state struct { + Current Compute `json:"current"` + } + if target.Instance.ComputePhase == "disabled" { + state.Current = p.Initial(reference) + } else if json.Unmarshal(target.Instance.ProviderState, &state) != nil || state.Current.ID == "" { + return runtimeobs.Sample{}, sandbox.ErrOwnership + } + if ValidateCompute(p.config, reference, state.Current) != nil { + return runtimeobs.Sample{}, sandbox.ErrOwnership + } + out, err := p.call(ctx, Request{Operation: "metrics", Reference: reference, Compute: state.Current}) + if errors.Is(err, sandbox.ErrNotFound) { + return runtimeobs.Sample{}, runtimeobs.ErrNotRunning + } + if err != nil { + return runtimeobs.Sample{}, err + } + if out.Metrics == nil { + return runtimeobs.Sample{}, ErrUnconfirmed + } + return sampleFromMetrics(p.config, *out.Metrics) +} + +func sampleFromMetrics(config Config, metrics Metrics) (runtimeobs.Sample, error) { + if metrics.ObservedAt.IsZero() || metrics.Uptime < 0 || metrics.MemoryLimitBytes == 0 || config.CPUs == 0 { + return runtimeobs.Sample{}, ErrUnconfirmed + } + startedAt := metrics.ObservedAt.Add(-metrics.Uptime) + if startedAt.Unix() < 0 || startedAt.After(metrics.ObservedAt) { + return runtimeobs.Sample{}, ErrUnconfirmed + } + cpuUsage := float64(metrics.VCPUTimeNs) / float64(time.Second) + cpuCapacity := float64(config.CPUs) + memoryUsage, memoryLimit := metrics.MemoryBytes, metrics.MemoryLimitBytes + return runtimeobs.Sample{ + ObservedAt: metrics.ObservedAt, StartedAt: &startedAt, + CPUUsageSecondsTotal: &cpuUsage, CPUCapacityCores: &cpuCapacity, + MemoryUsageBytes: &memoryUsage, MemoryLimitBytes: &memoryLimit, + }, nil +} diff --git a/services/agents-api/internal/sandbox/microsandbox/resources_test.go b/services/agents-api/internal/sandbox/microsandbox/resources_test.go new file mode 100644 index 000000000..d2b6be0f8 --- /dev/null +++ b/services/agents-api/internal/sandbox/microsandbox/resources_test.go @@ -0,0 +1,136 @@ +package microsandbox + +import ( + "context" + "encoding/json" + "errors" + "testing" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" +) + +func observationTarget(t *testing.T, compute Compute) runtimeobs.Target { + t.Helper() + state, err := json.Marshal(struct { + Current Compute `json:"current"` + }{Current: compute}) + if err != nil { + t.Fatal(err) + } + r := testRef() + return runtimeobs.Target{ + TenantID: r.TenantID, SessionID: "55555555-5555-4555-8555-555555555555", EnvironmentID: r.EnvironmentID, Mode: runtimeobs.ModeManaged, + Instance: runtimeobs.Instance{AllocationID: r.AllocationID, ProviderKey: testConfig().InstallationID, AllocationState: "running", ComputePhase: "running", ProviderState: state}, + } +} + +func TestObserveUsesPersistedExactComputeAndNormalizesMetrics(t *testing.T) { + config, reference := testConfig(), testRef() + compute := Compute{Name: Name(config, reference, 3), ID: "local:current", Generation: 3, RestoredFrom: func() *SnapshotIdentity { value := testSnapshot(); return &value }()} + // The snapshot fixture belongs to generation zero and is valid provenance for generation three. + observed := time.Date(2026, 9, 22, 12, 0, 0, 0, time.UTC) + provider, err := NewWithCaller(config, callerFunc(func(_ context.Context, request Request) (Response, error) { + if request.Operation != "metrics" || request.Reference != reference || request.Compute.ID != compute.ID || request.Compute.Generation != compute.Generation { + t.Fatalf("unexpected metrics request: %+v", request) + } + return Response{Version: ProtocolVersion, Metrics: &Metrics{ + ObservedAt: observed, Uptime: 5 * time.Minute, VCPUTimeNs: 2_500_000_000, + MemoryBytes: 1024, MemoryLimitBytes: 2 * 1024 * 1024, + }}, nil + })) + if err != nil { + t.Fatal(err) + } + sample, err := provider.Observe(deadline(t), observationTarget(t, compute)) + if err != nil { + t.Fatal(err) + } + if sample.StartedAt == nil || !sample.StartedAt.Equal(observed.Add(-5*time.Minute)) || sample.CPUUsageSecondsTotal == nil || *sample.CPUUsageSecondsTotal != 2.5 || sample.CPUCapacityCores == nil || *sample.CPUCapacityCores != 2 || sample.MemoryUsageBytes == nil || *sample.MemoryUsageBytes != 1024 || sample.MemoryLimitBytes == nil || *sample.MemoryLimitBytes != 2*1024*1024 { + t.Fatalf("bad normalized sample: %+v", sample) + } +} + +func TestObserveRejectsForeignOrMissingComputeBeforeHelper(t *testing.T) { + config, reference := testConfig(), testRef() + compute := Compute{Name: Name(config, reference, 0), ID: "local:current"} + calls := 0 + provider, _ := NewWithCaller(config, callerFunc(func(context.Context, Request) (Response, error) { calls++; return Response{}, nil })) + + foreign := observationTarget(t, compute) + foreign.Instance.ProviderKey = "66666666-6666-4666-8666-666666666666" + if _, err := provider.Observe(deadline(t), foreign); !errors.Is(err, sandbox.ErrOwnership) { + t.Fatalf("foreign installation accepted: %v", err) + } + missing := observationTarget(t, compute) + missing.Instance.ProviderState = json.RawMessage(`{}`) + if _, err := provider.Observe(deadline(t), missing); !errors.Is(err, sandbox.ErrOwnership) { + t.Fatalf("missing compute accepted: %v", err) + } + if calls != 0 { + t.Fatalf("invalid identity reached helper: %d calls", calls) + } +} + +func TestObserveDisabledUsesInitialComputeAndSuspendedDoesNotWake(t *testing.T) { + config, reference := testConfig(), testRef() + calls := 0 + provider, _ := NewWithCaller(config, callerFunc(func(_ context.Context, request Request) (Response, error) { + calls++ + if request.Compute != providerInitial(config, reference) { + t.Fatalf("disabled allocation did not use initial compute: %+v", request.Compute) + } + return Response{Version: ProtocolVersion, Metrics: &Metrics{ObservedAt: time.Now().UTC(), MemoryLimitBytes: 1}}, nil + })) + target := observationTarget(t, Compute{Name: Name(config, reference, 0), ID: "unused"}) + target.Instance.ComputePhase, target.Instance.ProviderState = "disabled", nil + if _, err := provider.Observe(deadline(t), target); err != nil { + t.Fatal(err) + } + target.Instance.ComputePhase = "suspended" + if _, err := provider.Observe(deadline(t), target); !errors.Is(err, runtimeobs.ErrNotRunning) { + t.Fatalf("suspended compute was not unavailable: %v", err) + } + if calls != 1 { + t.Fatalf("suspended compute reached helper: %d calls", calls) + } +} + +func providerInitial(config Config, reference sandbox.Reference) Compute { + return Compute{Name: Name(config, reference, 0)} +} + +func TestObserveMapsStoppedAndMetricsUnavailable(t *testing.T) { + config, reference := testConfig(), testRef() + compute := Compute{Name: Name(config, reference, 0), ID: "local:current"} + for _, test := range []struct { + name string + response Response + want error + }{ + {name: "stopped", response: Response{Version: ProtocolVersion, ErrorCode: "not_found"}, want: runtimeobs.ErrNotRunning}, + {name: "metrics unavailable", response: Response{Version: ProtocolVersion, ErrorCode: "metrics_unavailable"}, want: runtimeobs.ErrUnavailable}, + } { + t.Run(test.name, func(t *testing.T) { + provider, _ := NewWithCaller(config, callerFunc(func(context.Context, Request) (Response, error) { return test.response, nil })) + if _, err := provider.Observe(deadline(t), observationTarget(t, compute)); !errors.Is(err, test.want) { + t.Fatalf("error=%v want=%v", err, test.want) + } + }) + } +} + +func TestSampleFromMetricsRejectsInvalidTimingAndLimit(t *testing.T) { + config := testConfig() + for _, metrics := range []Metrics{ + {}, + {ObservedAt: time.Unix(1, 0), Uptime: -time.Second, MemoryLimitBytes: 1}, + {ObservedAt: time.Unix(1, 0), Uptime: 2 * time.Second, MemoryLimitBytes: 1}, + {ObservedAt: time.Unix(1, 0), MemoryLimitBytes: 0}, + } { + if _, err := sampleFromMetrics(config, metrics); err == nil { + t.Fatalf("invalid metrics accepted: %+v", metrics) + } + } +} diff --git a/services/agents-api/internal/sandbox/microsandbox/types.go b/services/agents-api/internal/sandbox/microsandbox/types.go index 5a1b7a3f7..4e394df35 100644 --- a/services/agents-api/internal/sandbox/microsandbox/types.go +++ b/services/agents-api/internal/sandbox/microsandbox/types.go @@ -9,7 +9,7 @@ import ( "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" ) -const ProtocolVersion = 1 +const ProtocolVersion = 2 const SDKVersion = "v0.7.2" const MaxOutputBytes = 1024 * 1024 const MaxRequestBytes = 72 * 1024 * 1024 @@ -65,9 +65,20 @@ type Response struct { Version int State *State Command *sandbox.CommandResult + Metrics *Metrics ErrorCode string } +// Metrics is the bounded provider-helper projection used by Core observability. +// It intentionally excludes instantaneous CPU percent and provider-native names. +type Metrics struct { + ObservedAt time.Time + Uptime time.Duration + VCPUTimeNs uint64 + MemoryBytes uint64 + MemoryLimitBytes uint64 +} + type Caller interface { Call(context.Context, Request) (Response, error) } diff --git a/services/agents-api/tools/microsandbox-provider/README.md b/services/agents-api/tools/microsandbox-provider/README.md index 484a44c0a..7c9ffd3bf 100644 --- a/services/agents-api/tools/microsandbox-provider/README.md +++ b/services/agents-api/tools/microsandbox-provider/README.md @@ -119,6 +119,14 @@ successful stdin completion and an explicit guest exit before returning a result Timeouts, output overflow and missing receipts return ErrCommandUnconfirmed. Closing an SDK exec handle alone does not prove the guest process exited. +The read-only metrics operation verifies the same exact allocation and compute +incarnation before calling the pinned SDK's point-in-time `SandboxHandle.Metrics`. +It returns only observation time, uptime, cumulative vCPU time and guest memory +usage/limit to Core. It does not connect to the guest, renew activity, resume paused +compute or mutate lifecycle state. SDK metrics-disabled and no-current-sample errors +are reduced to one safe unavailable code; raw diagnostics never cross the helper +boundary. + ## Acceptance boundary The feature is idle-only: Core must reserve a terminal Session with no pending diff --git a/services/agents-api/tools/microsandbox-provider/backend.go b/services/agents-api/tools/microsandbox-provider/backend.go index 9379631fb..92803d71c 100644 --- a/services/agents-api/tools/microsandbox-provider/backend.go +++ b/services/agents-api/tools/microsandbox-provider/backend.go @@ -5,6 +5,7 @@ package main import ( "context" "errors" + "time" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" wire "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox/microsandbox" @@ -23,6 +24,9 @@ func (b backend) run(ctx context.Context) (wire.Response, error) { case "inspect": _, s, e := b.inspect(ctx, b.q.Compute) return wire.Response{State: &s}, e + case "metrics": + m, e := b.metrics(ctx, b.q.Compute) + return wire.Response{Metrics: m}, e case "kill": return wire.Response{}, b.kill(ctx, b.q.Compute) case "resume_compute": @@ -60,6 +64,31 @@ func (b backend) run(ctx context.Context) (wire.Response, error) { } return wire.Response{}, sandbox.ErrInvalid } + +func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, error) { + h, state, err := b.inspect(ctx, c) + if err != nil { + return nil, err + } + if state.Status != string(sdk.SandboxStatusRunning) && state.Status != "draining" { + return nil, sandbox.ErrNotFound + } + metricsCtx, cancel := context.WithDeadline(ctx, b.q.Deadline) + defer cancel() + metrics, err := h.Metrics(metricsCtx) + if err != nil { + return nil, err + } + return projectMetrics(metrics, time.Now().UTC()), nil +} + +func projectMetrics(metrics *sdk.Metrics, observedAt time.Time) *wire.Metrics { + return &wire.Metrics{ + ObservedAt: observedAt, Uptime: metrics.Uptime, + VCPUTimeNs: metrics.VCPUTimeNs, MemoryBytes: metrics.MemoryBytes, + MemoryLimitBytes: metrics.MemoryLimitBytes, + } +} func (b backend) inspect(ctx context.Context, c wire.Compute) (*sdk.SandboxHandle, wire.State, error) { state := wire.State{Compute: c} h, e := sdk.GetSandbox(ctx, c.Name) diff --git a/services/agents-api/tools/microsandbox-provider/main.go b/services/agents-api/tools/microsandbox-provider/main.go index 250eb421d..d1276b411 100644 --- a/services/agents-api/tools/microsandbox-provider/main.go +++ b/services/agents-api/tools/microsandbox-provider/main.go @@ -75,6 +75,8 @@ func code(err error) string { return "not_found" case errors.Is(err, sandbox.ErrCommandUnconfirmed): return "command_unconfirmed" + case sdk.IsKind(err, sdk.ErrMetricsDisabled), sdk.IsKind(err, sdk.ErrMetricsUnavailable): + return "metrics_unavailable" default: return "unconfirmed" } diff --git a/services/agents-api/tools/microsandbox-provider/metrics_test.go b/services/agents-api/tools/microsandbox-provider/metrics_test.go new file mode 100644 index 000000000..cba5dd18c --- /dev/null +++ b/services/agents-api/tools/microsandbox-provider/metrics_test.go @@ -0,0 +1,23 @@ +//go:build linux + +package main + +import ( + "testing" + "time" + + sdk "github.com/superradcompany/microsandbox/sdk/go" +) + +func TestProjectMetricsKeepsOnlyCumulativeCommonFields(t *testing.T) { + observed := time.Date(2026, 9, 22, 12, 0, 0, 0, time.UTC) + projected := projectMetrics(&sdk.Metrics{ + CPUPercent: 72.5, VCPUTimeNs: 2_500_000_000, + MemoryBytes: 4096, MemoryLimitBytes: 8192, + DiskReadBytes: 100, DiskWriteBytes: 200, NetRxBytes: 300, NetTxBytes: 400, + Uptime: 5 * time.Minute, + }, observed) + if projected.ObservedAt != observed || projected.Uptime != 5*time.Minute || projected.VCPUTimeNs != 2_500_000_000 || projected.MemoryBytes != 4096 || projected.MemoryLimitBytes != 8192 { + t.Fatalf("bad projection: %+v", projected) + } +} From f6381b40d1fe0d420ba516a1a6eb938a27f821a2 Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 22:24:29 +0800 Subject: [PATCH 08/51] Keep microsandbox metrics helper dependency-light --- services/agents-api/cmd/server/main.go | 3 +- .../agents-api/internal/runtimeobs/source.go | 10 ++++- .../{ => storeresolver}/resolver.go | 42 +++++++++---------- .../{ => storeresolver}/resolver_test.go | 7 ++-- .../tools/microsandbox-provider/backend.go | 9 +++- .../microsandbox-provider/metrics_test.go | 6 +++ 6 files changed, 48 insertions(+), 29 deletions(-) rename services/agents-api/internal/runtimeobs/{ => storeresolver}/resolver.go (59%) rename services/agents-api/internal/runtimeobs/{ => storeresolver}/resolver_test.go (91%) diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index be0ce7bcf..59cdefd9e 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -29,6 +29,7 @@ import ( "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtime" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeenrollment" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/storeresolver" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" "github.com/jackc/pgx/v5/pgxpool" ) @@ -100,7 +101,7 @@ func run() error { observationSources[managed.InstallationID] = source } } - resolver, err := runtimeobs.NewResolver(executionStore) + resolver, err := storeresolver.NewResolver(executionStore) if err != nil { return err } diff --git a/services/agents-api/internal/runtimeobs/source.go b/services/agents-api/internal/runtimeobs/source.go index 4044456ab..3a12b4907 100644 --- a/services/agents-api/internal/runtimeobs/source.go +++ b/services/agents-api/internal/runtimeobs/source.go @@ -1,6 +1,14 @@ package runtimeobs -import "context" +import ( + "context" + "errors" +) + +var ( + ErrUnavailable = errors.New("Runtime observation unavailable") + ErrNotRunning = errors.New("Runtime is not running") +) // Source reads one provider-owned Runtime instance. Implementations must verify // ownership before returning data and must not renew, restart, or stop compute. diff --git a/services/agents-api/internal/runtimeobs/resolver.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go similarity index 59% rename from services/agents-api/internal/runtimeobs/resolver.go rename to services/agents-api/internal/runtimeobs/storeresolver/resolver.go index d18fb6e99..b09bff1bd 100644 --- a/services/agents-api/internal/runtimeobs/resolver.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go @@ -1,4 +1,4 @@ -package runtimeobs +package storeresolver import ( "context" @@ -6,14 +6,10 @@ import ( "errors" "fmt" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" ) -var ( - ErrUnavailable = errors.New("Runtime observation unavailable") - ErrNotRunning = errors.New("Runtime is not running") -) - type sessionStore interface { GetSession(context.Context, string, string) (store.Session, error) GetRuntimeAllocation(context.Context, string, string) (store.RuntimeAllocation, error) @@ -28,10 +24,10 @@ func NewResolver(s sessionStore) (*Resolver, error) { return &Resolver{store: s}, nil } -func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Target, error) { +func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (runtimeobs.Target, error) { session, err := r.store.GetSession(ctx, tenantID, sessionID) if err != nil { - return Target{}, fmt.Errorf("resolve Runtime Session: %w", err) + return runtimeobs.Target{}, fmt.Errorf("resolve Runtime Session: %w", err) } var configuration struct { Environment *struct { @@ -39,49 +35,49 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (Tar } `json:"environment"` } if err := json.Unmarshal(session.Configuration, &configuration); err != nil || configuration.Environment == nil { - return Target{}, errors.New("invalid stored Runtime environment configuration") + return runtimeobs.Target{}, errors.New("invalid stored Runtime environment configuration") } - target := Target{TenantID: session.TenantID, SessionID: session.ID, Mode: Mode(configuration.Environment.Type)} + target := runtimeobs.Target{TenantID: session.TenantID, SessionID: session.ID, Mode: runtimeobs.Mode(configuration.Environment.Type)} switch target.Mode { - case ModeNone: + case runtimeobs.ModeNone: if session.Environment != nil { - return Target{}, errors.New("environment:none unexpectedly has a durable Environment") + return runtimeobs.Target{}, errors.New("environment:none unexpectedly has a durable Environment") } return target, nil - case ModeSelfHosted: + case runtimeobs.ModeSelfHosted: if session.Environment == nil { - return Target{}, errors.New("self-hosted Session is missing its Environment") + return runtimeobs.Target{}, errors.New("self-hosted Session is missing its Environment") } if session.Environment.TenantID != session.TenantID || session.Environment.SessionID != session.ID { - return Target{}, errors.New("self-hosted Environment does not match resolved ownership") + return runtimeobs.Target{}, errors.New("self-hosted Environment does not match resolved ownership") } target.EnvironmentID = session.Environment.ID return target, nil - case ModeManaged: + case runtimeobs.ModeManaged: if session.Environment == nil { - return Target{}, errors.New("managed Session is missing its Environment") + return runtimeobs.Target{}, errors.New("managed Session is missing its Environment") } if session.Environment.TenantID != session.TenantID || session.Environment.SessionID != session.ID { - return Target{}, errors.New("managed Environment does not match resolved ownership") + return runtimeobs.Target{}, errors.New("managed Environment does not match resolved ownership") } target.EnvironmentID = session.Environment.ID allocation, err := r.store.GetRuntimeAllocation(ctx, tenantID, target.EnvironmentID) if errors.Is(err, store.ErrNotFound) { - return target, ErrUnavailable + return target, runtimeobs.ErrUnavailable } if err != nil { - return Target{}, fmt.Errorf("resolve Runtime allocation: %w", err) + return runtimeobs.Target{}, fmt.Errorf("resolve Runtime allocation: %w", err) } if allocation.TenantID != tenantID || allocation.SessionID != session.ID || allocation.EnvironmentID != target.EnvironmentID { - return Target{}, errors.New("Runtime allocation does not match resolved ownership") + return runtimeobs.Target{}, errors.New("Runtime allocation does not match resolved ownership") } - target.Instance = Instance{ + target.Instance = runtimeobs.Instance{ AllocationID: allocation.ID, ProviderKey: allocation.ProviderKey, DeviceID: allocation.DeviceID, AllocationState: allocation.State, AllocationCreatedAt: allocation.CreatedAt, ComputePhase: allocation.ComputePhase, ProviderState: append(json.RawMessage(nil), allocation.ComputeState...), } return target, nil default: - return Target{}, errors.New("invalid stored Runtime environment type") + return runtimeobs.Target{}, errors.New("invalid stored Runtime environment type") } } diff --git a/services/agents-api/internal/runtimeobs/resolver_test.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go similarity index 91% rename from services/agents-api/internal/runtimeobs/resolver_test.go rename to services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go index e46e1bb6c..7a6ba56a6 100644 --- a/services/agents-api/internal/runtimeobs/resolver_test.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go @@ -1,10 +1,11 @@ -package runtimeobs +package storeresolver import ( "context" "errors" "testing" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" ) @@ -37,7 +38,7 @@ func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { if err != nil { t.Fatal(err) } - if target.TenantID != "tenant" || target.SessionID != "session" || target.EnvironmentID != "environment" || target.Mode != ModeManaged || target.Instance.AllocationID != "allocation" || target.Instance.ProviderKey != "provider" || target.Instance.DeviceID != "device" { + if target.TenantID != "tenant" || target.SessionID != "session" || target.EnvironmentID != "environment" || target.Mode != runtimeobs.ModeManaged || target.Instance.AllocationID != "allocation" || target.Instance.ProviderKey != "provider" || target.Instance.DeviceID != "device" { t.Fatalf("incorrect managed identity binding: %+v", target) } if string(target.Instance.ProviderState) != `{"current":{"name":"sandbox"}}` { @@ -76,7 +77,7 @@ func TestResolverReportsManagedAllocationAsUnavailable(t *testing.T) { t.Fatal(err) } target, err := r.Resolve(t.Context(), "tenant", "session") - if !errors.Is(err, ErrUnavailable) || target.EnvironmentID != "environment" || target.Mode != ModeManaged { + if !errors.Is(err, runtimeobs.ErrUnavailable) || target.EnvironmentID != "environment" || target.Mode != runtimeobs.ModeManaged { t.Fatalf("allocation absence was not preserved: %+v %v", target, err) } } diff --git a/services/agents-api/tools/microsandbox-provider/backend.go b/services/agents-api/tools/microsandbox-provider/backend.go index 92803d71c..c46b73092 100644 --- a/services/agents-api/tools/microsandbox-provider/backend.go +++ b/services/agents-api/tools/microsandbox-provider/backend.go @@ -79,10 +79,17 @@ func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, er if err != nil { return nil, err } - return projectMetrics(metrics, time.Now().UTC()), nil + projected := projectMetrics(metrics, time.Now().UTC()) + if projected == nil { + return nil, wire.ErrUnconfirmed + } + return projected, nil } func projectMetrics(metrics *sdk.Metrics, observedAt time.Time) *wire.Metrics { + if metrics == nil { + return nil + } return &wire.Metrics{ ObservedAt: observedAt, Uptime: metrics.Uptime, VCPUTimeNs: metrics.VCPUTimeNs, MemoryBytes: metrics.MemoryBytes, diff --git a/services/agents-api/tools/microsandbox-provider/metrics_test.go b/services/agents-api/tools/microsandbox-provider/metrics_test.go index cba5dd18c..0355f177e 100644 --- a/services/agents-api/tools/microsandbox-provider/metrics_test.go +++ b/services/agents-api/tools/microsandbox-provider/metrics_test.go @@ -21,3 +21,9 @@ func TestProjectMetricsKeepsOnlyCumulativeCommonFields(t *testing.T) { t.Fatalf("bad projection: %+v", projected) } } + +func TestProjectMetricsRejectsMissingSDKSample(t *testing.T) { + if projected := projectMetrics(nil, time.Now()); projected != nil { + t.Fatalf("missing SDK sample projected: %+v", projected) + } +} From 1f1a446c774b4ee517c430ca6f9b00d5f9674b0c Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 22:40:39 +0800 Subject: [PATCH 09/51] Fence microsandbox observation identity and deadline --- contracts/agents-api/runtime-observability.md | 3 +++ .../agents-api/deploy/microsandbox/README.md | 3 +++ .../sandbox/microsandbox/resources.go | 9 +++++--- .../sandbox/microsandbox/resources_test.go | 23 +++++++------------ .../tools/microsandbox-provider/backend.go | 6 ++--- 5 files changed, 23 insertions(+), 21 deletions(-) diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 7d9fae351..91e100f52 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -57,6 +57,9 @@ suspended, metrics-disabled, and no-current-sample states are unavailable, never observed zero. The SDK also supplies instantaneous CPU percent, host RSS, disk, network, and overlay values; those are intentionally outside this public sample until their cross-provider semantics and API fields are designed. +Legacy suspension-disabled allocations without a persisted exact compute receipt +are also unavailable. A deterministic sandbox name is not an incarnation identity +and is never used as a sampling fallback. ## Duration boundaries diff --git a/services/agents-api/deploy/microsandbox/README.md b/services/agents-api/deploy/microsandbox/README.md index 984f94953..7780e2b0a 100644 --- a/services/agents-api/deploy/microsandbox/README.md +++ b/services/agents-api/deploy/microsandbox/README.md @@ -157,6 +157,9 @@ compute reports `unavailable` with `runtime_not_running`; metrics that are disab or have no current SDK sample report `sample_unavailable`. The SDK's instantaneous CPU percentage, host RSS, disk, network and overlay measurements are not yet exposed. Token usage continues to come from Session/Turn usage, not this provider. +An allocation without a persisted exact compute receipt also reports +`sample_unavailable`; the deterministic sandbox name is not sufficient to identify +one compute incarnation safely. Observation is operational evidence only. The idle suspension state machine uses its durable activity and compute-phase records and never consults Dashboard samples. diff --git a/services/agents-api/internal/sandbox/microsandbox/resources.go b/services/agents-api/internal/sandbox/microsandbox/resources.go index 2ab56d96f..4408dddb5 100644 --- a/services/agents-api/internal/sandbox/microsandbox/resources.go +++ b/services/agents-api/internal/sandbox/microsandbox/resources.go @@ -31,9 +31,12 @@ func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runti var state struct { Current Compute `json:"current"` } - if target.Instance.ComputePhase == "disabled" { - state.Current = p.Initial(reference) - } else if json.Unmarshal(target.Instance.ProviderState, &state) != nil || state.Current.ID == "" { + if json.Unmarshal(target.Instance.ProviderState, &state) != nil || state.Current.ID == "" { + // Suspension-disabled allocations predate durable compute receipts. A + // deterministic name is not an incarnation fence, so do not sample it. + if target.Instance.ComputePhase == "disabled" { + return runtimeobs.Sample{}, runtimeobs.ErrUnavailable + } return runtimeobs.Sample{}, sandbox.ErrOwnership } if ValidateCompute(p.config, reference, state.Current) != nil { diff --git a/services/agents-api/internal/sandbox/microsandbox/resources_test.go b/services/agents-api/internal/sandbox/microsandbox/resources_test.go index d2b6be0f8..02bb3d2a8 100644 --- a/services/agents-api/internal/sandbox/microsandbox/resources_test.go +++ b/services/agents-api/internal/sandbox/microsandbox/resources_test.go @@ -73,34 +73,27 @@ func TestObserveRejectsForeignOrMissingComputeBeforeHelper(t *testing.T) { } } -func TestObserveDisabledUsesInitialComputeAndSuspendedDoesNotWake(t *testing.T) { - config, reference := testConfig(), testRef() +func TestObserveWithoutExactReceiptAndSuspendedDoesNotWake(t *testing.T) { + config := testConfig() calls := 0 provider, _ := NewWithCaller(config, callerFunc(func(_ context.Context, request Request) (Response, error) { calls++ - if request.Compute != providerInitial(config, reference) { - t.Fatalf("disabled allocation did not use initial compute: %+v", request.Compute) - } - return Response{Version: ProtocolVersion, Metrics: &Metrics{ObservedAt: time.Now().UTC(), MemoryLimitBytes: 1}}, nil + return Response{}, nil })) - target := observationTarget(t, Compute{Name: Name(config, reference, 0), ID: "unused"}) + target := observationTarget(t, Compute{Name: Name(config, testRef(), 0), ID: "unused"}) target.Instance.ComputePhase, target.Instance.ProviderState = "disabled", nil - if _, err := provider.Observe(deadline(t), target); err != nil { - t.Fatal(err) + if _, err := provider.Observe(deadline(t), target); !errors.Is(err, runtimeobs.ErrUnavailable) { + t.Fatalf("receipt-less compute was sampled: %v", err) } target.Instance.ComputePhase = "suspended" if _, err := provider.Observe(deadline(t), target); !errors.Is(err, runtimeobs.ErrNotRunning) { t.Fatalf("suspended compute was not unavailable: %v", err) } - if calls != 1 { - t.Fatalf("suspended compute reached helper: %d calls", calls) + if calls != 0 { + t.Fatalf("unfenced compute reached helper: %d calls", calls) } } -func providerInitial(config Config, reference sandbox.Reference) Compute { - return Compute{Name: Name(config, reference, 0)} -} - func TestObserveMapsStoppedAndMetricsUnavailable(t *testing.T) { config, reference := testConfig(), testRef() compute := Compute{Name: Name(config, reference, 0), ID: "local:current"} diff --git a/services/agents-api/tools/microsandbox-provider/backend.go b/services/agents-api/tools/microsandbox-provider/backend.go index c46b73092..8141db349 100644 --- a/services/agents-api/tools/microsandbox-provider/backend.go +++ b/services/agents-api/tools/microsandbox-provider/backend.go @@ -66,15 +66,15 @@ func (b backend) run(ctx context.Context) (wire.Response, error) { } func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, error) { - h, state, err := b.inspect(ctx, c) + metricsCtx, cancel := context.WithDeadline(ctx, b.q.Deadline) + defer cancel() + h, state, err := b.inspect(metricsCtx, c) if err != nil { return nil, err } if state.Status != string(sdk.SandboxStatusRunning) && state.Status != "draining" { return nil, sandbox.ErrNotFound } - metricsCtx, cancel := context.WithDeadline(ctx, b.q.Deadline) - defer cancel() metrics, err := h.Metrics(metricsCtx) if err != nil { return nil, err From 7643824ea706f5abcc213410bb2dd9aa21c024fe Mon Sep 17 00:00:00 2001 From: sam Date: Tue, 22 Sep 2026 23:43:39 +0800 Subject: [PATCH 10/51] Refine Runtime dashboard visualizations --- apps/web/e2e/agents-lifecycle.spec.ts | 112 +++++ .../src/features/dashboard/DashboardView.css | 410 ++++++++++++++++++ .../features/dashboard/DashboardView.test.tsx | 24 +- .../dashboard/RuntimeObservabilityContent.tsx | 234 +++++++--- 4 files changed, 719 insertions(+), 61 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 11cbd4591..1ab51a13a 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2414,6 +2414,118 @@ test("presents Dashboard page-chain results and System boundaries without extra await attachScreenshot(page, testInfo, "narrow-system-contract-boundary"); }); +test("renders Runtime telemetry as visual snapshot panels with details on demand", async ({ page }, testInfo) => { + const baseline = 1_789_438_800; + const sessionId = "11111111-1111-4111-8111-111111111111"; + const environmentId = "22222222-2222-4222-8222-222222222222"; + const allocationId = "33333333-3333-4333-8333-333333333333"; + const runtimeSession = { + id: sessionId, + object: "agent.session", + agent: { + id: "agent_b", + model: "fixture/model", + name: "Runtime analyst", + instructions: null, + multi_agent: { enabled: false, max_concurrent_subagents: null }, + reasoning: {}, + service_tier: "auto", + text: { format: { type: "text" }, verbosity: "medium" }, + tools: [], + }, + environment: { + type: "openai_hosted", + id: environmentId, + capability_directories: [], + network: { access: "enabled", allowed_domains: [] }, + packages: { npm: [], python: [], system: [] }, + files: [], + plugins: [], + skills: [], + }, + status: "in_progress", + error: null, + metadata: { title: "Repository migration" }, + required_actions: [], + vault_ids: [], + usage: { + input_tokens: 420_000, + output_tokens: 66_000, + total_tokens: 486_000, + input_tokens_details: { cached_tokens: 180_000 }, + output_tokens_details: { reasoning_tokens: 22_000 }, + }, + created_at: baseline - 9_000, + last_active_at: baseline - 20, + }; + const list = (data: unknown[]) => ({ + object: "list", + data, + has_more: false, + first_id: data.length > 0 ? sessionId : null, + last_id: data.length > 0 ? sessionId : null, + }); + + await page.goto("/"); + const dashboard = page.locator(".dashboard-page"); + await expect(dashboard.getByRole("heading", { name: "Dashboard", exact: true })).toBeVisible(); + + await page.route("**/v1/agents/sessions*", async (route) => { + if (new URL(route.request().url()).pathname !== "/v1/agents/sessions") return route.fallback(); + await route.fulfill({ status: 200, contentType: "application/json", body: JSON.stringify(list([runtimeSession])) }); + }); + await page.route("**/v1/agents/runtime-observations*", async (route) => { + await route.fulfill({ + status: 200, + contentType: "application/json", + body: JSON.stringify(list([{ + id: sessionId, + object: "agent.runtime_observation", + session_id: sessionId, + environment_id: environmentId, + mode: "openai_hosted", + provider_type: "docker", + instance: { kind: "managed_allocation", allocation_id: allocationId, device_id: null, connection_generation: null }, + status: "observed", + reason: null, + allocation_created_at: baseline - 8_500, + resolved_at: baseline, + observed_at: baseline - 1, + started_at: baseline - 8_100, + cpu: { usage_seconds_total: 7_350, capacity_cores: 2, usage_cores: 0.72, utilization_ratio: 0.36 }, + memory: { usage_bytes: 1_288_490_188, limit_bytes: 2_147_483_648 }, + }])), + }); + }); + + await dashboard.getByRole("button", { name: "Refresh Dashboard snapshot" }).click(); + await expect(dashboard.getByRole("heading", { name: "Runtime health" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Resource load" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Longest-running Runtimes" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Token consumption" })).toBeVisible(); + await expect(dashboard.getByLabel("CPU now: 36%")).toBeVisible(); + await expect(dashboard.getByLabel("Memory now: 60%")).toBeVisible(); + await expect(dashboard.getByRole("table", { name: "Runtime targets" })).not.toBeVisible(); + for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); + await expect(page.getByRole("button", { name: "Close notification" })).toHaveCount(0); + await attachElementScreenshot(dashboard.locator(".dashboard-runtime-panel"), testInfo, "runtime-visual-dashboard"); + + await dashboard.getByText("Runtime targets", { exact: true }).click(); + await expect(dashboard.getByRole("table", { name: "Runtime targets" })).toBeVisible(); + + await page.setViewportSize({ width: 390, height: 844 }); + await expect(dashboard.getByRole("heading", { name: "Resource load" })).toBeVisible(); + for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); + const widths = await dashboard.locator(".dashboard-runtime-panel").evaluate((element) => ({ + viewport: innerWidth, + document: document.documentElement.scrollWidth, + panel: element.getBoundingClientRect().width, + })); + expect(widths.document).toBeLessThanOrEqual(widths.viewport); + expect(widths.panel).toBeLessThanOrEqual(widths.viewport); + await attachElementScreenshot(dashboard.locator(".dashboard-runtime-panel"), testInfo, "runtime-visual-dashboard-narrow"); +}); + test("publishes Dashboard counts only after every top-level Agent and Session page loads", async ({ page, request }) => { await resetFixture(request); const agentAfters: Array = []; diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 13778f0fb..c582b05d3 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -683,6 +683,370 @@ white-space: nowrap; } +.dashboard-runtime-visual-grid { + display: grid; + padding: 14px; + grid-template-columns: repeat(2, minmax(0, 1fr)); + gap: 12px; + background: color-mix(in srgb, var(--surface-subtle) 68%, var(--surface)); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-visual-card { + min-width: 0; + min-height: 248px; + padding: 15px 16px; + background: var(--surface); + border: 1px solid var(--line); + border-radius: 9px; + box-shadow: 0 1px 0 color-mix(in srgb, var(--fg) 2%, transparent); +} + +.dashboard-runtime-visual-card > header { + display: flex; + min-height: 37px; + align-items: flex-start; + justify-content: space-between; + gap: 12px; +} + +.dashboard-runtime-visual-card h3, +.dashboard-runtime-visual-card p { + margin: 0; +} + +.dashboard-runtime-visual-card h3 { + color: var(--fg); + font-size: 12px; + font-weight: 650; + line-height: 17px; +} + +.dashboard-runtime-visual-card header p, +.dashboard-runtime-visual-card header > span { + color: var(--fg-muted); + font-size: 9px; + line-height: 14px; +} + +.dashboard-runtime-visual-card header > svg { + color: color-mix(in srgb, var(--accent) 78%, var(--fg-muted)); +} + +.dashboard-runtime-donut-layout { + display: grid; + min-height: 174px; + grid-template-columns: minmax(126px, .8fr) minmax(150px, 1.2fr); + align-items: center; + gap: 22px; +} + +.dashboard-runtime-donut { + width: 142px; + max-width: 100%; + position: relative; + justify-self: center; + aspect-ratio: 1; +} + +.dashboard-runtime-donut svg { + width: 100%; + height: 100%; + overflow: visible; +} + +.dashboard-runtime-donut circle { + fill: none; + stroke-width: 10; +} + +.dashboard-runtime-donut-track { + stroke: var(--surface-subtle); +} + +.dashboard-runtime-donut-segment { + transform: rotate(-90deg); + transform-origin: center; + stroke-linecap: butt; +} + +.dashboard-runtime-donut-observed { + stroke: var(--success); +} + +.dashboard-runtime-donut-unavailable { + stroke: var(--warning); +} + +.dashboard-runtime-donut-unsupported { + stroke: color-mix(in srgb, var(--fg-muted) 55%, var(--surface)); +} + +.dashboard-runtime-donut > span { + display: grid; + position: absolute; + inset: 0; + place-content: center; + text-align: center; +} + +.dashboard-runtime-donut strong { + font-family: var(--font-mono); + font-size: 25px; + font-variant-numeric: tabular-nums slashed-zero; + font-weight: 600; + line-height: 29px; +} + +.dashboard-runtime-donut small { + color: var(--fg-muted); + font-size: 9px; +} + +.dashboard-runtime-chart-legend { + display: grid; + gap: 10px; +} + +.dashboard-runtime-chart-legend > div { + display: grid; + grid-template-columns: 7px minmax(0, 1fr) auto; + align-items: center; + gap: 8px; + color: var(--fg-muted); + font-size: 10px; +} + +.dashboard-runtime-chart-legend i { + width: 7px; + height: 7px; + background: var(--fg-muted); + border-radius: 50%; +} + +.dashboard-runtime-chart-legend .dashboard-runtime-legend-observed { + background: var(--success); +} + +.dashboard-runtime-chart-legend .dashboard-runtime-legend-unavailable { + background: var(--warning); +} + +.dashboard-runtime-chart-legend .dashboard-runtime-legend-unsupported { + background: color-mix(in srgb, var(--fg-muted) 55%, var(--surface)); +} + +.dashboard-runtime-chart-legend strong { + color: var(--fg); + font-family: var(--font-mono); + font-variant-numeric: tabular-nums; + font-weight: 600; +} + +.dashboard-runtime-resource-meters { + display: grid; + min-height: 174px; + align-content: center; + gap: 25px; +} + +.dashboard-runtime-resource-meter { + display: grid; + gap: 7px; +} + +.dashboard-runtime-resource-meter > div:first-child { + display: flex; + align-items: baseline; + justify-content: space-between; + gap: 12px; +} + +.dashboard-runtime-resource-meter span, +.dashboard-runtime-resource-meter small { + color: var(--fg-muted); + font-size: 9px; + line-height: 14px; +} + +.dashboard-runtime-resource-meter strong { + overflow: hidden; + font-family: var(--font-mono); + font-size: 12px; + font-variant-numeric: tabular-nums slashed-zero; + font-weight: 600; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-resource-track { + height: 9px; + overflow: hidden; + background: var(--surface-subtle); + border-radius: 999px; +} + +.dashboard-runtime-resource-track > i { + display: block; + height: 100%; + background: linear-gradient(90deg, color-mix(in srgb, var(--accent) 70%, var(--success)), var(--accent)); + border-radius: inherit; +} + +.dashboard-runtime-resource-track > .dashboard-runtime-resource-unknown { + width: 100%; + background: repeating-linear-gradient(135deg, transparent 0 5px, color-mix(in srgb, var(--fg-muted) 18%, transparent) 5px 8px); +} + +.dashboard-runtime-ranked-bars { + display: grid; + min-height: 174px; + padding-top: 11px; + align-content: center; + gap: 11px; +} + +.dashboard-runtime-ranked-row { + display: grid; + grid-template-columns: minmax(0, 1fr) minmax(90px, 1.1fr); + align-items: center; + gap: 3px 12px; +} + +.dashboard-runtime-ranked-row > div:first-child { + display: flex; + min-width: 0; + align-items: center; + justify-content: space-between; + gap: 8px; +} + +.dashboard-runtime-ranked-row span, +.dashboard-runtime-ranked-row strong, +.dashboard-runtime-ranked-row small { + overflow: hidden; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-ranked-row span { + font-size: 10px; + font-weight: 550; +} + +.dashboard-runtime-ranked-row strong { + flex: 0 0 auto; + font-family: var(--font-mono); + font-size: 9px; + font-variant-numeric: tabular-nums; + font-weight: 600; +} + +.dashboard-runtime-ranked-row small { + grid-column: 1; + color: var(--fg-muted); + font-size: 8px; + line-height: 11px; +} + +.dashboard-runtime-ranked-track { + height: 7px; + overflow: hidden; + background: var(--surface-subtle); + border-radius: 999px; +} + +.dashboard-runtime-ranked-track > i { + display: block; + height: 100%; + background: linear-gradient(90deg, color-mix(in srgb, var(--accent) 70%, var(--success)), var(--accent)); + border-radius: inherit; +} + +.dashboard-runtime-ranked-tokens .dashboard-runtime-ranked-track > i { + background: linear-gradient(90deg, color-mix(in srgb, var(--warning) 78%, var(--accent)), var(--warning)); +} + +.dashboard-runtime-chart-empty { + display: flex; + min-height: 174px; + align-items: center; + justify-content: center; + margin: 0; + padding: 20px; + color: var(--fg-muted); + font-size: 10px; + line-height: 15px; + text-align: center; + background: repeating-linear-gradient(135deg, transparent 0 8px, color-mix(in srgb, var(--fg-muted) 3%, transparent) 8px 9px); + border-radius: 6px; +} + +.dashboard-runtime-explorer { + min-width: 0; +} + +.dashboard-runtime-explorer > summary { + display: flex; + min-height: 62px; + align-items: center; + justify-content: space-between; + padding: 11px 14px; + gap: 16px; + cursor: pointer; + list-style: none; + border-bottom: 1px solid transparent; +} + +.dashboard-runtime-explorer > summary::-webkit-details-marker { + display: none; +} + +.dashboard-runtime-explorer[open] > summary { + background: var(--surface-subtle); + border-bottom-color: var(--line); +} + +.dashboard-runtime-explorer > summary > span { + display: grid; + min-width: 0; + gap: 2px; +} + +.dashboard-runtime-explorer > summary > span:last-child { + display: flex; + flex: 0 0 auto; + grid-auto-flow: column; + align-items: center; + gap: 6px; + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + font-variant-numeric: tabular-nums; +} + +.dashboard-runtime-explorer > summary strong { + font-size: 12px; + line-height: 17px; +} + +.dashboard-runtime-explorer > summary small { + overflow: hidden; + color: var(--fg-muted); + font-size: 9px; + line-height: 13px; + text-overflow: ellipsis; + white-space: nowrap; +} + +.dashboard-runtime-explorer > summary svg { + transition: transform 160ms var(--ease-settle); +} + +.dashboard-runtime-explorer[open] > summary svg { + transform: rotate(180deg); +} + .dashboard-runtime-insights { display: grid; grid-template-columns: minmax(0, 1.15fr) minmax(0, .85fr); @@ -909,6 +1273,14 @@ padding: 0 24px 0 8px; } +.dashboard-runtime-visible-count { + flex: 0 0 auto; + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + font-variant-numeric: tabular-nums; +} + .dashboard-runtime-table-scroll { overflow-x: auto; } @@ -1233,6 +1605,10 @@ flex-wrap: wrap; } + .dashboard-runtime-visual-grid { + grid-template-columns: 1fr; + } + .dashboard-runtime-search { width: 100%; flex-basis: 100%; @@ -1247,6 +1623,12 @@ min-width: 0; flex: 1; } + + .dashboard-runtime-visible-count { + display: flex; + width: 100%; + justify-content: flex-end; + } } @media (max-width: 520px) { @@ -1293,6 +1675,34 @@ grid-template-columns: 1fr; } + .dashboard-runtime-visual-grid { + padding: 10px; + gap: 10px; + } + + .dashboard-runtime-visual-card { + min-height: 224px; + } + + .dashboard-runtime-donut-layout { + grid-template-columns: 118px minmax(130px, 1fr); + gap: 14px; + } + + .dashboard-runtime-donut { + width: 118px; + } + + .dashboard-runtime-ranked-row { + grid-template-columns: minmax(0, 1fr) minmax(70px, .8fr); + } + + .dashboard-runtime-explorer > summary { + align-items: flex-start; + flex-direction: column; + gap: 6px; + } + .dashboard-runtime-targets > header { align-items: flex-start; flex-direction: column; diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index c0887eba7..ea2b4c1cd 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -242,18 +242,22 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("Runtime monitoring"); expect(html).toContain("1/1 managed observed"); - expect(html).toContain("CPU time / capacity"); + expect(html).toContain("Cumulative CPU / capacity"); expect(html).toContain("1m 13s / 2 cores"); expect(html).toContain("512 MiB / 2.00 GiB"); + expect(html).toContain('aria-label="Runtime snapshot visualizations"'); expect(html).toContain("Runtime health"); - expect(html).toContain("Observed · 2023-11-14 22:14 UTC"); - expect(html).not.toContain("Observed · 10s"); - expect(html).toContain("Usage coverage"); - expect(html).toContain("CPU 1/1"); - expect(html).toContain("Memory 1/1"); - expect(html).toContain("Tokens 1/1"); - expect(html).toContain("Unsupported 0"); + expect(html).toContain('aria-label="1 of 1 Runtimes observed"'); + expect(html).toContain("Resource load"); + expect(html).toContain("CPU now"); + expect(html).toContain("Instantaneous CPU is unavailable; cumulative CPU time remains in Explorer"); + expect(html).toContain('aria-label="CPU now: Unavailable"'); + expect(html).toContain('aria-label="Memory now: 25%"'); + expect(html).toContain("Longest-running Runtimes"); + expect(html).toContain("1m 20s"); + expect(html).toContain("Token consumption"); expect(html).toContain("Runtime targets"); + expect(html).toContain('
'); expect(html).toContain("Search Runtime targets"); expect(html).toContain("All statuses"); expect(html).toContain("All modes"); @@ -261,9 +265,9 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("CPU time"); expect(html).toContain("Managed research"); expect(html).toContain("Identity"); - expect(html).toContain("Unknown remains unknown, never zero"); + expect(html).toContain("unknown remains unknown, never zero"); expect(html).not.toContain("CPU %"); - expect(html).not.toContain("CPU now"); + expect(html).not.toContain("CPU now: 25%"); expect(html).not.toContain("historical chart"); }); diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index c580856cc..eb88e0e9c 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -1,9 +1,11 @@ import { + Activity, ChevronDown, ChevronLeft, ChevronRight, ChevronsUpDown, ChevronUp, + Clock3, Cpu, Gauge, MemoryStick, @@ -66,9 +68,175 @@ function percent(usage: number | null | undefined, limit: number | null | undefi return Math.min(100, Math.max(0, usage / limit * 100)); } -function coverage(known: number, total: number): string { - if (total === 0) return "No observed Runtimes"; - return `${Math.round(known / total * 100)}% coverage`; +function finiteNonNegative(value: unknown): number | null { + return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; +} + +function meterPercent(value: number | null): number | null { + return value === null ? null : Math.min(100, Math.max(0, value * 100)); +} + +function formatRatio(value: number | null): string { + return value === null ? "Unavailable" : `${(value * 100).toFixed(value >= 0.1 ? 0 : 1)}%`; +} + +function RuntimeHealthChart({ rows }: { rows: RuntimeDashboardRow[] }) { + const counts = { + observed: rows.filter((row) => row.observation.status === "observed").length, + unavailable: rows.filter((row) => row.observation.status === "unavailable").length, + unsupported: rows.filter((row) => row.observation.status === "unsupported").length, + }; + const total = rows.length; + const segments = [ + { key: "observed", label: "Observed", count: counts.observed }, + { key: "unavailable", label: "Unavailable", count: counts.unavailable }, + { key: "unsupported", label: "Unsupported", count: counts.unsupported }, + ]; + let offset = 0; + + return ( +
+
+

Runtime health

Current observation status

+
+
+
+ + + {segments.map((segment) => { + const size = total === 0 ? 0 : segment.count / total * 100; + const circle = ; + offset += size; + return circle; + })} + + {counts.observed}of {total} +
+
+ {segments.map((segment) => ( +
{segment.label}{segment.count}
+ ))} +
+
+
+ ); +} + +function ResourceMeter({ + label, + value, + detail, + ratio, +}: { + label: string; + value: string; + detail: string; + ratio: number | null; +}) { + const width = meterPercent(ratio); + return ( +
+
{label}{value}
+
+ {width === null ? : } +
+ {detail} +
+ ); +} + +function RuntimeResourceChart({ rows }: { rows: RuntimeDashboardRow[] }) { + const observed = rows.filter((row) => row.observation.status === "observed"); + const cpuSamples = observed.flatMap((row) => { + const usage = finiteNonNegative(row.observation.cpu?.usage_cores); + const capacity = finiteNonNegative(row.observation.cpu?.capacity_cores); + return usage !== null && capacity !== null && capacity > 0 ? [{ usage, capacity }] : []; + }); + const memorySamples = observed.flatMap((row) => { + const usage = finiteNonNegative(row.observation.memory?.usage_bytes); + const limit = finiteNonNegative(row.observation.memory?.limit_bytes); + return usage !== null && limit !== null && limit > 0 ? [{ usage, limit }] : []; + }); + const cpuUsage = cpuSamples.reduce((total, sample) => total + sample.usage, 0); + const cpuCapacity = cpuSamples.reduce((total, sample) => total + sample.capacity, 0); + const memoryUsage = memorySamples.reduce((total, sample) => total + sample.usage, 0); + const memoryLimit = memorySamples.reduce((total, sample) => total + sample.limit, 0); + const cpuRatio = cpuSamples.length > 0 && cpuCapacity > 0 ? cpuUsage / cpuCapacity : null; + const memoryRatio = memorySamples.length > 0 && memoryLimit > 0 ? memoryUsage / memoryLimit : null; + + return ( +
+
+

Resource load

Point-in-time provider samples

+ {observed.length} observed +
+
+ + +
+
+ ); +} + +interface RankedBarItem { + id: string; + label: string; + detail: string; + value: number; + formatted: string; +} + +function RankedBars({ items, empty, tone }: { items: RankedBarItem[]; empty: string; tone: "uptime" | "tokens" }) { + const maximum = Math.max(0, ...items.map((item) => item.value)); + if (items.length === 0) return

{empty}

; + return ( +
+ {items.map((item) => ( +
+
{item.label}{item.formatted}
+
+ {item.detail} +
+ ))} +
+ ); +} + +function RuntimeRankChart({ rows, kind }: { rows: RuntimeDashboardRow[]; kind: "uptime" | "tokens" }) { + const uptime = kind === "uptime"; + const items = rows.flatMap((row): RankedBarItem[] => { + const value = uptime ? row.computeUptimeSeconds : row.session.totalTokens; + if (value === null || !Number.isFinite(value) || value < 0) return []; + return [{ + id: row.observation.id, + label: row.session.title, + detail: uptime ? runtimeModeLabel(row) : dashboardStatusLabel(row.session.status), + value, + formatted: uptime ? formatDashboardDuration(value) : formatDashboardTokens(value), + }]; + }).sort((left, right) => right.value - left.value).slice(0, 5); + const title = uptime ? "Longest-running Runtimes" : "Token consumption"; + return ( +
+
+

{title}

{uptime ? "Current incarnation uptime" : "Session-reported total usage"}

+ {uptime ?
+ +
+ ); +} + +function RuntimeVisuals({ rows }: { rows: RuntimeDashboardRow[] }) { + return ( +
+ + + + +
+ ); } function runtimeModeLabel(row: RuntimeDashboardRow): string { @@ -79,20 +247,6 @@ function runtimeModeLabel(row: RuntimeDashboardRow): string { return dashboardEnvironmentLabel(row.session.environmentProfile); } -function observationTimestamp(row: RuntimeDashboardRow, loadedAt: number): number | null { - const timestamp = row.observation.status === "observed" - ? row.observation.observed_at - : row.observation.resolved_at; - if (!Number.isSafeInteger(timestamp) || timestamp < 0 || timestamp > Math.floor(loadedAt / 1_000)) return null; - return timestamp; -} - -function healthDetail(row: RuntimeDashboardRow, loadedAt: number): string { - const status = runtimeObservationStatusLabel(row.observation); - const timestamp = observationTimestamp(row, loadedAt); - return timestamp === null ? status : `${status} · ${formatDashboardTimestamp(timestamp)}`; -} - function SortHeader({ label, sorted, @@ -233,14 +387,7 @@ function RuntimeTargets({ const visibleRows = table.getFilteredRowModel().rows.length; return ( -
-
-
-

Runtime targets

-

Read-only Session navigation · missing measurements remain unknown

-
- {visibleRows.toLocaleString("en-US")} visible -
+
+ {visibleRows.toLocaleString("en-US")} visible
@@ -314,36 +462,20 @@ export function RuntimeObservabilityContent({ <>
} label="Active Runtimes" value={summary.observedRuntimeCount.toLocaleString("en-US")} detail={`${summary.managedRuntimeCount} managed · ${summary.unavailableRuntimeCount} unavailable`} /> - } label="CPU time / capacity" value={summary.cpuUsageSecondsTotal === null && summary.cpuCapacityCores === null ? "No current sample" : `${formatDashboardDuration(summary.cpuUsageSecondsTotal)} / ${summary.cpuCapacityCores?.toLocaleString("en-US") ?? "—"} cores`} detail={`${summary.cpuCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report CPU`} /> - } label="Memory" value={summary.memoryUsageBytes === null && summary.memoryLimitBytes === null ? "No current sample" : `${formatDashboardBytes(summary.memoryUsageBytes)} / ${formatDashboardBytes(summary.memoryLimitBytes)}`} detail={`${summary.memoryCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report memory`} /> + } label="Cumulative CPU / capacity" value={summary.cpuUsageSecondsTotal === null && summary.cpuCapacityCores === null ? "No current sample" : `${formatDashboardDuration(summary.cpuUsageSecondsTotal)} / ${summary.cpuCapacityCores?.toLocaleString("en-US") ?? "—"} cores`} detail={`${summary.cpuCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report CPU time`} /> + } label="Memory now" value={summary.memoryUsageBytes === null && summary.memoryLimitBytes === null ? "No current sample" : `${formatDashboardBytes(summary.memoryUsageBytes)} / ${formatDashboardBytes(summary.memoryLimitBytes)}`} detail={`${summary.memoryCoverageCount}/${summary.observedRuntimeCount} observed Runtimes report usage`} /> } label="Reported tokens" value={formatDashboardTokens(summary.totalTokens)} detail={`${summary.tokenCoverageCount}/${summary.sessionCount} Sessions report usage`} />
-
-
-

Runtime health

Status and observation freshness by Session

{summary.sessionCount} contexts
-
- {model.rows.slice(0, 8).map((row) => ( - - ))} - {model.rows.length > 8 ? +{model.rows.length - 8} more : null} -
-
-
-

Usage coverage

Unknown remains unknown, never zero

{stale ? "retained snapshot" : "current snapshot"}
-
-
CPU {summary.cpuCoverageCount}/{summary.observedRuntimeCount}{coverage(summary.cpuCoverageCount, summary.observedRuntimeCount)}
-
Memory {summary.memoryCoverageCount}/{summary.observedRuntimeCount}{coverage(summary.memoryCoverageCount, summary.observedRuntimeCount)}
-
Tokens {summary.tokenCoverageCount}/{summary.sessionCount}{coverage(summary.tokenCoverageCount, summary.sessionCount)}
-
Unavailable {summary.unavailableRuntimeCount}current observations
-
Unsupported {summary.unsupportedRuntimeCount}self-hosted or none
-
-
-
+ - +
+ + Runtime targetsSearch and inspect exact observations · unknown remains unknown, never zero + {model.rows.length.toLocaleString("en-US")} targets · {stale ? "retained snapshot" : "current snapshot"} + + +
); } From c2fce9eedc2bca8695ec9f74e061245a647dbce0 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 00:17:42 +0800 Subject: [PATCH 11/51] Fix microsandbox cumulative CPU observation --- .../tools/microsandbox-provider/backend.go | 31 ++++++++++++- .../microsandbox-provider/metrics_test.go | 46 +++++++++++++++++++ 2 files changed, 76 insertions(+), 1 deletion(-) diff --git a/services/agents-api/tools/microsandbox-provider/backend.go b/services/agents-api/tools/microsandbox-provider/backend.go index 8141db349..5168b1e78 100644 --- a/services/agents-api/tools/microsandbox-provider/backend.go +++ b/services/agents-api/tools/microsandbox-provider/backend.go @@ -16,6 +16,11 @@ const bootstrapLabel = "io.parsar.bootstrap" type backend struct{ q wire.Request } +type liveMetricsSource interface { + Metrics(context.Context) (*sdk.Metrics, error) + Detach(context.Context) error +} + func (b backend) run(ctx context.Context) (wire.Response, error) { switch b.q.Operation { case "create": @@ -75,7 +80,13 @@ func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, er if state.Status != string(sdk.SandboxStatusRunning) && state.Status != "draining" { return nil, sandbox.ErrNotFound } - metrics, err := h.Metrics(metricsCtx) + // v0.7.2's name-based handle metrics omit cumulative vCPU time. Connect to + // the already-running, identity-qualified instance so CPU usage remains a + // monotonic counter that Core can safely derive rates from. Connect never + // starts stopped compute, and Detach leaves the runtime lifecycle unchanged. + metrics, err := observeConnectedMetrics(metricsCtx, func(ctx context.Context) (liveMetricsSource, error) { + return h.Connect(ctx) + }) if err != nil { return nil, err } @@ -86,6 +97,24 @@ func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, er return projected, nil } +func observeConnectedMetrics(ctx context.Context, connect func(context.Context) (liveMetricsSource, error)) (*sdk.Metrics, error) { + live, err := connect(ctx) + if err != nil { + return nil, err + } + metrics, metricsErr := live.Metrics(ctx) + detachCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + detachErr := live.Detach(detachCtx) + cancel() + if metricsErr != nil { + return nil, metricsErr + } + if detachErr != nil { + return nil, detachErr + } + return metrics, nil +} + func projectMetrics(metrics *sdk.Metrics, observedAt time.Time) *wire.Metrics { if metrics == nil { return nil diff --git a/services/agents-api/tools/microsandbox-provider/metrics_test.go b/services/agents-api/tools/microsandbox-provider/metrics_test.go index 0355f177e..95dc25644 100644 --- a/services/agents-api/tools/microsandbox-provider/metrics_test.go +++ b/services/agents-api/tools/microsandbox-provider/metrics_test.go @@ -3,12 +3,58 @@ package main import ( + "context" + "errors" "testing" "time" sdk "github.com/superradcompany/microsandbox/sdk/go" ) +type fakeLiveMetrics struct { + metrics *sdk.Metrics + metricsErr error + detachErr error + metricsCalled bool + detachCalled bool + detachBounded bool +} + +func (f *fakeLiveMetrics) Metrics(context.Context) (*sdk.Metrics, error) { + f.metricsCalled = true + return f.metrics, f.metricsErr +} + +func (f *fakeLiveMetrics) Detach(ctx context.Context) error { + f.detachCalled = true + _, f.detachBounded = ctx.Deadline() + return f.detachErr +} + +func TestObserveConnectedMetricsSamplesLiveInstanceAndDetaches(t *testing.T) { + want := &sdk.Metrics{VCPUTimeNs: 123} + live := &fakeLiveMetrics{metrics: want} + connected := false + got, err := observeConnectedMetrics(context.Background(), func(context.Context) (liveMetricsSource, error) { + connected = true + return live, nil + }) + if err != nil || got != want || !connected || !live.metricsCalled || !live.detachCalled || !live.detachBounded { + t.Fatalf("unexpected observation: metrics=%+v err=%v connected=%t live=%+v", got, err, connected, live) + } +} + +func TestObserveConnectedMetricsPropagatesDetachFailure(t *testing.T) { + want := errors.New("detach failed") + live := &fakeLiveMetrics{metrics: &sdk.Metrics{VCPUTimeNs: 123}, detachErr: want} + got, err := observeConnectedMetrics(context.Background(), func(context.Context) (liveMetricsSource, error) { + return live, nil + }) + if got != nil || !errors.Is(err, want) || !live.detachCalled || !live.detachBounded { + t.Fatalf("detach failure not propagated: metrics=%+v err=%v live=%+v", got, err, live) + } +} + func TestProjectMetricsKeepsOnlyCumulativeCommonFields(t *testing.T) { observed := time.Date(2026, 9, 22, 12, 0, 0, 0, time.UTC) projected := projectMetrics(&sdk.Metrics{ From ff8d490ee70e176ca2a33efecf3688327260a864 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 00:38:36 +0800 Subject: [PATCH 12/51] Add Runtime live-window trend charts --- apps/web/e2e/agents-lifecycle.spec.ts | 65 ++++-- .../src/features/dashboard/DashboardView.css | 206 ++++++++++++++++++ .../features/dashboard/DashboardView.test.tsx | 26 +-- .../src/features/dashboard/DashboardView.tsx | 2 +- .../dashboard/RuntimeObservabilityContent.tsx | 184 +--------------- .../dashboard/RuntimeTrendCharts.test.tsx | 33 +++ .../features/dashboard/RuntimeTrendCharts.tsx | 194 +++++++++++++++++ .../features/dashboard/runtime-trends.test.ts | 147 +++++++++++++ .../src/features/dashboard/runtime-trends.ts | 180 +++++++++++++++ .../runtime-observability-design.md | 16 +- docs/web/README.md | 12 +- docs/web/README.zh-CN.md | 9 +- 12 files changed, 851 insertions(+), 223 deletions(-) create mode 100644 apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx create mode 100644 apps/web/src/features/dashboard/RuntimeTrendCharts.tsx create mode 100644 apps/web/src/features/dashboard/runtime-trends.test.ts create mode 100644 apps/web/src/features/dashboard/runtime-trends.ts diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 1ab51a13a..8dd744dd4 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2458,12 +2458,18 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand created_at: baseline - 9_000, last_active_at: baseline - 20, }; - const list = (data: unknown[]) => ({ + const runtimeSessions = Array.from({ length: 12 }, (_, index) => ({ + ...runtimeSession, + id: index === 0 ? sessionId : `11111111-1111-4111-8111-${String(index + 1).padStart(12, "0")}`, + metadata: { title: `Repository migration ${index + 1}` }, + usage: { ...runtimeSession.usage, input_tokens: 420_000 + index, total_tokens: 486_000 + index }, + })); + const list = (data: Array<{ id: string }>) => ({ object: "list", data, has_more: false, - first_id: data.length > 0 ? sessionId : null, - last_id: data.length > 0 ? sessionId : null, + first_id: data[0]?.id ?? null, + last_id: data.at(-1)?.id ?? null, }); await page.goto("/"); @@ -2472,39 +2478,52 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await page.route("**/v1/agents/sessions*", async (route) => { if (new URL(route.request().url()).pathname !== "/v1/agents/sessions") return route.fallback(); - await route.fulfill({ status: 200, contentType: "application/json", body: JSON.stringify(list([runtimeSession])) }); + await route.fulfill({ status: 200, contentType: "application/json", body: JSON.stringify(list(runtimeSessions)) }); }); await page.route("**/v1/agents/runtime-observations*", async (route) => { await route.fulfill({ status: 200, contentType: "application/json", - body: JSON.stringify(list([{ - id: sessionId, + body: JSON.stringify(list(runtimeSessions.map((runtime, index) => ({ + id: runtime.id, object: "agent.runtime_observation", - session_id: sessionId, + session_id: runtime.id, environment_id: environmentId, mode: "openai_hosted", provider_type: "docker", - instance: { kind: "managed_allocation", allocation_id: allocationId, device_id: null, connection_generation: null }, + instance: { + kind: "managed_allocation", + allocation_id: index === 0 ? allocationId : `33333333-3333-4333-8333-${String(index + 1).padStart(12, "0")}`, + device_id: null, + connection_generation: null, + }, status: "observed", reason: null, allocation_created_at: baseline - 8_500, resolved_at: baseline, observed_at: baseline - 1, started_at: baseline - 8_100, - cpu: { usage_seconds_total: 7_350, capacity_cores: 2, usage_cores: 0.72, utilization_ratio: 0.36 }, + cpu: { usage_seconds_total: 7_350 + index, capacity_cores: 2, usage_cores: 0.72 + index / 100, utilization_ratio: 0.36 + index / 200 }, memory: { usage_bytes: 1_288_490_188, limit_bytes: 2_147_483_648 }, - }])), + })))), }); }); - await dashboard.getByRole("button", { name: "Refresh Dashboard snapshot" }).click(); - await expect(dashboard.getByRole("heading", { name: "Runtime health" })).toBeVisible(); - await expect(dashboard.getByRole("heading", { name: "Resource load" })).toBeVisible(); - await expect(dashboard.getByRole("heading", { name: "Longest-running Runtimes" })).toBeVisible(); - await expect(dashboard.getByRole("heading", { name: "Token consumption" })).toBeVisible(); - await expect(dashboard.getByLabel("CPU now: 36%")).toBeVisible(); - await expect(dashboard.getByLabel("Memory now: 60%")).toBeVisible(); + const refresh = dashboard.getByRole("button", { name: "Refresh Dashboard snapshot" }); + await refresh.click(); + await expect(dashboard.getByRole("heading", { name: "CPU usage" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Memory usage" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Compute uptime" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Token throughput" })).toBeVisible(); + await page.waitForTimeout(20); + await refresh.click(); + await page.waitForTimeout(20); + await refresh.click(); + await expect(dashboard.getByLabel("CPU usage: 3 live samples")).toBeVisible(); + await expect(dashboard.getByLabel("Memory usage: 3 live samples")).toBeVisible(); + await expect(dashboard.getByLabel("Compute uptime: 3 live samples")).toBeVisible(); + await expect(dashboard.getByLabel("Token throughput: 3 live samples")).toBeVisible(); + await expect(dashboard).not.toContainText("Collecting live samples"); await expect(dashboard.getByRole("table", { name: "Runtime targets" })).not.toBeVisible(); for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); await expect(page.getByRole("button", { name: "Close notification" })).toHaveCount(0); @@ -2512,9 +2531,19 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await dashboard.getByText("Runtime targets", { exact: true }).click(); await expect(dashboard.getByRole("table", { name: "Runtime targets" })).toBeVisible(); + await expect(dashboard.getByText("Page 1 of 2")).toBeVisible(); + const runtimeSearch = dashboard.getByPlaceholder("Search Session, provider, or identity"); + await runtimeSearch.fill("Repository migration 12"); + await expect(dashboard.getByText("1 visible")).toBeVisible(); + await expect(dashboard.getByRole("table", { name: "Runtime targets" }).getByRole("row", { name: /Repository migration 12/ })).toBeVisible(); + await runtimeSearch.fill(""); + await dashboard.getByRole("button", { name: "Tokens" }).click(); + await expect(dashboard.getByRole("columnheader", { name: "Tokens" })).toHaveAttribute("aria-sort", "descending"); + await dashboard.getByRole("button", { name: "Next" }).click(); + await expect(dashboard.getByText("Page 2 of 2")).toBeVisible(); await page.setViewportSize({ width: 390, height: 844 }); - await expect(dashboard.getByRole("heading", { name: "Resource load" })).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Memory usage" })).toBeVisible(); for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); const widths = await dashboard.locator(".dashboard-runtime-panel").evaluate((element) => ({ viewport: innerWidth, diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index c582b05d3..48426170d 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -982,6 +982,192 @@ border-radius: 6px; } +.dashboard-runtime-trend-grid { + display: grid; + padding: 14px; + grid-template-columns: repeat(2, minmax(0, 1fr)); + gap: 12px; + background: color-mix(in srgb, var(--surface-subtle) 68%, var(--surface)); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-trend-card { + min-width: 0; + padding: 14px 15px 10px; + background: var(--surface); + border: 1px solid var(--line); + border-radius: 9px; +} + +.dashboard-runtime-trend-card > header { + display: flex; + min-height: 48px; + align-items: flex-start; + justify-content: space-between; + gap: 12px; +} + +.dashboard-runtime-trend-card h3, +.dashboard-runtime-trend-card p { + margin: 0; +} + +.dashboard-runtime-trend-card h3 { + font-size: 12px; + font-weight: 650; + line-height: 17px; +} + +.dashboard-runtime-trend-card p { + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + line-height: 14px; +} + +.dashboard-runtime-trend-legend { + display: flex; + max-width: 58%; + flex-wrap: wrap; + justify-content: flex-end; + gap: 5px 10px; + color: var(--fg-muted); + font-size: 9px; +} + +.dashboard-runtime-trend-legend span { + display: inline-flex; + min-width: 0; + align-items: center; + gap: 5px; +} + +.dashboard-runtime-trend-legend i { + width: 11px; + height: 2px; + flex: 0 0 auto; + background: var(--fg-muted); +} + +.dashboard-runtime-trend-orange, +.dashboard-runtime-trend-stroke-orange { + color: #f59e52; + stroke: #f59e52; + background: #f59e52 !important; +} + +.dashboard-runtime-trend-green, +.dashboard-runtime-trend-stroke-green { + color: #50d5a0; + stroke: #50d5a0; + background: #50d5a0 !important; +} + +.dashboard-runtime-trend-blue, +.dashboard-runtime-trend-stroke-blue { + color: #78a7ff; + stroke: #78a7ff; + background: #78a7ff !important; +} + +.dashboard-runtime-trend-purple, +.dashboard-runtime-trend-stroke-purple { + color: #b998f4; + stroke: #b998f4; + background: #b998f4 !important; +} + +.dashboard-runtime-chart-frame { + min-height: 220px; + position: relative; +} + +.dashboard-runtime-chart-frame svg { + display: block; + width: 100%; + height: 220px; + overflow: visible; +} + +.dashboard-runtime-trend-gridline { + stroke: color-mix(in srgb, var(--line) 88%, transparent); + stroke-width: 1; + vector-effect: non-scaling-stroke; +} + +.dashboard-runtime-trend-axis { + fill: var(--fg-muted); + font-family: var(--font-mono); + font-size: 8px; +} + +.dashboard-runtime-trend-band { + opacity: .07; +} + +.dashboard-runtime-trend-band-safe { + fill: var(--success); +} + +.dashboard-runtime-trend-band-warning { + fill: var(--warning); +} + +.dashboard-runtime-trend-band-danger { + fill: var(--danger); +} + +.dashboard-runtime-trend-line { + fill: none; + stroke-width: 2.25; + stroke-linecap: round; + stroke-linejoin: round; + vector-effect: non-scaling-stroke; +} + +.dashboard-runtime-trend-dot { + fill: currentColor; + stroke-width: 0; +} + +.dashboard-runtime-chart-collecting { + display: grid; + padding: 9px 12px; + position: absolute; + right: 18px; + bottom: 29px; + gap: 1px; + color: var(--fg); + background: color-mix(in srgb, var(--surface) 91%, transparent); + border: 1px solid var(--line); + border-radius: 6px; + box-shadow: var(--shadow-control); + pointer-events: none; +} + +.dashboard-runtime-chart-collecting strong { + font-size: 9px; + line-height: 13px; +} + +.dashboard-runtime-chart-collecting span { + color: var(--fg-muted); + font-size: 8px; + line-height: 12px; +} + +.dashboard-runtime-trend-accessible { + position: absolute; + width: 1px; + height: 1px; + padding: 0; + margin: -1px; + overflow: hidden; + clip: rect(0, 0, 0, 0); + white-space: nowrap; + border: 0; +} + .dashboard-runtime-explorer { min-width: 0; } @@ -1609,6 +1795,10 @@ grid-template-columns: 1fr; } + .dashboard-runtime-trend-grid { + grid-template-columns: 1fr; + } + .dashboard-runtime-search { width: 100%; flex-basis: 100%; @@ -1680,6 +1870,22 @@ gap: 10px; } + .dashboard-runtime-trend-grid { + padding: 10px; + gap: 10px; + } + + .dashboard-runtime-trend-card > header { + align-items: flex-start; + flex-direction: column; + gap: 5px; + } + + .dashboard-runtime-trend-legend { + max-width: 100%; + justify-content: flex-start; + } + .dashboard-runtime-visual-card { min-height: 224px; } diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index ea2b4c1cd..9e3c8c4df 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -233,7 +233,7 @@ describe("Dashboard loaded-result presentation", () => { resolved_at: 1_700_000_100, observed_at: 1_700_000_090, started_at: 1_700_000_010, - cpu: { usage_seconds_total: 73.5, capacity_cores: 2, usage_cores: null, utilization_ratio: null }, + cpu: { usage_seconds_total: 73.5, capacity_cores: 2, usage_cores: 3, utilization_ratio: 1.5 }, memory: { usage_bytes: 536_870_912, limit_bytes: 2_147_483_648 }, }; const html = render({ @@ -245,17 +245,17 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("Cumulative CPU / capacity"); expect(html).toContain("1m 13s / 2 cores"); expect(html).toContain("512 MiB / 2.00 GiB"); - expect(html).toContain('aria-label="Runtime snapshot visualizations"'); - expect(html).toContain("Runtime health"); - expect(html).toContain('aria-label="1 of 1 Runtimes observed"'); - expect(html).toContain("Resource load"); - expect(html).toContain("CPU now"); - expect(html).toContain("Instantaneous CPU is unavailable; cumulative CPU time remains in Explorer"); - expect(html).toContain('aria-label="CPU now: Unavailable"'); - expect(html).toContain('aria-label="Memory now: 25%"'); - expect(html).toContain("Longest-running Runtimes"); - expect(html).toContain("1m 20s"); - expect(html).toContain("Token consumption"); + expect(html).toContain('aria-label="Runtime live-window charts"'); + expect(html).toContain("CPU usage"); + expect(html).toContain("Memory usage"); + expect(html).toContain("Compute uptime"); + expect(html).toContain("Token throughput"); + expect(html).toContain("150%"); + expect(html).toContain("Collecting live samples"); + expect(html).toContain("1/2 minimum · no history is synthesized"); + expect(html).toContain("CPU usage collecting live samples; 1 of 2 minimum"); + expect(html).toContain("Latest value"); + expect(html).toContain("Missing samples"); expect(html).toContain("Runtime targets"); expect(html).toContain('
'); expect(html).toContain("Search Runtime targets"); @@ -267,7 +267,7 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("Identity"); expect(html).toContain("unknown remains unknown, never zero"); expect(html).not.toContain("CPU %"); - expect(html).not.toContain("CPU now: 25%"); + expect(html).not.toContain("CPU now"); expect(html).not.toContain("historical chart"); }); diff --git a/apps/web/src/features/dashboard/DashboardView.tsx b/apps/web/src/features/dashboard/DashboardView.tsx index 4fd6cc68c..6f23d2151 100644 --- a/apps/web/src/features/dashboard/DashboardView.tsx +++ b/apps/web/src/features/dashboard/DashboardView.tsx @@ -382,7 +382,7 @@ export function DashboardView({

Runtime monitoring

-

Current provider samples joined to an exact complete Session snapshot · no historical series

+

Current provider samples · browser-local live window · no durable history

{runtimeModel ? ( diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index eb88e0e9c..9e59b7747 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -1,18 +1,16 @@ import { - Activity, ChevronDown, ChevronLeft, ChevronRight, ChevronsUpDown, ChevronUp, - Clock3, Cpu, Gauge, MemoryStick, Search, Server, } from "lucide-react"; -import { useMemo, useState, type ReactNode } from "react"; +import { useEffect, useMemo, useState, type ReactNode } from "react"; import { flexRender, getCoreRowModel, @@ -37,6 +35,8 @@ import { type RuntimeDashboardRow, } from "./dashboard-model"; import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; +import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; +import { appendRuntimeTrendSample, type RuntimeTrendSample } from "./runtime-trends"; const PAGE_SIZE = 10; @@ -68,177 +68,6 @@ function percent(usage: number | null | undefined, limit: number | null | undefi return Math.min(100, Math.max(0, usage / limit * 100)); } -function finiteNonNegative(value: unknown): number | null { - return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; -} - -function meterPercent(value: number | null): number | null { - return value === null ? null : Math.min(100, Math.max(0, value * 100)); -} - -function formatRatio(value: number | null): string { - return value === null ? "Unavailable" : `${(value * 100).toFixed(value >= 0.1 ? 0 : 1)}%`; -} - -function RuntimeHealthChart({ rows }: { rows: RuntimeDashboardRow[] }) { - const counts = { - observed: rows.filter((row) => row.observation.status === "observed").length, - unavailable: rows.filter((row) => row.observation.status === "unavailable").length, - unsupported: rows.filter((row) => row.observation.status === "unsupported").length, - }; - const total = rows.length; - const segments = [ - { key: "observed", label: "Observed", count: counts.observed }, - { key: "unavailable", label: "Unavailable", count: counts.unavailable }, - { key: "unsupported", label: "Unsupported", count: counts.unsupported }, - ]; - let offset = 0; - - return ( -
-
-

Runtime health

Current observation status

-
-
-
- - - {segments.map((segment) => { - const size = total === 0 ? 0 : segment.count / total * 100; - const circle = ; - offset += size; - return circle; - })} - - {counts.observed}of {total} -
-
- {segments.map((segment) => ( -
{segment.label}{segment.count}
- ))} -
-
-
- ); -} - -function ResourceMeter({ - label, - value, - detail, - ratio, -}: { - label: string; - value: string; - detail: string; - ratio: number | null; -}) { - const width = meterPercent(ratio); - return ( -
-
{label}{value}
-
- {width === null ? : } -
- {detail} -
- ); -} - -function RuntimeResourceChart({ rows }: { rows: RuntimeDashboardRow[] }) { - const observed = rows.filter((row) => row.observation.status === "observed"); - const cpuSamples = observed.flatMap((row) => { - const usage = finiteNonNegative(row.observation.cpu?.usage_cores); - const capacity = finiteNonNegative(row.observation.cpu?.capacity_cores); - return usage !== null && capacity !== null && capacity > 0 ? [{ usage, capacity }] : []; - }); - const memorySamples = observed.flatMap((row) => { - const usage = finiteNonNegative(row.observation.memory?.usage_bytes); - const limit = finiteNonNegative(row.observation.memory?.limit_bytes); - return usage !== null && limit !== null && limit > 0 ? [{ usage, limit }] : []; - }); - const cpuUsage = cpuSamples.reduce((total, sample) => total + sample.usage, 0); - const cpuCapacity = cpuSamples.reduce((total, sample) => total + sample.capacity, 0); - const memoryUsage = memorySamples.reduce((total, sample) => total + sample.usage, 0); - const memoryLimit = memorySamples.reduce((total, sample) => total + sample.limit, 0); - const cpuRatio = cpuSamples.length > 0 && cpuCapacity > 0 ? cpuUsage / cpuCapacity : null; - const memoryRatio = memorySamples.length > 0 && memoryLimit > 0 ? memoryUsage / memoryLimit : null; - - return ( -
-
-

Resource load

Point-in-time provider samples

- {observed.length} observed -
-
- - -
-
- ); -} - -interface RankedBarItem { - id: string; - label: string; - detail: string; - value: number; - formatted: string; -} - -function RankedBars({ items, empty, tone }: { items: RankedBarItem[]; empty: string; tone: "uptime" | "tokens" }) { - const maximum = Math.max(0, ...items.map((item) => item.value)); - if (items.length === 0) return

{empty}

; - return ( -
- {items.map((item) => ( -
-
{item.label}{item.formatted}
-
- {item.detail} -
- ))} -
- ); -} - -function RuntimeRankChart({ rows, kind }: { rows: RuntimeDashboardRow[]; kind: "uptime" | "tokens" }) { - const uptime = kind === "uptime"; - const items = rows.flatMap((row): RankedBarItem[] => { - const value = uptime ? row.computeUptimeSeconds : row.session.totalTokens; - if (value === null || !Number.isFinite(value) || value < 0) return []; - return [{ - id: row.observation.id, - label: row.session.title, - detail: uptime ? runtimeModeLabel(row) : dashboardStatusLabel(row.session.status), - value, - formatted: uptime ? formatDashboardDuration(value) : formatDashboardTokens(value), - }]; - }).sort((left, right) => right.value - left.value).slice(0, 5); - const title = uptime ? "Longest-running Runtimes" : "Token consumption"; - return ( -
-
-

{title}

{uptime ? "Current incarnation uptime" : "Session-reported total usage"}

- {uptime ?
- -
- ); -} - -function RuntimeVisuals({ rows }: { rows: RuntimeDashboardRow[] }) { - return ( -
- - - - -
- ); -} - function runtimeModeLabel(row: RuntimeDashboardRow): string { if (row.observation.mode === "openai_hosted") { const provider = row.observation.provider_type; @@ -457,6 +286,11 @@ export function RuntimeObservabilityContent({ }) { const model = useMemo(() => buildRuntimeDashboardModel(snapshot.sessions, snapshot.observations), [snapshot]); const summary = model.summary; + const [trendSamples, setTrendSamples] = useState(() => appendRuntimeTrendSample([], snapshot)); + + useEffect(() => { + setTrendSamples((current) => appendRuntimeTrendSample(current, snapshot)); + }, [snapshot]); return ( <> @@ -467,7 +301,7 @@ export function RuntimeObservabilityContent({ } label="Reported tokens" value={formatDashboardTokens(summary.totalTokens)} detail={`${summary.tokenCoverageCount}/${summary.sessionCount} Sessions report usage`} /> - +
diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx new file mode 100644 index 000000000..3f73ca7b1 --- /dev/null +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -0,0 +1,33 @@ +import { renderToStaticMarkup } from "react-dom/server"; +import { describe, expect, it } from "vitest"; + +import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; +import type { RuntimeTrendSample } from "./runtime-trends"; + +function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample { + return { + sampledAt, + targets: [{ + sessionId: "session-1", + label: "Runtime worker", + cpuRatio, + uptimeSeconds: 120, + }], + memoryUsageBytes: 512, + memoryLimitBytes: 1_024, + tokenTotals: [], + inputTokensPerMinute: null, + outputTokensPerMinute: null, + }; +} + +describe("Runtime live-window chart accessibility", () => { + it("reports a current gap as unavailable instead of announcing a stale value as latest", () => { + const html = renderToStaticMarkup( + , + ); + + expect(html).toContain("Runtime worker
"); + expect(html).not.toContain("Runtime worker"); + }); +}); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx new file mode 100644 index 000000000..daf0277ba --- /dev/null +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -0,0 +1,194 @@ +import { useMemo } from "react"; + +import { formatDashboardBytes, formatDashboardDuration, formatDashboardTokens } from "./dashboard-model"; +import { tokenThroughput, type RuntimeTrendSample } from "./runtime-trends"; + +interface TrendPoint { + sampledAt: number; + value: number | null; +} + +interface TrendSeries { + id: string; + label: string; + tone: "orange" | "green" | "blue" | "purple"; + points: TrendPoint[]; +} + +interface TrendBand { + from: number; + to: number; + tone: "safe" | "warning" | "danger"; +} + +const WIDTH = 640; +const HEIGHT = 220; +const PLOT = { left: 52, right: 16, top: 22, bottom: 34 }; + +function finite(values: readonly (number | null)[]): number[] { + return values.filter((value): value is number => value !== null && Number.isFinite(value) && value >= 0); +} + +function timeLabel(value: number): string { + return new Date(value).toLocaleTimeString([], { hour: "2-digit", minute: "2-digit", second: "2-digit" }); +} + +function segments(points: readonly TrendPoint[]): TrendPoint[][] { + const result: TrendPoint[][] = []; + let current: TrendPoint[] = []; + for (const point of points) { + if (point.value === null || !Number.isFinite(point.value)) { + if (current.length > 0) result.push(current); + current = []; + } else current.push(point); + } + if (current.length > 0) result.push(current); + return result; +} + +function TrendChart({ + title, + subtitle, + samples, + series, + maximum, + formatValue, + bands = [], +}: { + title: string; + subtitle: string; + samples: readonly RuntimeTrendSample[]; + series: TrendSeries[]; + maximum: number; + formatValue: (value: number) => string; + bands?: TrendBand[]; +}) { + const timestamps = samples.map((sample) => sample.sampledAt); + const start = timestamps[0] ?? 0; + const end = timestamps.at(-1) ?? start; + const range = Math.max(1, end - start); + const plotWidth = WIDTH - PLOT.left - PLOT.right; + const plotHeight = HEIGHT - PLOT.top - PLOT.bottom; + const yMaximum = Math.max(1, maximum); + const x = (sampledAt: number) => PLOT.left + (sampledAt - start) / range * plotWidth; + const y = (value: number) => PLOT.top + (1 - Math.min(yMaximum, Math.max(0, value)) / yMaximum) * plotHeight; + const hasLine = series.some((entry) => segments(entry.points).some((segment) => segment.length >= 2)); + const ticks = [1, .66, .33, 0]; + const xTicks = [start, start + range / 2, end]; + + return ( +
+
+

{title}

{subtitle}

+
+ {series.map((entry) => {entry.label})} +
+
+
+ + {bands.map((band) => ( + + ))} + {ticks.map((tick) => { + const value = yMaximum * tick; + return ( + + + {formatValue(value)} + + ); + })} + {samples.length > 0 ? xTicks.map((tick, index) => ( + {timeLabel(tick)} + )) : null} + {series.flatMap((entry) => segments(entry.points).map((segment, index) => { + const points = segment.map((point) => `${x(point.sampledAt)},${y(point.value ?? 0)}`).join(" "); + return segment.length === 1 + ? + : ; + }))} + + {!hasLine ?
Collecting live samples{samples.length}/2 minimum · no history is synthesized
: null} +
+
Unavailable150%1
+ + + + {series.map((entry) => { + const latest = entry.points.at(-1)?.value ?? null; + const missing = entry.points.filter((point) => point.value === null).length; + return ; + })} + +
{hasLine ? `${title} live trend available` : `${title} collecting live samples; ${samples.length} of 2 minimum`}
SeriesLatest valueMissing samples
{entry.label}{latest === null ? "Unavailable" : formatValue(latest)}{missing}
+
+ ); +} + +function targetIds(samples: readonly RuntimeTrendSample[], field: "cpuRatio" | "uptimeSeconds"): string[] { + const latest = new Map(); + for (const sample of samples) { + for (const target of sample.targets) { + const value = target[field]; + if (value !== null) latest.set(target.sessionId, value); + } + } + return [...latest.entries()].sort((left, right) => right[1] - left[1]).slice(0, 3).map(([id]) => id); +} + +function targetLabel(samples: readonly RuntimeTrendSample[], id: string): string { + for (let index = samples.length - 1; index >= 0; index -= 1) { + const target = samples[index]?.targets.find((candidate) => candidate.sessionId === id); + if (target) return target.label; + } + return "Runtime"; +} + +const tones: TrendSeries["tone"][] = ["orange", "green", "blue"]; + +export function RuntimeTrendCharts({ samples }: { samples: readonly RuntimeTrendSample[] }) { + const charts = useMemo(() => { + const cpuIds = targetIds(samples, "cpuRatio"); + const uptimeIds = targetIds(samples, "uptimeSeconds"); + const cpu = cpuIds.map((id, index): TrendSeries => ({ + id, + label: targetLabel(samples, id), + tone: tones[index] ?? "blue", + points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.sessionId === id)?.cpuRatio ?? null })).map((point) => ({ ...point, value: point.value === null ? null : point.value * 100 })), + })); + const memoryUsed = samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.memoryUsageBytes })); + const memoryLimit = samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.memoryLimitBytes })); + const uptime = uptimeIds.map((id, index): TrendSeries => ({ + id, + label: targetLabel(samples, id), + tone: tones[(index + 2) % tones.length] ?? "blue", + points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.sessionId === id)?.uptimeSeconds ?? null })), + })); + const throughput = tokenThroughput(samples); + return { + cpu, + memory: [ + { id: "used", label: "used", tone: "purple", points: memoryUsed }, + { id: "limit", label: "configured limit", tone: "green", points: memoryLimit }, + ] satisfies TrendSeries[], + uptime, + tokens: [ + { id: "input", label: "input", tone: "orange", points: throughput.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.inputPerMinute })) }, + { id: "output", label: "output", tone: "green", points: throughput.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.outputPerMinute })) }, + ] satisfies TrendSeries[], + }; + }, [samples]); + const cpuMaximum = Math.max(100, ...finite(charts.cpu.flatMap((series) => series.points.map((point) => point.value)))); + const memoryMaximum = Math.max(1, ...finite(charts.memory.flatMap((series) => series.points.map((point) => point.value)))); + const uptimeMaximum = Math.max(1, ...finite(charts.uptime.flatMap((series) => series.points.map((point) => point.value)))); + const tokenMaximum = Math.max(1, ...finite(charts.tokens.flatMap((series) => series.points.map((point) => point.value)))); + + return ( +
+ `${Math.round(value)}%`} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} /> + formatDashboardBytes(Math.round(value))} /> + formatDashboardDuration(value)} /> + `${formatDashboardTokens(Math.round(value))}/min`} /> +
+ ); +} diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts new file mode 100644 index 000000000..a3000c959 --- /dev/null +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -0,0 +1,147 @@ +import { describe, expect, it } from "vitest"; + +import type { AgentSession, RuntimeObservation } from "@agents-core-web/agents-client"; + +import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; +import { + appendRuntimeTrendSample, + RUNTIME_TREND_MAX_TARGETS, + runtimeTrendSample, + tokenThroughput, +} from "./runtime-trends"; + +function snapshot(at: number, options: { + input?: number; + output?: number; + cpuRatio?: number | null; + memory?: number; + startedAt?: number; +} = {}): RuntimeDashboardSnapshot { + const sessionId = "11111111-1111-4111-8111-111111111111"; + const observedAt = Math.floor(at / 1_000); + const input = options.input ?? 100; + const output = options.output ?? 20; + const session = { + id: sessionId, + object: "agent.session", + metadata: { title: "Runtime worker" }, + agent: { + id: "agent-1", + name: "Worker", + model: "fixture/model", + instructions: null, + multi_agent: { enabled: false, max_concurrent_subagents: null }, + reasoning: {}, + service_tier: "auto", + text: { format: { type: "text" }, verbosity: "medium" }, + tools: [], + }, + environment: { type: "none" }, + status: "idle", + error: null, + required_actions: [], + vault_ids: [], + usage: { + input_tokens: input, + output_tokens: output, + total_tokens: input + output, + input_tokens_details: { cached_tokens: 0 }, + output_tokens_details: { reasoning_tokens: 0 }, + }, + created_at: observedAt - 180, + last_active_at: observedAt, + } as AgentSession; + const observation = { + id: sessionId, + session_id: sessionId, + status: "observed", + provider_type: "docker", + observed_at: observedAt, + started_at: options.startedAt ?? observedAt - 120, + cpu: { utilization_ratio: options.cpuRatio ?? .25, usage_cores: .5, capacity_cores: 2, usage_seconds_total: 30 }, + memory: { usage_bytes: options.memory ?? 512, limit_bytes: 1_024 }, + } as RuntimeObservation; + return { sessions: [session], observations: [observation], loadedAt: at }; +} + +describe("Runtime live-window trends", () => { + it("projects only honest point-in-time and cumulative Session values", () => { + const sample = runtimeTrendSample(snapshot(120_000)); + expect(sample).toMatchObject({ + sampledAt: 120_000, + tokenTotals: [{ + sessionId: "11111111-1111-4111-8111-111111111111", + inputTokens: 100, + outputTokens: 20, + }], + memoryUsageBytes: 512, + memoryLimitBytes: 1_024, + }); + expect(sample.targets).toEqual([expect.objectContaining({ + label: "Runtime worker", + cpuRatio: .25, + uptimeSeconds: 120, + })]); + }); + + it("deduplicates refreshes and bounds the rolling window", () => { + let samples = appendRuntimeTrendSample([], snapshot(60_000), 120_000, 2); + samples = appendRuntimeTrendSample(samples, snapshot(120_000)); + samples = appendRuntimeTrendSample(samples, snapshot(120_000, { memory: 768 }), 120_000, 2); + samples = appendRuntimeTrendSample(samples, snapshot(180_000), 120_000, 2); + expect(samples.map((sample) => sample.sampledAt)).toEqual([120_000, 180_000]); + expect(samples[0]?.memoryUsageBytes).toBe(768); + }); + + it("derives throughput only between monotonic cumulative samples", () => { + let samples = appendRuntimeTrendSample([], snapshot(60_000, { input: 100, output: 20 })); + samples = appendRuntimeTrendSample(samples, snapshot(120_000, { input: 220, output: 50 })); + samples = appendRuntimeTrendSample(samples, snapshot(180_000, { input: 10, output: 5 })); + expect(tokenThroughput(samples)).toEqual([ + { sampledAt: 60_000, inputPerMinute: null, outputPerMinute: null }, + { sampledAt: 120_000, inputPerMinute: 120, outputPerMinute: 30 }, + { sampledAt: 180_000, inputPerMinute: null, outputPerMinute: null }, + ]); + }); + + it("matches token counters by Session without creating churn spikes", () => { + const first = snapshot(60_000, { input: 100, output: 20 }); + const second = snapshot(120_000, { input: 220, output: 50 }); + second.sessions.push({ + ...second.sessions[0]!, + id: "22222222-2222-4222-8222-222222222222", + usage: { ...second.sessions[0]!.usage!, input_tokens: 5_000, output_tokens: 800, total_tokens: 5_800 }, + }); + const third = snapshot(180_000, { input: 250, output: 10 }); + let samples = appendRuntimeTrendSample([], first); + samples = appendRuntimeTrendSample(samples, second); + samples = appendRuntimeTrendSample(samples, third); + expect(tokenThroughput(samples)).toEqual([ + { sampledAt: 60_000, inputPerMinute: null, outputPerMinute: null }, + { sampledAt: 120_000, inputPerMinute: 120, outputPerMinute: 30 }, + { sampledAt: 180_000, inputPerMinute: 30, outputPerMinute: null }, + ]); + }); + + it("bounds retained target detail and keeps raw token counters only for the latest sample", () => { + const many = snapshot(60_000); + many.sessions = Array.from({ length: 20 }, (_, index) => ({ + ...many.sessions[0]!, + id: `session-${index}`, + metadata: { title: `Runtime ${index}` }, + })); + many.observations = many.sessions.map((session, index) => ({ + ...many.observations[0]!, + id: session.id, + session_id: session.id, + cpu: { ...many.observations[0]!.cpu!, utilization_ratio: index / 10 }, + started_at: 60 - index, + } as RuntimeObservation)); + const later = { ...many, loadedAt: 120_000 }; + let samples = appendRuntimeTrendSample([], many); + samples = appendRuntimeTrendSample(samples, later); + expect(samples.every((sample) => sample.targets.length <= RUNTIME_TREND_MAX_TARGETS)).toBe(true); + expect(samples[0]?.tokenTotals).toEqual([]); + expect(samples[1]?.tokenTotals).toHaveLength(20); + }); +}); diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts new file mode 100644 index 000000000..ebc1c6295 --- /dev/null +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -0,0 +1,180 @@ +import type { AgentSession, RuntimeObservation } from "@agents-core-web/agents-client"; + +import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; + +export const RUNTIME_TREND_WINDOW_MS = 60 * 60 * 1_000; +export const RUNTIME_TREND_MAX_SAMPLES = 120; +export const RUNTIME_TREND_SERIES_LIMIT = 3; +export const RUNTIME_TREND_MAX_TARGETS = RUNTIME_TREND_SERIES_LIMIT * 2; + +export interface RuntimeTrendTarget { + sessionId: string; + label: string; + cpuRatio: number | null; + uptimeSeconds: number | null; +} + +export interface RuntimeTrendSample { + sampledAt: number; + targets: RuntimeTrendTarget[]; + memoryUsageBytes: number | null; + memoryLimitBytes: number | null; + tokenTotals: RuntimeTrendTokenTotal[]; + inputTokensPerMinute: number | null; + outputTokensPerMinute: number | null; +} + +export interface RuntimeTrendTokenTotal { + sessionId: string; + inputTokens: number; + outputTokens: number; +} + +export interface TokenThroughputSample { + sampledAt: number; + inputPerMinute: number | null; + outputPerMinute: number | null; +} + +function finiteNonNegative(value: unknown): number | null { + return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; +} + +function safeInteger(value: unknown): number | null { + return typeof value === "number" && Number.isSafeInteger(value) && value >= 0 ? value : null; +} + +function sessionTitle(session: AgentSession): string { + const title = session.metadata?.title; + if (typeof title === "string" && title.length > 0) return title; + return session.agent.name ?? session.agent.model; +} + +function cpuRatio(observation: RuntimeObservation): number | null { + if (observation.status !== "observed") return null; + const reported = finiteNonNegative(observation.cpu?.utilization_ratio); + if (reported !== null) return reported; + const usage = finiteNonNegative(observation.cpu?.usage_cores); + const capacity = finiteNonNegative(observation.cpu?.capacity_cores); + return usage !== null && capacity !== null && capacity > 0 ? usage / capacity : null; +} + +function uptimeSeconds(observation: RuntimeObservation): number | null { + if (observation.status !== "observed") return null; + const startedAt = safeInteger(observation.started_at); + const observedAt = safeInteger(observation.observed_at); + return startedAt !== null && observedAt !== null && observedAt >= startedAt + ? observedAt - startedAt + : null; +} + +function tokenTotals(sessions: readonly AgentSession[]): RuntimeTrendTokenTotal[] { + return sessions.flatMap((session): RuntimeTrendTokenTotal[] => { + const inputTokens = safeInteger(session.usage?.input_tokens); + const outputTokens = safeInteger(session.usage?.output_tokens); + return inputTokens === null || outputTokens === null + ? [] + : [{ sessionId: session.id, inputTokens, outputTokens }]; + }); +} + +export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeTrendSample { + const sessions = new Map(snapshot.sessions.map((session) => [session.id, session])); + const observed = snapshot.observations.flatMap((observation) => { + const session = sessions.get(observation.session_id); + if (!session || observation.status !== "observed") return []; + return [{ + sessionId: observation.session_id, + label: sessionTitle(session), + cpuRatio: cpuRatio(observation), + memoryUsageBytes: finiteNonNegative(observation.memory?.usage_bytes), + memoryLimitBytes: finiteNonNegative(observation.memory?.limit_bytes), + uptimeSeconds: uptimeSeconds(observation), + }]; + }); + const targetIds = new Set([ + ...observed.filter((target) => target.cpuRatio !== null) + .sort((left, right) => (right.cpuRatio ?? 0) - (left.cpuRatio ?? 0)) + .slice(0, RUNTIME_TREND_SERIES_LIMIT) + .map((target) => target.sessionId), + ...observed.filter((target) => target.uptimeSeconds !== null) + .sort((left, right) => (right.uptimeSeconds ?? 0) - (left.uptimeSeconds ?? 0)) + .slice(0, RUNTIME_TREND_SERIES_LIMIT) + .map((target) => target.sessionId), + ]); + const targets = observed.filter((target) => targetIds.has(target.sessionId)).map((target) => ({ + sessionId: target.sessionId, + label: target.label, + cpuRatio: target.cpuRatio, + uptimeSeconds: target.uptimeSeconds, + })); + const pairedMemory = observed.filter((target) => ( + target.memoryUsageBytes !== null && target.memoryLimitBytes !== null + )); + return { + sampledAt: snapshot.loadedAt, + targets, + memoryUsageBytes: pairedMemory.length === 0 + ? null + : pairedMemory.reduce((total, target) => total + (target.memoryUsageBytes ?? 0), 0), + memoryLimitBytes: pairedMemory.length === 0 + ? null + : pairedMemory.reduce((total, target) => total + (target.memoryLimitBytes ?? 0), 0), + tokenTotals: tokenTotals(snapshot.sessions), + inputTokensPerMinute: null, + outputTokensPerMinute: null, + }; +} + +function tokenRate( + previous: RuntimeTrendSample, + next: RuntimeTrendSample, +): Pick { + const elapsedMinutes = (next.sampledAt - previous.sampledAt) / 60_000; + if (elapsedMinutes <= 0) return { inputTokensPerMinute: null, outputTokensPerMinute: null }; + const previousTotals = new Map(previous.tokenTotals.map((total) => [total.sessionId, total])); + const pairs = next.tokenTotals.flatMap((current): Array<[RuntimeTrendTokenTotal, RuntimeTrendTokenTotal]> => { + const prior = previousTotals.get(current.sessionId); + return prior ? [[prior, current]] : []; + }); + const inputDeltas = pairs.map(([prior, current]) => current.inputTokens - prior.inputTokens); + const outputDeltas = pairs.map(([prior, current]) => current.outputTokens - prior.outputTokens); + const inputDelta = inputDeltas.length > 0 && inputDeltas.every((delta) => delta >= 0) + ? inputDeltas.reduce((total, delta) => total + delta, 0) + : null; + const outputDelta = outputDeltas.length > 0 && outputDeltas.every((delta) => delta >= 0) + ? outputDeltas.reduce((total, delta) => total + delta, 0) + : null; + return { + inputTokensPerMinute: inputDelta === null ? null : inputDelta / elapsedMinutes, + outputTokensPerMinute: outputDelta === null ? null : outputDelta / elapsedMinutes, + }; +} + +export function appendRuntimeTrendSample( + current: readonly RuntimeTrendSample[], + snapshot: RuntimeDashboardSnapshot, + windowMs = RUNTIME_TREND_WINDOW_MS, + maximum = RUNTIME_TREND_MAX_SAMPLES, +): RuntimeTrendSample[] { + const next = runtimeTrendSample(snapshot); + const retained = current + .filter((sample) => sample.sampledAt !== next.sampledAt) + .filter((sample) => sample.sampledAt >= next.sampledAt - windowMs && sample.sampledAt <= next.sampledAt) + .sort((left, right) => left.sampledAt - right.sampledAt) + .slice(-(maximum - 1)); + const previous = retained.at(-1); + if (previous) Object.assign(next, tokenRate(previous, next)); + return [ + ...retained.map((sample) => ({ ...sample, tokenTotals: [] })), + next, + ].slice(-maximum); +} + +export function tokenThroughput(samples: readonly RuntimeTrendSample[]): TokenThroughputSample[] { + return samples.map((sample) => ({ + sampledAt: sample.sampledAt, + inputPerMinute: sample.inputTokensPerMinute, + outputPerMinute: sample.outputTokensPerMinute, + })); +} diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 35a8ba55c..3da1a5bcf 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -292,17 +292,19 @@ seams separate: 3. A feature-local state model retains `last_complete`, current refresh status, local filters, and the selected time range. It aborts an overlapping refresh and marks old data stale after a failed or incomplete refresh. -4. Presentational components render summary metrics, a compact health matrix, - explicit usage coverage, the Runtime table, and an identity detail surface. - Trend components and chart dependencies are absent unless a later history - capability and contract are configured. +4. Presentational components render summary metrics, a browser-local live window, + the Runtime table, and an identity detail surface. The live window contains only + complete snapshots collected while this Dashboard instance is mounted; it is + bounded, ephemeral, and never presented as durable operator history. The initial implementation uses a 30-second Web cadence plus up to five seconds of jitter and a 15-second whole-refresh budget. These are Web configuration, not API guarantees. Web pauses periodic reads when hidden, refreshes when visibility returns, and adds jitter so multiple browsers do not synchronize. Filtering is local to the last complete snapshot and never changes tenant authorization or -provider selection. +provider selection. The same complete snapshots feed the one-hour, 120-sample +browser-local live window; reload, navigation, or connection replacement may reset +it, and no point is interpolated or persisted by Core. ## 12. Token usage boundary @@ -367,12 +369,14 @@ this does not add an upstream OpenAI operation. - Publish only complete traversals and retain the previous snapshot on failure. - Join existing Session Usage and Turn status by exact Session ID. - Add responsive, keyboard-accessible current-resource views. +- Build an explicitly ephemeral live window from complete Web snapshots. ### Phase 4: optional history - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. -- Add trend charts only when this capability is advertised. +- Replace or extend the ephemeral live window with explicitly advertised durable + history reads; never merge the two retention semantics implicitly. ### Phase 5: additional sources diff --git a/docs/web/README.md b/docs/web/README.md index 398ccacc1..cbb1d7c83 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -30,12 +30,12 @@ credentials or execution into the browser. ### Dashboard Dashboard is the starting point. It summarizes the current Agent and Session results, -loads a complete tenant-scoped Runtime observation snapshot, shows current Docker -resource evidence and coverage without inventing missing values, highlights Sessions -that need attention, and links directly to Agent creation or a new Session. Runtime -health and coverage cards sit above a searchable, filterable, sortable, paginated -semantic table. Historical charts remain absent until an operator configures a -separate history capability. +loads complete tenant-scoped Runtime observation snapshots, and retains a bounded +browser-local live window for CPU, memory, compute-uptime, and token-throughput +charts without inventing missing values. The searchable, filterable, sortable, +paginated semantic table remains available on demand. This live window starts when +the Dashboard opens and is not durable history; cross-browser retention still +requires a separate operator history capability. ### Agents diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index 77b804710..961a5509e 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -28,10 +28,11 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 ### Dashboard Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载完整的租户级 Runtime -观测快照,并在不把缺失值伪装成 0 的前提下展示当前 Docker 资源和数据覆盖率。页面也 -展示 Runtime health 与 coverage 卡片,以及支持搜索、状态/模式筛选、排序、分页的语义 -表格;同时展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。 -在运维方配置独立历史能力之前,页面不会伪造历史趋势图,也不会引入趋势图组件。 +观测快照,并保留有界的浏览器本地 live window,展示 CPU、内存、compute uptime 和 +token throughput 趋势,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的 +语义表格按需展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 +Session 的流程。live window 从打开 Dashboard 后开始采集,不是跨浏览器持久历史; +持久保留仍需运维方配置独立 history 能力。 ### Agents From 15cccdd92eb41cc297d5eb47e8aff1f6e1d6ec86 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 00:53:47 +0800 Subject: [PATCH 13/51] Fix Runtime observation error assertion --- .../agents-api/internal/api/runtime_observations_test.go | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/services/agents-api/internal/api/runtime_observations_test.go b/services/agents-api/internal/api/runtime_observations_test.go index 29768be94..90dd4ab7f 100644 --- a/services/agents-api/internal/api/runtime_observations_test.go +++ b/services/agents-api/internal/api/runtime_observations_test.go @@ -217,8 +217,10 @@ func TestRuntimeObservationListRejectsWholePageOnIntegrityFailure(t *testing.T) if response.Code != http.StatusInternalServerError { t.Fatalf("integrity failure returned %d: %s", response.Code, response.Body) } - var envelope v1.ErrorResponse - if err := json.Unmarshal(response.Body.Bytes(), &envelope); err != nil || envelope.Error.Code == "" { + var envelope struct { + Error map[string]json.RawMessage `json:"error"` + } + if err := json.Unmarshal(response.Body.Bytes(), &envelope); err != nil || len(envelope.Error) == 0 { t.Fatalf("integrity failure leaked a partial page: %s", response.Body) } } From e2656aab378fa40fb0b5baa61b31cede216ab3c7 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 01:38:20 +0800 Subject: [PATCH 14/51] Derive live CPU utilization from cumulative samples --- apps/web/e2e/agents-lifecycle.spec.ts | 16 ++- .../features/dashboard/DashboardView.test.tsx | 4 +- .../dashboard/RuntimeTrendCharts.test.tsx | 1 + .../features/dashboard/RuntimeTrendCharts.tsx | 9 +- .../features/dashboard/runtime-trends.test.ts | 90 ++++++++++++++- .../src/features/dashboard/runtime-trends.ts | 104 +++++++++++++++++- .../runtime-observability-design.md | 15 ++- contracts/agents-api/runtime-observability.md | 6 +- docs/web/README.md | 5 +- docs/web/README.zh-CN.md | 4 +- 10 files changed, 231 insertions(+), 23 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 8dd744dd4..452323cfb 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2480,7 +2480,11 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand if (new URL(route.request().url()).pathname !== "/v1/agents/sessions") return route.fallback(); await route.fulfill({ status: 200, contentType: "application/json", body: JSON.stringify(list(runtimeSessions)) }); }); + let runtimeObservationReads = 0; await page.route("**/v1/agents/runtime-observations*", async (route) => { + const sampleIndex = runtimeObservationReads; + runtimeObservationReads += 1; + const observedAt = baseline - 1 + sampleIndex * 30; await route.fulfill({ status: 200, contentType: "application/json", @@ -2500,10 +2504,15 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand status: "observed", reason: null, allocation_created_at: baseline - 8_500, - resolved_at: baseline, - observed_at: baseline - 1, + resolved_at: observedAt + 1, + observed_at: observedAt, started_at: baseline - 8_100, - cpu: { usage_seconds_total: 7_350 + index, capacity_cores: 2, usage_cores: 0.72 + index / 100, utilization_ratio: 0.36 + index / 200 }, + cpu: { + usage_seconds_total: 7_350 + index + sampleIndex * 30, + capacity_cores: 2, + usage_cores: null, + utilization_ratio: null, + }, memory: { usage_bytes: 1_288_490_188, limit_bytes: 2_147_483_648 }, })))), }); @@ -2523,6 +2532,7 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard.getByLabel("Memory usage: 3 live samples")).toBeVisible(); await expect(dashboard.getByLabel("Compute uptime: 3 live samples")).toBeVisible(); await expect(dashboard.getByLabel("Token throughput: 3 live samples")).toBeVisible(); + await expect(dashboard.getByText("CPU usage live trend available")).toBeAttached(); await expect(dashboard).not.toContainText("Collecting live samples"); await expect(dashboard.getByRole("table", { name: "Runtime targets" })).not.toBeVisible(); for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 9e3c8c4df..1e6b2f843 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -252,8 +252,8 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("Token throughput"); expect(html).toContain("150%"); expect(html).toContain("Collecting live samples"); - expect(html).toContain("1/2 minimum · no history is synthesized"); - expect(html).toContain("CPU usage collecting live samples; 1 of 2 minimum"); + expect(html).toContain("1/2 valid points · 1 snapshots · no history is synthesized"); + expect(html).toContain("CPU usage collecting live samples; 1 of 2 valid points from 1 snapshots"); expect(html).toContain("Latest value"); expect(html).toContain("Missing samples"); expect(html).toContain("Runtime targets"); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index 3f73ca7b1..5cee52939 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -13,6 +13,7 @@ function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample cpuRatio, uptimeSeconds: 120, }], + cpuCandidates: [], memoryUsageBytes: 512, memoryLimitBytes: 1_024, tokenTotals: [], diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index daf0277ba..7a9313c8d 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -73,6 +73,9 @@ function TrendChart({ const x = (sampledAt: number) => PLOT.left + (sampledAt - start) / range * plotWidth; const y = (value: number) => PLOT.top + (1 - Math.min(yMaximum, Math.max(0, value)) / yMaximum) * plotHeight; const hasLine = series.some((entry) => segments(entry.points).some((segment) => segment.length >= 2)); + const validPoints = Math.max(0, ...series.map((entry) => entry.points.filter((point) => ( + point.value !== null && Number.isFinite(point.value) + )).length)); const ticks = [1, .66, .33, 0]; const xTicks = [start, start + range / 2, end]; @@ -108,10 +111,10 @@ function TrendChart({ : ; }))} - {!hasLine ?
Collecting live samples{samples.length}/2 minimum · no history is synthesized
: null} + {!hasLine ?
Collecting live samples{validPoints}/2 valid points · {samples.length} snapshots · no history is synthesized
: null} - + {series.map((entry) => { @@ -185,7 +188,7 @@ export function RuntimeTrendCharts({ samples }: { samples: readonly RuntimeTrend return (
- `${Math.round(value)}%`} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} /> + `${Math.round(value)}%`} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} /> formatDashboardBytes(Math.round(value))} /> formatDashboardDuration(value)} /> `${formatDashboardTokens(Math.round(value))}/min`} /> diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts index a3000c959..e80670d09 100644 --- a/apps/web/src/features/dashboard/runtime-trends.test.ts +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -14,11 +14,16 @@ function snapshot(at: number, options: { input?: number; output?: number; cpuRatio?: number | null; + cpuUsageCores?: number | null; + cpuUsageSecondsTotal?: number; + cpuCapacity?: number; memory?: number; startedAt?: number; + allocationId?: string; + observedAt?: number; } = {}): RuntimeDashboardSnapshot { const sessionId = "11111111-1111-4111-8111-111111111111"; - const observedAt = Math.floor(at / 1_000); + const observedAt = options.observedAt ?? Math.floor(at / 1_000); const input = options.input ?? 100; const output = options.output ?? 20; const session = { @@ -53,12 +58,29 @@ function snapshot(at: number, options: { } as AgentSession; const observation = { id: sessionId, + object: "agent.runtime_observation", session_id: sessionId, + environment_id: "22222222-2222-4222-8222-222222222222", + mode: "openai_hosted", status: "observed", + reason: null, provider_type: "docker", + instance: { + kind: "managed_allocation", + allocation_id: options.allocationId ?? "33333333-3333-4333-8333-333333333333", + device_id: null, + connection_generation: null, + }, + allocation_created_at: observedAt - 180, + resolved_at: observedAt, observed_at: observedAt, started_at: options.startedAt ?? observedAt - 120, - cpu: { utilization_ratio: options.cpuRatio ?? .25, usage_cores: .5, capacity_cores: 2, usage_seconds_total: 30 }, + cpu: { + utilization_ratio: Object.hasOwn(options, "cpuRatio") ? options.cpuRatio ?? null : .25, + usage_cores: Object.hasOwn(options, "cpuUsageCores") ? options.cpuUsageCores ?? null : .5, + capacity_cores: options.cpuCapacity ?? 2, + usage_seconds_total: options.cpuUsageSecondsTotal ?? 30, + }, memory: { usage_bytes: options.memory ?? 512, limit_bytes: 1_024 }, } as RuntimeObservation; return { sessions: [session], observations: [observation], loadedAt: at }; @@ -104,6 +126,68 @@ describe("Runtime live-window trends", () => { ]); }); + it("derives real CPU utilization from cumulative samples within one Runtime incarnation", () => { + const cumulative = (at: number, usage: number) => snapshot(at, { + cpuRatio: null, + cpuUsageCores: null, + cpuUsageSecondsTotal: usage, + cpuCapacity: 2, + startedAt: 0, + }); + let samples = appendRuntimeTrendSample([], cumulative(60_000, 10)); + samples = appendRuntimeTrendSample(samples, cumulative(120_000, 70)); + samples = appendRuntimeTrendSample(samples, cumulative(180_000, 130)); + expect(samples.map((sample) => sample.targets[0]?.cpuRatio ?? null)).toEqual([null, .5, .5]); + }); + + it("does not derive CPU across restarts, allocation changes, or counter regressions", () => { + const base = snapshot(60_000, { + cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 100, startedAt: 0, + }); + for (const next of [ + snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 1 }), + snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 0, allocationId: "44444444-4444-4444-8444-444444444444" }), + snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 10, startedAt: 0 }), + ]) { + const samples = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), next); + expect(samples.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); + } + }); + + it("does not connect directly reported CPU across Runtime incarnation fences", () => { + const base = snapshot(60_000, { cpuRatio: .25, startedAt: 0 }); + for (const next of [ + snapshot(120_000, { cpuRatio: .5, startedAt: 1 }), + snapshot(120_000, { cpuRatio: .5, startedAt: 0, allocationId: "44444444-4444-4444-8444-444444444444" }), + snapshot(120_000, { cpuRatio: .5, startedAt: 0, observedAt: 60 }), + ]) { + const samples = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), next); + expect(samples.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); + } + }); + + it("rejects non-finite CPU ratios produced by finite provider inputs", () => { + const direct = runtimeTrendSample(snapshot(60_000, { + cpuRatio: null, + cpuUsageCores: Number.MAX_VALUE, + cpuCapacity: Number.MIN_VALUE, + })); + expect(direct.targets[0]?.cpuRatio ?? null).toBeNull(); + + const cumulative = (at: number, usage: number) => snapshot(at, { + cpuRatio: null, + cpuUsageCores: null, + cpuUsageSecondsTotal: usage, + cpuCapacity: Number.MIN_VALUE, + startedAt: 0, + }); + const samples = appendRuntimeTrendSample( + appendRuntimeTrendSample([], cumulative(60_000, 0)), + cumulative(120_000, Number.MAX_VALUE), + ); + expect(samples.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); + }); + it("matches token counters by Session without creating churn spikes", () => { const first = snapshot(60_000, { input: 100, output: 20 }); const second = snapshot(120_000, { input: 220, output: 50 }); @@ -141,6 +225,8 @@ describe("Runtime live-window trends", () => { let samples = appendRuntimeTrendSample([], many); samples = appendRuntimeTrendSample(samples, later); expect(samples.every((sample) => sample.targets.length <= RUNTIME_TREND_MAX_TARGETS)).toBe(true); + expect(samples[0]?.cpuCandidates).toEqual([]); + expect(samples[1]?.cpuCandidates).toHaveLength(20); expect(samples[0]?.tokenTotals).toEqual([]); expect(samples[1]?.tokenTotals).toHaveLength(20); }); diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts index ebc1c6295..82430b613 100644 --- a/apps/web/src/features/dashboard/runtime-trends.ts +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -14,9 +14,18 @@ export interface RuntimeTrendTarget { uptimeSeconds: number | null; } +export interface RuntimeTrendCPUCandidate extends RuntimeTrendTarget { + observedAt: number | null; + incarnationKey: string | null; + usageSecondsTotal: number | null; + capacityCores: number | null; + reportedRatio: number | null; +} + export interface RuntimeTrendSample { sampledAt: number; targets: RuntimeTrendTarget[]; + cpuCandidates: RuntimeTrendCPUCandidate[]; memoryUsageBytes: number | null; memoryLimitBytes: number | null; tokenTotals: RuntimeTrendTokenTotal[]; @@ -50,13 +59,21 @@ function sessionTitle(session: AgentSession): string { return session.agent.name ?? session.agent.model; } -function cpuRatio(observation: RuntimeObservation): number | null { +function reportedCpuRatio(observation: RuntimeObservation): number | null { if (observation.status !== "observed") return null; const reported = finiteNonNegative(observation.cpu?.utilization_ratio); if (reported !== null) return reported; const usage = finiteNonNegative(observation.cpu?.usage_cores); const capacity = finiteNonNegative(observation.cpu?.capacity_cores); - return usage !== null && capacity !== null && capacity > 0 ? usage / capacity : null; + return usage !== null && capacity !== null && capacity > 0 + ? finiteNonNegative(usage / capacity) + : null; +} + +function incarnationKey(observation: RuntimeObservation): string | null { + if (observation.status !== "observed") return null; + const startedAt = safeInteger(observation.started_at); + return startedAt === null ? null : `${observation.instance.kind}:${observation.instance.allocation_id}:${startedAt}`; } function uptimeSeconds(observation: RuntimeObservation): number | null { @@ -86,7 +103,11 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT return [{ sessionId: observation.session_id, label: sessionTitle(session), - cpuRatio: cpuRatio(observation), + cpuRatio: reportedCpuRatio(observation), + observedAt: safeInteger(observation.observed_at), + incarnationKey: incarnationKey(observation), + usageSecondsTotal: finiteNonNegative(observation.cpu?.usage_seconds_total), + capacityCores: finiteNonNegative(observation.cpu?.capacity_cores), memoryUsageBytes: finiteNonNegative(observation.memory?.usage_bytes), memoryLimitBytes: finiteNonNegative(observation.memory?.limit_bytes), uptimeSeconds: uptimeSeconds(observation), @@ -114,6 +135,24 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT return { sampledAt: snapshot.loadedAt, targets, + cpuCandidates: observed.flatMap((target): RuntimeTrendCPUCandidate[] => ( + target.cpuRatio !== null || ( + target.observedAt !== null && target.incarnationKey !== null && + target.usageSecondsTotal !== null && target.capacityCores !== null && target.capacityCores > 0 + ) + ? [{ + sessionId: target.sessionId, + label: target.label, + cpuRatio: target.cpuRatio, + uptimeSeconds: target.uptimeSeconds, + observedAt: target.observedAt, + incarnationKey: target.incarnationKey, + usageSecondsTotal: target.usageSecondsTotal, + capacityCores: target.capacityCores, + reportedRatio: target.cpuRatio, + }] + : [] + )), memoryUsageBytes: pairedMemory.length === 0 ? null : pairedMemory.reduce((total, target) => total + (target.memoryUsageBytes ?? 0), 0), @@ -126,6 +165,58 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT }; } +function cpuRatios(previous: RuntimeTrendSample, next: RuntimeTrendSample): Map { + const previousCandidates = new Map(previous.cpuCandidates.map((candidate) => [candidate.sessionId, candidate])); + const ratios = new Map(); + for (const current of next.cpuCandidates) { + const prior = previousCandidates.get(current.sessionId); + if (current.reportedRatio !== null && !prior) { + ratios.set(current.sessionId, current.reportedRatio); + continue; + } + if ( + !prior || prior.incarnationKey === null || prior.incarnationKey !== current.incarnationKey || + prior.observedAt === null || current.observedAt === null || current.observedAt <= prior.observedAt + ) continue; + if (current.reportedRatio !== null) { + ratios.set(current.sessionId, current.reportedRatio); + continue; + } + if ( + prior.usageSecondsTotal === null || current.usageSecondsTotal === null || + current.usageSecondsTotal < prior.usageSecondsTotal || + current.capacityCores === null || current.capacityCores <= 0 + ) continue; + const usageCores = (current.usageSecondsTotal - prior.usageSecondsTotal) / + (current.observedAt - prior.observedAt); + const ratio = finiteNonNegative(usageCores / current.capacityCores); + if (ratio !== null) ratios.set(current.sessionId, ratio); + } + return ratios; +} + +function applyCPURatios(sample: RuntimeTrendSample, ratios: ReadonlyMap): void { + const uptime = sample.targets.filter((target) => target.uptimeSeconds !== null) + .sort((left, right) => (right.uptimeSeconds ?? 0) - (left.uptimeSeconds ?? 0)) + .slice(0, RUNTIME_TREND_SERIES_LIMIT) + .map((target) => ({ ...target, cpuRatio: null })); + const cpu = sample.cpuCandidates.flatMap((candidate): RuntimeTrendTarget[] => { + const ratio = ratios.get(candidate.sessionId); + return ratio === undefined ? [] : [{ + sessionId: candidate.sessionId, + label: candidate.label, + cpuRatio: ratio, + uptimeSeconds: candidate.uptimeSeconds, + }]; + }).sort((left, right) => (right.cpuRatio ?? 0) - (left.cpuRatio ?? 0)) + .slice(0, RUNTIME_TREND_SERIES_LIMIT); + const selected = new Map( + uptime.map((target) => [target.sessionId, target]), + ); + for (const target of cpu) selected.set(target.sessionId, target); + sample.targets = [...selected.values()]; +} + function tokenRate( previous: RuntimeTrendSample, next: RuntimeTrendSample, @@ -164,9 +255,12 @@ export function appendRuntimeTrendSample( .sort((left, right) => left.sampledAt - right.sampledAt) .slice(-(maximum - 1)); const previous = retained.at(-1); - if (previous) Object.assign(next, tokenRate(previous, next)); + if (previous) { + Object.assign(next, tokenRate(previous, next)); + applyCPURatios(next, cpuRatios(previous, next)); + } return [ - ...retained.map((sample) => ({ ...sample, tokenTotals: [] })), + ...retained.map((sample) => ({ ...sample, cpuCandidates: [], tokenTotals: [] })), next, ].slice(-maximum); } diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 3da1a5bcf..7d750011d 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -168,9 +168,11 @@ The API projection may additionally expose `usage_cores` and `utilization_ratio` only when the service has two ordered samples for the same Runtime incarnation. A future process-local observation cache may keep the previous cumulative value for this calculation. Phase 2 intentionally leaves both derived fields null -because it has only one provider sample per request. Cache loss must make the -derived fields temporarily null; it must never change the cumulative source -measurement or lifecycle state. +because it has only one provider sample per request. The browser-local live window +therefore derives interval utilization from adjacent cumulative samples only when +Session, allocation, compute `started_at`, and provider observation order still +match. Restart, replacement, counter regression, missing capacity, or cache loss +creates a gap; none changes the cumulative source measurement or lifecycle state. ## 8. Duration semantics @@ -297,6 +299,11 @@ seams separate: complete snapshots collected while this Dashboard instance is mounted; it is bounded, ephemeral, and never presented as durable operator history. +The live window retains the latest complete cumulative CPU counters only long +enough to calculate the next interval. Historical chart points contain the bounded +derived series, not all Runtime identities. This makes real Docker and microsandbox +CPU charts work without moving lifecycle authority or durable history into Web. + The initial implementation uses a 30-second Web cadence plus up to five seconds of jitter and a 15-second whole-refresh budget. These are Web configuration, not API guarantees. Web pauses periodic reads when hidden, refreshes when visibility @@ -365,6 +372,8 @@ this does not add an upstream OpenAI operation. ### Phase 3: Web Dashboard +Implemented for the browser-local current-snapshot live window. + - Add Runtime observations as a third independent Dashboard collection. - Publish only complete traversals and retain the previous snapshot on failure. - Join existing Session Usage and Turn status by exact Session ID. diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 91e100f52..73490e9c8 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -85,7 +85,7 @@ duplication, Kubernetes/E2B source, or telemetry-driven lifecycle action. The internal source interface admits Docker and microsandbox without changing Session attribution or the existing sandbox lifecycle interface. -The review proposal for later API and Web phases is split into the +The current API and browser-local Web live window are documented in the [full design](runtime-observability-design.md) and the -[proposed public extension](runtime-observability-api.md). Neither document marks -those later phases as implemented. +[public extension](runtime-observability-api.md). Durable history, additional +providers, self-hosted telemetry, and idle-policy authority remain later phases. diff --git a/docs/web/README.md b/docs/web/README.md index cbb1d7c83..f9cf419cc 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -35,7 +35,10 @@ browser-local live window for CPU, memory, compute-uptime, and token-throughput charts without inventing missing values. The searchable, filterable, sortable, paginated semantic table remains available on demand. This live window starts when the Dashboard opens and is not durable history; cross-browser retention still -requires a separate operator history capability. +requires a separate operator history capability. When a provider reports only +cumulative CPU time, Web derives interval utilization only across adjacent samples +from the same verified Runtime incarnation; restarts and counter regressions create +gaps instead of false spikes. ### Agents diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index 961a5509e..21b727108 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -32,7 +32,9 @@ Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载 token throughput 趋势,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的 语义表格按需展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。live window 从打开 Dashboard 后开始采集,不是跨浏览器持久历史; -持久保留仍需运维方配置独立 history 能力。 +持久保留仍需运维方配置独立 history 能力。若 provider 只报告累计 CPU 时间,Web 仅在 +相邻样本属于同一已验证 Runtime incarnation 时计算区间利用率;重启或计数回退会形成 +数据缺口,不会制造峰值。 ### Agents From fd8a778683b19319a003a0bd31e6162e596894cc Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 02:05:56 +0800 Subject: [PATCH 15/51] Add Runtime live range controls --- apps/web/e2e/agents-lifecycle.spec.ts | 4 + .../src/features/dashboard/DashboardView.css | 80 ++++++++++++++++++- .../features/dashboard/DashboardView.test.tsx | 4 + .../dashboard/RuntimeObservabilityContent.tsx | 36 ++++++++- .../features/dashboard/runtime-trends.test.ts | 14 ++++ .../src/features/dashboard/runtime-trends.ts | 15 ++++ .../runtime-observability-design.md | 8 +- docs/web/README.md | 5 +- docs/web/README.zh-CN.md | 5 +- 9 files changed, 162 insertions(+), 9 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 452323cfb..9383c0b19 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2524,6 +2524,10 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard.getByRole("heading", { name: "Memory usage" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Compute uptime" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Token throughput" })).toBeVisible(); + const liveRange = dashboard.getByRole("group", { name: "Runtime live range" }); + await expect(liveRange.getByRole("button", { name: "1h" })).toHaveAttribute("aria-pressed", "true"); + await liveRange.getByRole("button", { name: "15m" }).click(); + await expect(liveRange.getByRole("button", { name: "15m" })).toHaveAttribute("aria-pressed", "true"); await page.waitForTimeout(20); await refresh.click(); await page.waitForTimeout(20); diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 48426170d..f4563cb86 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -982,13 +982,85 @@ border-radius: 6px; } +.dashboard-runtime-live { + min-width: 0; + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-live-toolbar { + display: flex; + min-height: 54px; + align-items: center; + justify-content: space-between; + padding: 10px 14px; + gap: 16px; + background: var(--surface); + border-bottom: 1px solid var(--line); +} + +.dashboard-runtime-live-toolbar h3, +.dashboard-runtime-live-toolbar p { + margin: 0; +} + +.dashboard-runtime-live-toolbar h3 { + font-size: 12px; + font-weight: 650; + line-height: 17px; +} + +.dashboard-runtime-live-toolbar p { + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + line-height: 14px; +} + +.dashboard-runtime-range { + display: inline-flex; + flex: 0 0 auto; + padding: 2px; + gap: 2px; + background: var(--surface-subtle); + border: 1px solid var(--line); + border-radius: 7px; +} + +.dashboard-runtime-range button { + min-width: 38px; + padding: 4px 8px; + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + font-variant-numeric: tabular-nums; + background: transparent; + border: 0; + border-radius: 5px; + cursor: pointer; +} + +.dashboard-runtime-range button:hover, +.dashboard-runtime-range button:focus-visible { + color: var(--fg); +} + +.dashboard-runtime-range button:focus-visible { + outline: 2px solid color-mix(in srgb, var(--accent) 50%, transparent); + outline-offset: 1px; +} + +.dashboard-runtime-range button[aria-pressed="true"] { + color: var(--fg); + background: var(--surface); + box-shadow: var(--shadow-control); +} + .dashboard-runtime-trend-grid { display: grid; padding: 14px; grid-template-columns: repeat(2, minmax(0, 1fr)); gap: 12px; background: color-mix(in srgb, var(--surface-subtle) 68%, var(--surface)); - border-bottom: 1px solid var(--line); } .dashboard-runtime-trend-card { @@ -1881,6 +1953,12 @@ gap: 5px; } + .dashboard-runtime-live-toolbar { + align-items: flex-start; + flex-direction: column; + gap: 7px; + } + .dashboard-runtime-trend-legend { max-width: 100%; justify-content: flex-start; diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 1e6b2f843..51cd7dd43 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -246,6 +246,10 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("1m 13s / 2 cores"); expect(html).toContain("512 MiB / 2.00 GiB"); expect(html).toContain('aria-label="Runtime live-window charts"'); + expect(html).toContain("Live resource trends"); + expect(html).toContain("Browser-local samples · no durable history"); + expect(html).toContain('aria-label="Runtime live range"'); + expect(html).toContain('aria-pressed="true">1h'); expect(html).toContain("CPU usage"); expect(html).toContain("Memory usage"); expect(html).toContain("Compute uptime"); diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index 9e59b7747..19f06b320 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -36,7 +36,14 @@ import { } from "./dashboard-model"; import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; -import { appendRuntimeTrendSample, type RuntimeTrendSample } from "./runtime-trends"; +import { + appendRuntimeTrendSample, + RUNTIME_TREND_RANGES, + RUNTIME_TREND_WINDOW_MS, + runtimeTrendRange, + type RuntimeTrendRange, + type RuntimeTrendSample, +} from "./runtime-trends"; const PAGE_SIZE = 10; @@ -287,6 +294,11 @@ export function RuntimeObservabilityContent({ const model = useMemo(() => buildRuntimeDashboardModel(snapshot.sessions, snapshot.observations), [snapshot]); const summary = model.summary; const [trendSamples, setTrendSamples] = useState(() => appendRuntimeTrendSample([], snapshot)); + const [selectedTrendRange, setSelectedTrendRange] = useState(RUNTIME_TREND_WINDOW_MS); + const visibleTrendSamples = useMemo( + () => runtimeTrendRange(trendSamples, selectedTrendRange), + [selectedTrendRange, trendSamples], + ); useEffect(() => { setTrendSamples((current) => appendRuntimeTrendSample(current, snapshot)); @@ -301,7 +313,27 @@ export function RuntimeObservabilityContent({ } label="Reported tokens" value={formatDashboardTokens(summary.totalTokens)} detail={`${summary.tokenCoverageCount}/${summary.sessionCount} Sessions report usage`} />
- +
+
+
+

Live resource trends

+

Browser-local samples · no durable history

+
+
+ {RUNTIME_TREND_RANGES.map((range) => ( + + ))} +
+
+ +
diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts index e80670d09..fa33ec427 100644 --- a/apps/web/src/features/dashboard/runtime-trends.test.ts +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -6,6 +6,7 @@ import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; import { appendRuntimeTrendSample, RUNTIME_TREND_MAX_TARGETS, + runtimeTrendRange, runtimeTrendSample, tokenThroughput, } from "./runtime-trends"; @@ -115,6 +116,19 @@ describe("Runtime live-window trends", () => { expect(samples[0]?.memoryUsageBytes).toBe(768); }); + it("selects a live range relative to the newest complete snapshot", () => { + const samples = [ + runtimeTrendSample(snapshot(0)), + runtimeTrendSample(snapshot(10 * 60_000)), + runtimeTrendSample(snapshot(30 * 60_000)), + runtimeTrendSample(snapshot(60 * 60_000)), + ]; + expect(runtimeTrendRange(samples, 15 * 60_000).map((sample) => sample.sampledAt)) + .toEqual([60 * 60_000]); + expect(runtimeTrendRange(samples, 60 * 60_000).map((sample) => sample.sampledAt)) + .toEqual([0, 10 * 60_000, 30 * 60_000, 60 * 60_000]); + }); + it("derives throughput only between monotonic cumulative samples", () => { let samples = appendRuntimeTrendSample([], snapshot(60_000, { input: 100, output: 20 })); samples = appendRuntimeTrendSample(samples, snapshot(120_000, { input: 220, output: 50 })); diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts index 82430b613..76667bf4f 100644 --- a/apps/web/src/features/dashboard/runtime-trends.ts +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -6,6 +6,11 @@ export const RUNTIME_TREND_WINDOW_MS = 60 * 60 * 1_000; export const RUNTIME_TREND_MAX_SAMPLES = 120; export const RUNTIME_TREND_SERIES_LIMIT = 3; export const RUNTIME_TREND_MAX_TARGETS = RUNTIME_TREND_SERIES_LIMIT * 2; +export const RUNTIME_TREND_RANGES = [ + { label: "15m", milliseconds: 15 * 60 * 1_000 }, + { label: "1h", milliseconds: RUNTIME_TREND_WINDOW_MS }, +] as const; +export type RuntimeTrendRange = typeof RUNTIME_TREND_RANGES[number]["milliseconds"]; export interface RuntimeTrendTarget { sessionId: string; @@ -272,3 +277,13 @@ export function tokenThroughput(samples: readonly RuntimeTrendSample[]): TokenTh outputPerMinute: sample.outputTokensPerMinute, })); } + +export function runtimeTrendRange( + samples: readonly RuntimeTrendSample[], + range: RuntimeTrendRange, +): RuntimeTrendSample[] { + const newest = samples.at(-1)?.sampledAt; + if (newest === undefined) return []; + const start = newest - range; + return samples.filter((sample) => sample.sampledAt >= start && sample.sampledAt <= newest); +} diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 7d750011d..be3153ada 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -310,8 +310,11 @@ API guarantees. Web pauses periodic reads when hidden, refreshes when visibility returns, and adds jitter so multiple browsers do not synchronize. Filtering is local to the last complete snapshot and never changes tenant authorization or provider selection. The same complete snapshots feed the one-hour, 120-sample -browser-local live window; reload, navigation, or connection replacement may reset -it, and no point is interpolated or persisted by Core. +browser-local live window. Operators can select a 15-minute or one-hour view +without discarding the retained buffer. The range is measured back from the newest +complete snapshot rather than browser wall-clock time. Reload, navigation, or +connection replacement may reset the window, and no point is interpolated or +persisted by Core. ## 12. Token usage boundary @@ -379,6 +382,7 @@ Implemented for the browser-local current-snapshot live window. - Join existing Session Usage and Turn status by exact Session ID. - Add responsive, keyboard-accessible current-resource views. - Build an explicitly ephemeral live window from complete Web snapshots. +- Expose 15-minute and one-hour views with an explicit browser-local source label. ### Phase 4: optional history diff --git a/docs/web/README.md b/docs/web/README.md index f9cf419cc..53757cea1 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -34,8 +34,9 @@ loads complete tenant-scoped Runtime observation snapshots, and retains a bounde browser-local live window for CPU, memory, compute-uptime, and token-throughput charts without inventing missing values. The searchable, filterable, sortable, paginated semantic table remains available on demand. This live window starts when -the Dashboard opens and is not durable history; cross-browser retention still -requires a separate operator history capability. When a provider reports only +the Dashboard opens, offers 15-minute and one-hour views, and is not durable +history; cross-browser retention still requires a separate operator history +capability. When a provider reports only cumulative CPU time, Web derives interval utilization only across adjacent samples from the same verified Runtime incarnation; restarts and counter regressions create gaps instead of false spikes. diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index 21b727108..ee30d39bb 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -31,8 +31,9 @@ Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载 观测快照,并保留有界的浏览器本地 live window,展示 CPU、内存、compute uptime 和 token throughput 趋势,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的 语义表格按需展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 -Session 的流程。live window 从打开 Dashboard 后开始采集,不是跨浏览器持久历史; -持久保留仍需运维方配置独立 history 能力。若 provider 只报告累计 CPU 时间,Web 仅在 +Session 的流程。live window 从打开 Dashboard 后开始采集,可选择 15 分钟或 1 小时 +视图,但不是跨浏览器持久历史;持久保留仍需运维方配置独立 history 能力。若 provider +只报告累计 CPU 时间,Web 仅在 相邻样本属于同一已验证 Runtime incarnation 时计算区间利用率;重启或计数回退会形成 数据缺口,不会制造峰值。 From 07c85797aba1cfbfc7f33c97c7a54d5fb2185183 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 02:30:53 +0800 Subject: [PATCH 16/51] Document Runtime monitoring dashboard --- docs/web/README.md | 2 ++ docs/web/README.zh-CN.md | 2 ++ docs/web/images/runtime-dashboard.png | Bin 0 -> 78256 bytes 3 files changed, 4 insertions(+) create mode 100644 docs/web/images/runtime-dashboard.png diff --git a/docs/web/README.md b/docs/web/README.md index 53757cea1..b040e0650 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -41,6 +41,8 @@ cumulative CPU time, Web derives interval utilization only across adjacent sampl from the same verified Runtime incarnation; restarts and counter regressions create gaps instead of false spikes. +![Runtime monitoring live trends](images/runtime-dashboard.png) + ### Agents Agents are reusable working profiles. Create one from scratch or begin with a starter diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index ee30d39bb..d4c4b161f 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -37,6 +37,8 @@ Session 的流程。live window 从打开 Dashboard 后开始采集,可选择 相邻样本属于同一已验证 Runtime incarnation 时计算区间利用率;重启或计数回退会形成 数据缺口,不会制造峰值。 +![Runtime 监控实时趋势](images/runtime-dashboard.png) + ### Agents Agent 是可以反复使用的工作配置。可以从空白 Agent 开始,也可以使用事故响应、 diff --git a/docs/web/images/runtime-dashboard.png b/docs/web/images/runtime-dashboard.png new file mode 100644 index 0000000000000000000000000000000000000000..544242f8943ae009a5ad4b2ec752dc41e01f5adb GIT binary patch literal 78256 zcmb@uRa{kD8#k(SN=bKvbVxTyhbYq0DGf@8$O0DK-H4!qbW3+ETDqh`nuTdJ2V`_Oh?OJR}&8jPg?qg zR3a0?xahdgVe|Q+)s|DsP5h$x&!Tcv&jZZiq#<=CzJxTdW$Pt|3?YSRisT?Rol*vk zY%#q`i!cIa^+c99MhLD98Z88uPN4q%&UkkO@ko}a$JyR25<0$6tFpZlZU3nGx7S52 zzLzK9ml?N}4rG;w&x0?LV~>0G=9!sE8n2~@*9im0P_p$sGwh|MVkLu8e5R0#jJN84 zpIk*CKnw!kKI_D^diIT}RdMBKg&6&Rt_yxz#i#EdFVb+T9kW6olC@-gU6}Zc;NO3Q z+M(dH+m09X$1}<%z9qE!_Uem>-&OVhZhDnWrk?%7pI^|*F38|;mZ=L36MQ83-*u6J zLg`GI{`YV49slQxA|Yx2eL{`|o$ucjfk3(Ot^Ru~JW~ePe_uw|0pWnfGpVKTN#Zj_ zTE0Es%y5dP5atDUp?wpA?|*-Lo;Vf$xI3JXT-4pZSo;V2zu&HFNX?8*wI_-^MdL8kuy&Ovk=t&nWM}*f zAGN6aUMjspwBVBG+CUOpp7pEmTHmT|C$zu54E$s8-x_D6mItS;9Jb!8$I?h=i+P=F z3{F>BkKA6La~al;V&v%8yvx%_<}k2==!HI^fD-CXk zfKmDV@oM)dA1RMnSL26W?Gjy6^nahj(Gfq9MRT<+5HXLM+jg9R&Z@?4T7cTvXTK(# zkZm|au*|sCAM6Nqu*+}Hmz`R!z|ZFgixE?QCW=NDJ(k5yJCLH4lb%h2Er`a+VbpYy zpdr!lcgg_IGllc5i2I)IU@`}|w{#AJ+N;02^fTUn|2+3NSTO4jr=1+m6qaR*5j#Qq z((HT5OBXwlRz-tZ{H9pG7MzI7yX(V%>wQ6MTh@`X*1JZxt>K9x4Pp-cL;^qfkf2;+ zx}Z~Z^jwv-#&qdhRvgyd$zt%EpQ>rR;FvsVKJCB+YvkPi)Er!;ZrNW{llKX=u)8Q9|)sIb29sPAuz>2icHHsC9v)JrbCls0(~bONwT9{0=fS`dys8<1*? zs>1DFWwSmO%JeQVKVH*qz!_NUf0cyE?U}Z-h}sg8i9G^mJ`4Mv^Jn^K@JD&xb*d zI^V0EqC!f=9Ckfk4Am$hJg<24o&7TXYMa2Ob#qXg=BsTx<(HbgGz-%BteM!J+JYmq zC;Q)F6T}*OYSrZ|;&GtbR}5-SGM5n11}&tI%6psE0-cvhr_>NF5=Y{~WGMZ4@?l}i zR!+c@4;t-fR4UJpBzU2@!Lcm+^L07AmKerKtgk~_vB?G2+1&F-jl6dl&q!9*Kk?&S z(tJ{o2)MgsZs$p8kqSh(VMIM6sHJ1+-HDB`X(-XFO7GMF^>5kloR%HkE$XcsRu>gda+s@T{cxO2@FJ9ouU%?2_QqkV zgcRAJ<{jaFTW!P2>wjyhOAwx#X^6Kp^Md=lL$ZcC?F0BJR-bFA@=6A^8aSA6&DtERZmnMDR#l3>nSZw!Hp8KR}-B?Dc+^bZ5; z%MS0G7c3{|MLGVfT;E}RFX*FGl?e!B9ghW#TYRMl44XV%20c29E`_{KECz``RS`sj zA~ssVB7euJTi$kpynA6#polkk-a(kq_Th8kzvXeYhbEotetWW*c>HU?{heY7SWr&_ z&*g%Evpu~kyeo3vmo7_LXBjv{4hfr={8gdYZfqtjmqH%+s8zlgqrPGRcVYpr3lvQc zs`;t2&_Bh3U2#XoXnU(;=hkVRn3h^rdf`l!E%MtayyXp-iewN4yK zhS~&CkAq6fLH+whmW$(F>hs;{!dNYI98wvrI6krg9J1?K|K)NmK;y8b0NE5&pCt z2c1z(waqq0zU;}hw)|F`uGNbtAJX1Gp%nSO zX9TwOREb`|?cbEFc>a28a6%`u#eEDF7!_lG2NN4frLN5X?ymUd*iG*Ihb6v>PPz@7 zO3ag+u)35MJZaPI*}SbTn|?KT5K!K$Z^k|=wBB9WHC=%cjLRymnJMgsQkTgREQpP| z1p0!Zg1h5vB{KC)2H)bx;aFx3s&QWc0LTUGX|@!oyck{=s_v0Kf0u|3tCupI1V_tb`yozlX*TdBs?8Wy!3J%ca&8u(ksrw`&%tC&*VNKt=x%J z?K&4Cq6FP?>sCt}Qi{B1@xFAX<$Bv_QOr~|fUhm;`OUDwwKtklr_4y)r0vlOGfODO z7es6_{yzMNCIm%-NL-!nRDteKQLba`mn06}K~ET{+<3!mOZ#?Z%~o6@I^SMxl>R)p zI_kt}R#@qnEHe%`S??$1wRkv>DaQV`{`D*pN86ucLO5ZdmU-+4lVXz-n7WUMpRS01)pm!5wWJ#4Dd{4CGX2Jj9Y0!I$F>}6^nHgCIF`_qkv zQJqj6>ZKMx(RXRiW5mY3E8+tgJ@dlNS|z%1;m(}3wn&iKPso@rLxwU0dsP&i^AT|I zY-%0mYgbCQ29hGE)MfSZ;=m5EX}O_C^Dlep{8hgCO`7n+%_lclSU=SD4Q)s~ObW^l zu$rnnJFW9dLnF&Ik9p6({TvTL;}*c65wLv#Oc=_}l9-XoSn(wwub95ovlJM4?-|n( z%omfex7?azgVj!2#@!j?1X_p-`uOjsL#aH76~Df}VS_t5kK<5^BxW9m9Phw4hg5j0 zF?^bAM?*ms=66B^U7R-hc{_G|oG?!&s2IpfFKg}f41LFOZ_!S*FlQx>+7KafUrz*= zea(>&eZGFll;`&Xj5CB1$ORv8k$aMYS}ms{LvWxEJ~GC-pVC z4N-cB7zg2M1I%E-6l`^*bC`f@#5T3KHGm8L6P}{hEUD*bn#2RPMx=_9fNW?4RQXbl zU?im8wN*o3$C^gMm;JA!E&*U91VGgPs2!CQ=N|?j2a&R;kn~?T)Y%wJ>AkP@Bjm>_ z{eJoO3yp&1NVfQ!?^>JZ;fA@+FC{KyNcpTNc+BJ)TsLNvLY$ia5(N^wyUGuqdy?}W z2H{Aw#VGu~R12Dj8E>=gxfd8AUPrfv8dN!Co8iA_tKrI zf<1aIg02ClV+vYRVuP|D(X6v(1_6o&FSLv5ljOxgW<+@;2T80HG}*K>1dkr9#ZyxX zI6D>@ewI>qvqFV#3=M<9mP%{$=o6tRC=wgXQxYZfC4ta`jgA`oS^T<*J7HrTG6<fm%90ZFNV*8Jo!6wo%HMO%$qux5vk|%tQo{k!q%) z-GrqDb)k9$VxUXY1SO1S+`3!vS{;BFl6!uYD)T4xvn{f~1( zS)tUKhkz*0lo`;7$^5oF`rR%a&HLSaxy4V~$TSL{n=6M;IYd$Gr~5vxSe4aKhb6PR zS$tk1`?8;<=bt71O$ouR;_1kqYXJ8nnbmS*EY9>08m2ejGu1vY{Od)W%~FB8bbC!3 z$mT)IwIlx)ctL<}v<3E&T%RdklQsd2JfHSimQ2twvb=x?u&{UPTO#4%scy&t%pFV3 zK9cu0@ZWWMcvrG21lxSn=tFtxsMut3k0usUx71|YK4uC(f056?G7svVkhuqR$$qvn z4;@`BU0ddti|wdu7d|*#DhoJ2w#5ew#I6NnZJR-}c2;Y~p#HHTlCtshmXKDVm+x)I z$y+X+*Pat{=ngWZdolh2>8VwzqZHVd!7+}OPyZoCJP3X=V&LsEG7X#4c*~D=x*y+; z*E$$vE|xF#rUKmPd)0b>M_wfMvac^=voDT55y?S{X0YpBEUIq=#$0QFleurO{`88$M3KZk$(s1ZL z_#f)5yH0EZlq;)A@V1~zZzUJ{udsm@^FTU8K$7{U= zc;8gLyfi_fT*igw@<@FH%(J4&{{j z`Tp79c_SJB_c}3tOz#*@3!`mb*L6iW#1z_fw|f%&z)ZT-)B#||vl{bH^GU?5EdTok zYCB4@iR=yNaGuH8&pW7nEUbku*1~hY9nsH1=BVq#Pcjz)>-A6;x;Vz^s&{=l(bzgK zy<+GXND_2-Q39RKj(W%0jn!CHb8}W|CPpRP*O<@oEQgm=zj`K?Z8&GKspoC5#xLXM z5LC0gz*vef1n1zze;U@pMYZE1pveoMoX!E9cC^yTLAy%upMbS$73e$)=ygZ9Ihu=u zRO*YZ4eCwHbgz?jo^!jNfN)|i^}sA2=$9q85piwES8z7!M-aTf8Mn3&YMQ-zG+hCU z-|UsGXW2yMG*bld+V0m^JWF16>sD1(I=P?J-sF-4>Xw3)YS14HfVeMkQ;2(OeOK)= z=+M^ogm5(-A>0`PR-f&;ir(6aOHANQ^dNPyX)7IeHZTJ|lEkJ{ONS&;4Y$6|&kFnR zlHhhrbu02WeqKrYkMO+jV|W^^R+F!Nd3n5wxIf4pC687hf;al++O_!%G%C+-64P`ak7U^@aSf> zbRcmhv*{eHc1M7wqTAqV*%^ZEw%!*9_%(cMB!*hje}Aq9XdsNZ(?H&uHV#;naEhT4 ze?lz*1S8Y@kq03M98{s$SW3r@MAp|rf8rFVlrjW3Y+EijGt4AHo=^@1E;fF+-Ysi= zRiy4<^dkZ6boc$a<1zmjjIDQ|)B}VMu5-6vw;YeSCE{~tUuxKZ@`S>s6}1Yy=P0u0 z%84vv>Cf$09_67ALs>Sy0);ERl>!K+-wYfZf+_l%0JnJhcaW=9%zck!a=F4>;qK-V zpbI&(t=%BQA8|m(RiO4uoXUL!w2VcddwnvF27Qp2^%bNh2P{U8kV}W&2|%c?euE1- z%#@pg1pol}Ef6iHznJ$%gL4aLns)>d2g_ukTEGfgjGLSRQmwdR;-47O&;%(vCih1QGSg z4A+*+n-#!D@As-k>T|*2dtlCI`|~+oo20|!>&+tkcc3KgJUUIoqn97W=C$biy_YfU@?u2*FR^t%jkAA#;dPec?f=QC&hJy4B`))OH{@3Lz?y!8`xnnXRKw~l^?c;_;VK-^;RB#Aymwp^`K;S=u^IsoE4Ic{1>BgN;F}nUD;t1* z0QK$Y?3|ZvYyx{|7#`2jzXXItIiOQW*j&nBUnJ%L+3JC_Py$ZjbH3=IDg97B18J&N zH!o_?UcfUDDPAPvXeuaQqtpU8`MYQpb&YP#GTlux@_ad9`ErTmW*G2#h6nCzbFV{V&RdZco35WWEM%jLTS}nyrO5>G@=D#U|UrEO_yAUfu?u<-+hO?&j>u z#=ZmL!cI$P?GL$df3oWX#~6mJTvRw~<2Aij)Kp=*Rf%hckaG?jk6`ovQuIC_n*%5X};1D{0Xc&^P*>8EUm|a3+i;YBct`XRe?}%ZXJ=kpt+u!%GxP=))9ba|s`MMVS^D%* z>*!Z@3TXtFr(xWfae+Xt28K?WBHof@^z$boE{_n{hja3NKfNenprNipt&2wlbu>KT zon*c&3_>Ri;B5@j8)y$n&R?(>Xcx4ay+35yUNb3C*i_1_J&d+svxf6_L>mIEWWgN{iA+HV`_4>iaPzvM{?|+9&DD)X&m9_%s?f_abYHl zjXzxcnSvtsJWstyEhFTPXD%eMtEAotN-*dlC}8Zeh+g^OWW`r2OdmJ#^Hlwg#c;!( zc=aQ?gyPM*hEHGxp==mEmib1typM+!v$aACHzT!hY2V%FL)$0YN1ttOF#)yw9cM)y z#;=q{0uF{^8Ht%g7(w7-ff9jEg)lNB5GxxR21*@wRCo46{VRNg;#mGb+pOjtr4A=i zTXjZf%ez$2e`-G4;&N?ltaHP<`VlNHg$7^p+RA#bFaKrP82+z}NQq*~qwWCa7(f$Ea{LCjVAH<22()8ve^vsnFR@QsL3BWuOOS zAml0??Vw|_>sR;b35b5-;2Pr6kL3u?Q?x2NXGTWFp=z?|$Ctqc!iFIz$>@~)_A{}6 zgmS2BVtA>3yq=zA{?NkNJx|7xmM0tfssI!y!FRv<^d8;VT4WS!ZK4Y9>CZFs3eWG$ zhuCl%S{E#WQ-VVw^zw)P^UwT8walui9|iBb2sCnKkM3BTSn*F>cc*?NYedak>TV%o zo2LZYnd)R>4|h}j^io5<70_|c)EPAt`mUr6#n^g|{9XRT+2OLrq~eHPW_OhsWn^mn zTa%xn4c|Xf$^Hh&mE<^N-8@KN6fVa}_x#7Blpuo3{bZEZ+tDkhJ)zfg%MzFC322_m zyElPP{$2K-Ql?9x(2QcAYmV=A8G0?s@{>dAX$)G-7?QG#OD6FJ?n1fnGN1Y) z(SEHd{d(wV5v+cWxR1N`X%?$+EI-qi*#@6;$Ki`!>Hr`ik;E)uNC7nwP23M?u9@OM zfANBE-oi&DADkflPK?K|#lHW(u|#mPqlvOgYX&27fn$AK;K(Qp>KrBrR}Qhj?oCMmM-s@->Xt*2spTA zN0eiwP_2Zv$)eM6xlGxro}K=g#8i-6*sQ!e$4JLin9zKtyzNcQLYOC_75-gQ`a7wZ zb$Z+=F8St?dg~jhxi0{>13L++M2O3KgfLh(HrL`)X17yab?eEQEc_!KIe!)#n+Up& z8>m2#tbL!5ACGh-F9@+9#Xj@q+=?7n{7*t31T2(_f13^#+TjK@`w9-VM3Y*EfY7N> zO&7M+PXm1jQaL`9oOKI(0OT&iePS{^Q*U9_vOAG0nMl5Diqs( z?QMW!fZeB4tj*yo2S$fY>|tomK2 zOueAv*IdY}yz|EQI~_Xfwts5;Z@sms)x{7(h`5!5nBL9Ta-_%d;)<9DbpZ&qCX$Yu z4{wofeTf-9R-{4gDkVqjI-`Wvy|==iNe7l*Oea=FRe1N;H`(k`V~Amul^USu-+lqc z@QrH{E1Pd}Z>d?-FKi$kcB}o^!3sfpJ%`VF*yd&`r_Nx`^LhmUjRzY{`9;#(JmwI> zx1~VhCib1VzBTeYB*fAZ+tM>!%qZLAAt@k2(Z|{-ye9bvf0`-s2bj!V!+kf;8eVE| zqpST!LkZBvPFw%wIRXesDv$YSMpTRU|Fi&m3PQzeJXN$~bzvvvGo=1mDVk5OX1kw! zA=+eV+6T&MnoZ;~TApBz6zeTMLxtBNv|)t7y$(rA%!HB`fi$6WMjI5j&39sNPONn_DyYX3hzdBC38CA6yVpH~pnArnF6wTBr z2VT&ou@M|*vm8%#)G&8#VyN4H<89UJRz`0|+Rhgeku9) z{pMX&pyW@YcxhLkhUoI=@R+0DcLclr^!O2x#DfxwbJ%muFpoO{$N{+hy}SY#U=M&c zl{hOVJ$u7R4ftej{tJHfXOI*q4!^t}(6&kOwCw@c4F&-KWDS5c&$)T@OyjvtJ7j_} zfJO}n6N8lo5G96K9*n5j&Vu%~Y8=x5Q1jkiF1QYI0E}L>nz}uy*@axPYTWpq)jgEt z3CpZ}TKp5G`}YS*;86h7V5fCtrxO~mm;If!zm^URV!p}HY>B7WERHDo^srZjsyRHb z9Z^CEi^~&bYAvXgS3!&dFvI7H%#NU18l``zsb=l`8%RR+P=RMFH-K|ESEP|M%3-Mr z9wwsi3BNXDw$DKWfD}4E^lN~d1D?0i42yn{HOhYG5;!yS^H|hu6tS5qE zdK}#gVCQ7{WF`-*(152F$cIPjU115wbe)kPCcnB+*vxCn$pudaoYsPg`?Gl`UsxUc!*UNuJ| z>jjaF5zr%iZvRf}IhbDOg;JS)y#qFpTw69^c7Qn%({GV+F|l<#NnLti`?tQZ6fR?s z&?5P0p8#+2g3shKSoM_W(*e#J;)&AN4U#dp7i%#9#PEKwKrp@B1lH&USov=qKr&^eb!kPxCc~l-8pU_f z=u!1rb%e%wv5dHB7i&BG{Zrh(Eez;T=&bEIX(Pqr8jzZ#QOKGT^Q5IdS}gm*hrk1G z>jNYCi-9XUVgqn>!A+L}dR)2=9^#$4%e+IK>EndUR8f!8oikneM!~QOv&`VU^I7Yh zI9kj^k}e=eOTuA%kuBhb>B&IkyA;wK(7ey~!ak((Sw|0^!`0W4t=YMaW`an~l6!7u zUke6xJlE25wgqV(fV~AB7q;0^g63>@nmVT#)}Dga1`;3x;8-;!NP(zUD$;O{q@@3i z=O(t8ysuVa^mDr@6GL_vS==nB9w?LMY*hQ7##A;nnL2-%Hj`#(0ZF*32yDB0m-?_% zldA)euJZrg_Bdn#WD1VvVXORWy?xoQz9>5^H3zC)?Kb~dv^FWgJSH* z<4#>m;RzuxU9JQskbsM%7~J|{1gs^nYdaCmzjYoeDfJ)GV+S>WtP{20DPH)XTpcUl zNepj7At(g^L_K@k_#mvraQW>1sMB99WUfTSo;RN-%xnT?Q6Aei{4`*2ORmxZdH-RX zN!iinNC3?qD7?|npPrd~09zUk^icg4Q%hJR=?|tm znctm*AmOWXYnfX4YzHM&F@gO^oM()83 zU2~7`5&qy4`VR1BAtJ*NnJ3~B2XRRqUaXGnQ0Xg)WjJ!V6f8EvrFJ zmv+=OfIiPVwPs;IX)nMNG>4dAR*39bPu+OrUHFY#1n}__3mtr){i%Z}%ZV>|- zm%j!^$-7JmzjRAiF|h?v?a?7;%a1(9AQ3_+7_IlQm7*`!iUohRs!L770+p;eju4d% z3E}UNR4hkrRJ)?@N7|3VJhCAb=0&E21S&zIP%#4#2Hha#a#uctR<#FMh9!3U8}`5a z^#qfy3o*;bkRYONy+B$PpKzhv6p<~xPeAoAP~uWQ!>-~sGE~uSX}Rm9yGR&UYhbiZEz+(3DRWaP_t!1R}!R5hEr z*l}GSdD#eDV!aHS3OjV_-5zg7(VdxXS$n$0RN9YTi$GVf{QdO3q8V~#V=T^@feb{p zOM&zcvPU!v(GzmYnX<#J^qyw(Kz>mBE#cWALP1}M4lC?R-bk*z^lYqHL>W?46*i}XvNTH(qxlp&?>X8tEdGSL)tOU>2{`^;F z-_3K}mjVID;|NK~6KU9dOhd-JE4WYA=?h(eWx_x%w98q)~g3Er&8H$v}U) zPRZpy0AhG2Nw+3~Yw7*%_5>uo6E9x-fdt)*%#%7}Y_rQ%bE}&Vh^)p1J|Ed z3^W1u>p)42oMecajnS)_m?7s(&4TWkhOri@P=u3{uZs|2Q{$~brY+saUz$6bbM2r%s)_1M284*oaKK5WV^`t+|Y+47y@>BLNNVRV)GsO4m%&;&o zLFMO{&I6^p#$leQO0MX6ab&%X<|eL9A9o>lEYQg*9djx;u_)X^3kYNufFNt8nUmn% z8<7<``;l;nh(kX$3lMEf>HFhE>qV_*l00OO%Cl2%1uj-ieS%KRWaai6oQfWJ}eQsI&Lx;$+3gEADMG zjrvIXvAyF2^|Jh>8C593YU3J!93{X|bE`349(f++)BA}{(C3dk_8ar>Ft@(VGJ^SW zWU|_IIB0F|+oO-&hz1xy5*QI{Bb3^|k7PPR){(^Q6mfzR8ngj>;|<$(w;qu{K8s%O zpAd8s&O8h0MTSdaG>g|1_QCiF8D%|q(;!sB9on3EbhF z3+5>Wej_r8gs;$t_F2E;;A{{kVlB8gAm^IxtrSBZ!*lY2rsO&U;C8bfdB^J z!Sl5-e~Vu|wmS&e^FV5D04xKP;3pF3BrHo48j?e38-_-ISxR5$!x1Mu62Dg2uF)Od zHC;ZdvX?Z|U4x(}tT%kQX=U@rNR&;l7CvGMpk8*Hp9(NECS-p<-O9#fFh&P-38rB? z#owKmzr%^e)W;zg>}qt>rSmz45Z71C^+CfXlyAxpO^5{X$1dIuOCe&9zEw*2BH?j% zu`x&-L(xKZ)?PsBlT2%V2cSQ*O7_`JL&93XJ`oEX7nKaLt4Uy>VPeuY2-l*|ldt_o zZWbPtnf}3VxLY}jkPw7@7}D&1@eyc$dll?iBM4rbOY-9Yvd01-9gE>)-G|jiTk2Zb zyAZxJuu%G3MIC_ib1961I}WGS@B~jxjfMHc>dDe498FqYEB2RPyq(M}Ny^dJymfUB5k{z#a%da_&qy2hW z*xhhhj%9Jg6T#dB?wd5iLba!KAtnVEq6b^UUwOCgDmSQq7E^rs!EdbR^06xCI;-c$ z*Pc%kCSD50vF+qHZKH4=$ZAUXZiEn?)qN>Nr=$uR-aLpfAHdS+bvTXbw5={3jbS>u z-}&07hWMD(eAAtjWR=Ga=#V}&`94PjDcW+gpjioaJjr@Sr{G`Ts-j7zF1n6cjAU#M z?)tKd_RbvKN<90ac=YVFtBmqVe#c6BhdNyY7t>#QWyZVW1ggS**cw80X$693$LMvO zl*)83`y+_+J?*T`mdRR`xFBd*YrP(7*{29CzvDU$phWq zz5S*7ntw7kfgomuFWcJHhlC4jpq8P)!zl-?CqFT)%_n~26}01l9kRJYaHWO(_x=nD zPnX?hO?F!Jm)?oI?0|$_SPQ`3cG=^$Sg93Qe4Aow#-)4j>r&2GnkDh(jP|?zTw=;l zb}+nEdsKBn3*ZMfSG#6hdawm9H1XA;&eE1qw-tWPD2G03=>#omSt9cVVH#{^Omg*L zZ~S$jQnJ=#MEQJ7SyOo8B<^=*az?@Ugi<$PWp$Q9Sh-LsgH365J*{e? z(R$=W>szHoU+hTAT|FZNN@_h!p;za)TFFRXFwixm3a)rc>bP>a5P%YL!c_bz_+7kF zhbpvOF+)HncwiCa=W{@n!z3QIUlfCb4JJxZ9(wDmk}mT1YUFxrVy*W>P7W;Ro5N?>f|!Hn+#YQ_2CvQrqzM1fL**0)zBUt}0{I|N!a?jkPPhwSkQ%<_t6lHBs77AIX>hXGNy&VF zJ8hs4WA8A2S`6N2vV?5QyXRNFqV>aBU*;hyWT++^SaqYeX2o%W4-G%kjXd zLXv>FLnVc%M$mVEP(={=I=pyVst82GcxH_jpf>SYFL*#UU}*&NfXxZ_K+m`M&6mjg zB>~6)zz;i-Mm7ZNeILk+UA?6Mg5WI3cle)fj{(6*f7P2t9WonDiOJqu3^C6F;S-OY z@r`5y$GpztxR!@OviapRFeL;O04xR{9Lh$3m0<6aai@htah#~m-L?RdPG_pcw%!6sHQ+_vGskDto7{8cZ*B^0pZOKtOy~~4hRrLu}HX00M~0L znEw7|7f>F>xYc{RML=3wj(f<3OiL59Kn8CtKUirL|hmOsEg5Wc+&!&%P z9T&}s*!AIKfIY;$EqwCjlxABk;gO#eYV~+%>N(s_Hd2gGC;hVb+Nbkh6G*w%D+w36Zsq z*(K0aX|wz`Z3RgbfGv2ah5y3R_*jh+QSX4#^%dcI?TbnZkocvPfD&r`wl1~KTP|&{ z)HTbD1y-n;A6X)I>f2arse2ZJ}l!+?jR&64M?l zrh!BV9HSH;{TUXW`NNf!Qm z`X?9ojq0NBE<2DtAwX1e-S_<+hzhH6<5otv)tgT5!^rEWe@DZV+G_5hpb{YW%^ z9qL^uA{PU5FGTyItNWrRf1Zgatp+OM59U8=pYatHBG$*RoKIwfLWC3l1_-;A)*}yb zlDMJ|i!xp8iN?`)s|l5uTLG3Y(1?`LFhxbt@79aU?v|ZGHwTgiCj&;8j33?0y}==wm~Aevm&ta|btL40yct<81*fQ4inUi2Ot*$f<>z9<~9bJZ8?o$NqrA@Yr7j1A+M(=)U{>Kp^$T`ukn@W^Q?d zag{dEJ}vM-LANKI^I_yM&A>Ds$kZS#GiXmwxH*zzTy3N26EkE%9mq>NUIWJCz_4e= zb{D0`f@!dZVdV05jS`kM$9RSctU%cmT1Bhvpymkw^;!dY;PHnma0c9Lp_3Wpq|ju@ zhh0zKHGK);TC;A9j!3I{A5k-!y|%er=^zlfmyO96O=8jP*joVc*`-b6fZJ|k_e$u; z3#AlJ%D6XfI;Je(AOI%9M`;CI74X)X_)YpwG`>RReE$ zK2-ANsME@PZ>B=d5V-pqA-Zl4ujf*`2h{UyQ4IhSmc41Ui3r{Zc%N>A+!ib7iz{@tB6kXc{>)b5}=LT8$6)Vww3&~~X^=rBjae7B)ez+uir z#~A@+pVn`ayIVo6_!vA7Slcbgi)jn-g8+to6Q~X#(MA`beZnmRg0pVkCP2To3T6aj zm6-^;yq-rUY^q??PMRsTbfr^PpZf|(yMQ~oRW+NmBVavyQR!Q}=uWoqO(wYX?hy}M zm4AH}3{*@3^322;4f2C~%K`Ul^r>aeM9bd09K*q#y*glYW9_qehUuId>~TQdSDg4P zU_p@L#01_NvWg-889W2y+#oTitbE-xWB*5{UpV}1eW2y~;Ae@E9-S*von;89<5cHa z5_^pV3{Rbxhx}?+sFoVw_rbit@|Z2w;i`(-QW)Wx$0B(8DT}K8qo|q&Z^9* zgRgdnpQ$(KX`Dg?mUZ1OYhN z{nN3a=^eb0!NZP62C|Ny>muaC!lZeI%s#R&4;7acuQl931SDBh=3NT{x{$PU(dnx1!p#`k$)<^nv^wpH>W zl}as#i#RT&sr%i{&QsQLK-D`A>v%K`EglqkOwOyf5u@(g?x0kHvx(D=n!e$)1Y4>N zcp;ML!5tzr+B9+{dI^%sQ7CN<(zX!78&l~qdlQ(*`)w1PSK1nx{YP7o`zueuInM;G zsy>Db)P84dS>O8ZH<|e%%smZzZ{!fe%8`?rmz+*alSN!2`&*6GWADO8< zJ(!dnKYMzKfvgLL?yg-Z@l#QlLq5jXxmTjTj^g$?o3+l;P>W3^^tPzdtTLpr{uSIQ zjIrjIfLX;7s#_gKXn=xdi}2hYi(CM9`CA*>RHPFn`#&w<>Q8Yb4GOPxHW-Zx_0QcG z70I!6xvKffJ3)^9m9q76G=~jzWtoF6whlj+7bP*qY$*S}#nR7d1MYh>rpsqacEE%Y zw8}tvUt6T7c81B}q|)`$@i0;dNaQG-q;n6UM=&oMb0Vyi0lW&nGFRK}9-9oh7vbe@ zAc#85{~n8zdXPxo2PhWa@Xg{v;lw2pq^ef{!Qox?|1^<4>LHd1mw2;@Tb4aA{5225y@|)Uw?XHPw zdnD-If}iVKpd*ObLniV>R*$x_YqrYG6gHWBA~AO3y1e#u5=e9U#{kCF zbUfw@<0ycfRa)wr%qypGI$Rz^KMW7fhTTa&jB!Wf1`;xXOijfb(mT>T{Pf``(x(TD z^Iv)QL`?b3jhz}ck^#bLJdxIpvjRpg%Nsa8w{H!IyBM_4w1JeT+(oy=e3O@(8=SBM zMsW?~#5$+tyDX2TcIo35sXZ(DyLny;-a+eDTxRW50GfkYiZ?v7pS`!A48&APWm^nL zJwO6kP7d)Qq1-^vKi`S~%k`p4J%3tNkeHTpf~xfjMjyJ~W@7trZJax5Vt-MCC7aziY96 z`T%UK8pJ_x#(Bwsze-D@(v3a^T=IcS=x1o1oI)z4cxA?Mn1liep-TCJq?AMXs!o?z*o*2jGLseo?d*6@0pVpl)z62)skG*4Xi0=mDkS{PXP9V*_5CY8Ga^Y{sMs$!zT*3WD}Ox@ z+Wd-u;$1aBea`e}c(f9cyxx^=lo)RvPQ4OwH{Uj#1o(de6FflxHex1br%jP3oy?O5 znYZ>J$_AaT;$C|_uEPE88w3?^{yZC*Yzo6O_|6&b8Z8o8@*3 zi)h#|2e4<;T-OLKJGMcizowBC1B}ab%JR0bW+(Wy;y#OG4em!uVFBdL>RU#ueQ5lT z=J1`HZszM0>pkMa3}W<|Rsxe{DNEUi+2ol8aw|u>3g5U2W51b0Tke?7*;cogl@>~3 zL16x4-YKamP1xtGs~yo`%?ujHk!Ha28QAeRw+}IQFy-w*MZr5ii1R@IgfRl5+t_a> z{QDtpqG*K#JxADr_3(=EE;!6q#x=MBYmn!>`6bDRdZ_?Kn{Mu1cRIvwmsHaJ z9jc@DJqmfnuxd1f1x{d7wG?hpM zOs~zO@`FXfIJ$H$@10mv*8LU|n4Q+=^FK0cmG;Xm;tM$r z?S*q(-91$|Cxcx3Ww5nhsYY;ET?0He(LNM^sET~(FZ9#K(Z}OvSV(iruAgM!G~WBW zo;Xam-krVKsx1b|4(Xgjxsl{Lr= z?CGlE!eo`#kGhGCoi8}*7R=77_I=mfY4&Kq#D&Pn1EjbKd)-=M)J!ILIwZI(2hf{3 z7_a?{8xs|AOx`Jg?lN@}aSQ#a*G0R22mX@>eN(f8PxK7Kzn@o`ficY=yMcYtm7p}F zBeR?7l{`(&qAe{<%NPOiIR_?H`r_@k!{`A(dr-Uu|JjH$-QcwByUFjCEg0kEL7gXo z4PEOY^56mI0XL=+>7MD7S%91KTylFxWn!L&(BhRr@3@Za z7d!s87?zK|#j3SI{WkPyZ8>eCw{J8YTQ|=$&b;J?bUw5d`+&|mmaQQ`L%iIo{CE^9e z`eBE*xI;@_;c1@X^>pAwyhyE;ag{9U8r@;D0I&3Wx+P4B_{ z#)|1$vTBCFP&g-y&t@Gi$=e*~``SRq#VeY-}y|3$6SHKGWZw3+M zNRW*ExRszJDgI%_W2XahzS}C(s`uj|L(Enz7+RuHNf=}hA2Xb^Lx ztx3S5{tD0zk%_FpMUnYeugvUw%KpHO^{LX*9LzL>KOe*b;4!V5MT3aJ-!Im7sD?hH zrd4;i@bs;LpzpU2{muTOvo6P=+V#j+ekWMne21Um2%4lPqUJs7k9`W$WpCUf#`^={ zNw$tonDoQnLHk^0aL##Eg#12szSIVVJ}T#qM2U=ex6_sT5R7s7o;jEOLtu9Ye?y>e*>fAdkB`3vTk%C*Cxlizi>)(C;QuQO`M_+zI)?Sxq47hSUkoP4B z>K#GW`Xcxm-2~}GVj}dGX zt*pQl1CEbr9K@hNE`}s*i^2(m8jaw}9knS<0(C5_z{e14mzZtsq~cvxyPR17r$nHg zF@HfEHHqB{SO^6qtJnmZ1G~4-{d4AS-9Db1w-+lS&YzLjy-xjWqJEv%g2IHxUuN6! zP4Lf+u$TlX5jsSiYwg27A=5)4BNGS9dSqhKGyuHV&mSe8$hN&O0$VCS(n9|v;do${ zM$v6jLYu9H6y0S8VKsMKKiA$)qyNi;EacAM&%JB*4$%G_0~yFN)l(p0zzeP@iTL}u z>Ep_PWULDAFD*fnls!Mh<7`SL7kFPJ&Ugt4H&0)Yc^)<< z>0d8&R>f=PcYuV{Ma6EUyQ+>ruEgqtt*$K&NkjV+P`8wqf4uQ#Asm zxc2sKD8d=qlkE@!9M>BVs54m4U?16!4fK#hTKf)j_+wBerHsy4OYSX|;*Hu;H6{T1 z0TDt$j8Lfv?Qk`l{GFm+%(Be=6A2}V!eShFT|bm1P56fk>n=)i4sM{K6TL-8Iy5CD z`aN&=TSqW+XCir}mD^_n&Wr7Lv-T&@MEb0@J&|&%=5%bu@3tOeqmJ~sx?1(g7}`eg zBQG3#hRihjIM-0w>{n54Jme&H{nGh6=CSL%Cw)W%Yn!B78q@GCte{6);EjFqgNMiG z80U}Aam4el5>^vEV#;ti{;q+Qf92^|-j~O>{AJv8=cX3ta31Cr_uOAGQf2$nw$p8^ z&c`28%y8Og9BkeqUiqm~&aoC>lHDN6&9*ZB2r{fTG;wsMtlA09`6|3YB(2X=Y5Hf* z4B}R7O5cMFzKj%Vh*EFI{;o)%<{gn1Xn)AVH*0M~t$q}V;}7oOH4jevl8)e?7Z>^0 zYD}Qyp1=8~MwYemv2n)$tT}<4;(#!&GZLc z%$zAc&4>|Ju?SrM=({bD?~h3T#r=Bs=)g|CSiT;hyL75=4o5Spzisz)HnoBwWk;r%sLPv;ofM@&@ zfXf=gg{`ZT*>IWA4)q+pvHl{sBuyp>nP-*f_gG$&gVoMHgBtG;fXe}j&o^9RKA&?r z?eTH%Gd2YyrB9nRAoW`lwX-{A47;>ZNYJj*s>dAR#h*ZoM^yyX1UF0D(zFkd)PNG* z1K;I@6^qu+2gKZev1{(yawk;?HT-L$OQI7igQ|b~79yyYwAKILE5949BF%u!$3B2M zAV;mrk!Wy6cE?>yhm;|BbuXwGQ?`rN5;6Mg{NBu=_ zY!j9Obu$hsL~sE@_|W~aXL^`v5UgGT)GCUM8LJ|;>mm{kF#^5?1Phz9jnYQAJw6a5 zP=y2uaZi3QU={G1RxQzp`~WCG^r?IhfN~(aY^1?PQbBWO*%s2$Lg;!6aCoR;kxTr= z(DXt^%sZnNs9A%s05S=kpeI;Gq0xnGvIum~FyshTt38dCa^%9j*WK_=i5HZ;L3>~H z2+sfh4*w4_8CdmIFp7ZQVb#7veqfca@Ye%BZ7ZN3AWfYLWEut!q#)}DmU|# z+5ws!`VxMzL`WgRRc8>3$v9~bqzOZ}c^5y@>}E7b zIKbn0eZ6cL_^O5zh&(z5*iX!=*(DOsKGbSxX#%)|a{urNhnN`8Gh*qYi*R3}qj569 zt{&1cuRa5tC_2f#mJpQD67bU^XJZHOx7((lcUpGdA)h@2HE~q{5KModfA_k6a=6vx z{GNm1Lw^uZWZ+M1-xoBm6Nx^Bg2-vLJEnRp2)w)?G=$&4A4+IQUUR%tDgn3hySQY6 z`_YjBW4YjB_17XwFC3Tyn6b_rpkbhR;Jh6DT?b+f)C?p5rm5C+g13iBIz5e^#fhZ# zV$f05s|bjr`VFSnER~oCpP>|iA3+P_%$50^*xY$%P#h~}(QuiSMY)G&6dHS7Y3Rmh zZd+5|t3|?|+HnP-o8`3=Ej44o*jFbD-@31q4W^(v6X4zW$$nd9i&x^*sVzJ~@-bquUWR5{b#3kgx^E(+q&b zW7wRog9&SAe%l;`v2H1sKK#o=V;(udbLUs-D)$0~UE~`q#V>qp4gLi>Y7@C1ltWsQ z3dzjj>A7JvQXzVs5s#eT9esu-RGM($@#4e9NJfPqXz5q!ppbQ1lB!`YDVaL%tIPT^ z`R;t-K@>SF!!vuMd8WEYd_fU*gYGQ{rexRGO)1y$5|j?Xi9p)ve7tmyiLNK_?&pRP zGLqI4XrLq<=DxD~bov*zLBtOaD2F$Bz3SaRH#Csp++4$Cin7KhOg@39@Zi<$pOuij zV4GUy_&1{5rW%MsFztbn-D}D^)%X&E=7ueJfa%`v1temaeEo;}Vm*_-9)q`+k6_e1 z!WS5GpQDE8ZMQpaV@0`-dm!So_Pr@l6EfTp4N)LHC`fD8ahWTk|e1)QKzS z)Zj$oAozRZ8YscLI)wc)i6e^P=uzSq-ow6pv02c+<`I0+53nTNzQkXoW5HAnbYXJ- zUo>^hL>BrzPCEhluQtCC_^Min4@%n%S!81f%N8^ftS8W=Or^cmmo@%4Uv4B>kYx^e zPe1EWrLrUyQ~AK_;uBWA+j@FZ5yC#WJn2@Vv{Vce^A!nz*_WtJ`MJQ2h(Tp;C$U_7 zKCUu)Pr?Q;0FL0we^t}_d5xn>aH8nsk7i`8GuvY-eEF+8(DrlAR;%_9u!)83qi~25 z)*ut{ew8!OunAO7hN-r!NReY`E!JC)^jt*w9arrx3YJk=QC}28&f}y?>>rYJv8t~x^p8oy7yCM_klH^CB{ zV8K^QTMyd8wEimIG%U)vq`d~pxB~O+yHWJN$v=xj=8*3m$h?rPIqB=U&tj!Y$Xkfy zNnLf`>(Ha}u(~9L&bZbojcqswX*pC6&*iMKJ0jHx7%xyzzW=y(oas0WEjP`>){V-h z^*_cGdOhcPB5~CFcXyZ^Zw%~-npfRDE~U1kz;JvhF|D}8+v&cDZh^`z)8j@q!kF~% zT9#A^SOh_6Z97vQ>p!^tvcQtYxdoZzI2E%Pg%_7C)ot_6?#odo+m4##y@qK+BAGnB zWlSi8#CYUgJJI>QxUGx!%hWc*-NvPA1Dzsw&hxE`+7NRVPD+`wH8>QSD{JCbBz8u` z_t!g72&_lA9<&p3E~rwE+PTpvJB_3eEIJ?O%r%xf7_da(w1mCtxrroX6Zcqi`}mSR zTDo#5*-Nd>F5Q@>i=gxA-5Ij1#nm@$y41}VgA~Hv`ijz~@=1Qqs4)tOl5@>&7{2kz zJjCPt-OUGH<*tTah0X&gCrakvh~K%21p=`SCYum>Fms$1j-JST z?|0qIM|>8L?H&c$NVRb8Zw+h;Ll)w%T^I9uMW`*8I+7&?J+r8IE!bxb9vGzWH`lX{ zG>JUK`-CaMET$87YMRd3lOMh-eE5L$zYXnAUi8Y!N}fhdw~kCZ$SrP&Dw&w*3G0^? zPNaD5xtf9%2OyiuZ;mD+?QJ@`pQBo}@M9w-Kd=N3m-l4liFBi%&P5#+mKDvTB)^{d zdVLaMK8GYdRh53AC^SdBCM$nZ^V_(}Ms1Vtj6PmV-&s$A)PQWriYh;k5EMl5f~K2y*8_fqLDgD zsO@}U*1Ef482)@&1fN!ro|b|qCe?_?ZzYuaeVe%SYdG)kE(P=Ep_rT~maJ619na*WkvS*rPzgSJE=> z>(5Iit;immXE>ceA;PsB>E5YR$#sV$P}A)biA2kI)B2&WM5Rc0-g$~3Bg2e(A1N^- z>Pdcw#`@f60&JgK0`F(qhy4&MJ8M}Ar*L|ZP5x%#!v_kTtRh5Dg5JH1d2=;^QY(yV zWn)K>>jm|l+}EgUeI$RmWe>}g=7kGATVW;=drmjb@>}pPY5Braxa)PVJ8)2(OA4KM zvoma;8KZ3H-shQ%^j{j|QjdvW`b%5!ePB~)`}QNpXH}~)Lv~`0vb-Z`f+E}Wu5FnO zsejH&f3J^&8`@GK=v6Yr;mj)oDI8lb9wOV#=7kfe;+!tEP*P~`Io2_b4q(V*&1nKXC|`4NoU{BK89@$5(*xr zQu2a1cQ;9RdjIm@-)T7`Qpr_)ICpwruAa+HKoC!xn?1UIWg$$q8;&xjoNpo-mTuGS zzJ-e#;wb@6(fGItz1y;*R}cn6#d@)lWh3By&=2VLapjjNA?ShKgQxUc)gax zv&pu)S>)6M!Vo_gsx;48)FphB2~ughcAKF3!LXI>j7ZZ>lv&;x2*@3pZScQFgqGe{ z;!G!gnQ&fZKjW3p#ojm8Q}0D~0iQtY>PuP>ZyK*wx0#;HWuH5F9Ug!^H0KSQeIobG zR&RE1heP4!0D#PW@pS8gRuraDxHMV)q`Rj$_jX0G|Q8isZy>X1;zp{(903vPLd`0lEYo^h0}F z&?j$rYVB)o0{D#`ZZrP~r>J!Q@+@~=Lu3ybl& z&VUuM;}qLA`QI?Ux=ki{f&_)ckf6`Kt3C~lcR!vUkQRE=hId@>A{n^MjYqCGIMHj- zB2V@Aru`|nZDD(uL;MnSIu5$_WA;N^1v2c>4N~m|O;%Gnyoohw(%~P<-Ikh7u9T~a z5f1es7m`5ZD?Kw>=(oc`s(S` z8nHC$mBwcNPFE_+gFVQmTZ0(f#qqk^<30ZQQaFp=XDz3lHm?VNlxtKR*PC!D}G4);Y^W(A*1y^(LHZeRHtl!o8(9C#+)cOR=bpY*yqsIimEdo5`~Sn=ftEXCpZJp*heL$U)F&c25M zN!fCD0pa}2^vFq(bB1;xxyuzBNsa#N%RUe%#vE~r_dg(2dg*)|t%7Z}BJXE<{xz~~ z4E*@^^c&%_0<-5SQ~ftaAZXy0KfESK8%KHkVJ<=e& zK01Iy+R3nh^=3nBDZAh~X^ssg^dxSkCskKwi|6u{_GjH#?_Q;Z`8p#gr5pJA8 z77*LDnX$5A>Dh--9L>C5vTTX zGsQ)N6xz#1=}8Z<4;rMo8`@T3_hzD5?+xWS7guM{^Tgdb@bovQ}1h> z)pcUZK0lw)ICx1DYI!l@hN~fGzganYh+De0D;VXNRLy5=>_#E%h2JPjlAZcm_pIOM zQ17fV{eZN6=Ju!aV>uJ!jk;etTStWlXta57_?$-n{8` zD5ilNZ~D2jU^yDu*d*jHA{hGEh0@%E_$))1OI1&Mg+4 zv5K^MJPo~XX{I4x%UYJ%3Xj&){vB156_1moKzbQ}Zg2`;KeU2X-(BWp?x>P|8xm+p z_qG+!35XqYZ`5Z>#80^95yrKaae+1p7zbKi~Ne&iWLYDYxT0OU)WEjn?;cs^XyS>UF=w>zi8ND^W` z^YO1{Wr6aqp9B4ost@S~l53wnC8oZ4tReK+hc|mYQ=apXP`u(km%q7`?i6k#1NSfu zf?snc+3Q}jY8x|r?e(yzjUQpECKaI`!-}{tl{vCKhMBU2^$#~F4 z3xVPw_o(Ls7o)h&VCd&GtaJtEHA)d-@xcRqG7DZ^9E|wY|2)BlMO|=6pWt_gGC_&~;w2`|zVhgDzLtKDaWCkvjUE9$8tG5; zmaDm53UF(0Y|Lm+g4cI!G5`RdSaI0r5gC^en|F_Jpa6ek_ZURaAVE9>!CURWA3?2` zf{2thb+V*Ype~Ew2;>icL{+k6%T6;_k)Z91_?@f7C}t5;s{rueJ&o(2b;l-a15OUg z&bogAd+uZfFFSLT*#y;9#`Xf{zbJmLq@i~bn7vCt4-Aq{?A#0pH@?3)xeK=}`XRuY z@ZE0%egkn30M}%2YY;>%!(|_Gv`DzN2)7wxp$;%VP_RB`S2_d~^#SPLeO3mvt|N}I zX52x>^gz%K1HspK2#G?#*#xg*n5B3f!G}QI1<{1xr)t)~u8fWFg`9ODVVfou7gr>cJ-!ApTu*6txBks3bQ1(vk#+G2guzem!9;_96G zMz;MhtI%BrG4u}JIPe1+6r~1F+@YNMz!n;Mqz4Qy?a^{(0mf2uYhD_Xo zhq=){T(^^zAx{j{wtpa~JoJ-#7Mqe^Iiy)XXCG^Z5LOU#WP?lr5**!|w(8(-w^X#A zz^?{926qSvBt3v=twtK_a{Qbn-EEM6Ns$%HB^L#N3`bY!8MW4QgBQl_eVekIeyWF zZ2U8tr;GgZvxuyFXH=ven0}CkB|igEAU2er6F$ovqL(y@aL}ON#fIW%Zrh4(hek|e z^MD}6gIWsU=N&+0vDx$4G23*HdOt)$D-(C;z_SW|Ha6r$%z=!hz# zMm+!y<=bS+PNWV4R1ZWaG4xwJ!hD`JkI3*ZVAZ`X5L&f7w^w6wzevW!{I>H6y z!B_M>*gcMV5C47JqL2(unu#OaL>X`d6o_GC>Fs+d4zcV>j7-`pEZQ7iq5~{z@yc~0 zD65O6JmzLQ(6E2YfenL{6Ce!7o9=rQ7bouy@*U?>t=ihpSw} zE&`ehc$Al_fC!-))N{ZzNq5Hj&~XLmp>SYntpmgEB3 zCr^;$4ts!F1N6s=av*oG{S7BxHK>Qc#V-oEKS0v-=R@%PxW_P&qOT=@u?oCuVfa8K zanXH(EkjJpUl4<3l;IHNxpg)F(rpvOJFJxD26MaMk^!7YToK(maXK%+C-D8J2yB1e zKX44t+!2O=K!}C`B7oje(fGk>Fr>Zf!Pp}Q=P28cDTivb5Q}pk-av4oKWUja*ZDW# z$;*doPTGloJt9PKmlN`5{giL~1RrZwBE*D&5_A&g9YhY?nW$usi(E(zIA1iMyZE9k z|6=fojBzo5{rSL!L80s<;2p`v@$l5kk3-21)2E)6ic1H)yVeKr4-Ed?1b5gIDmT!h z@s%7{V}=6afbjy!VuZt&s0-#m!ol%ZB(es`dw(*KIq2mrBPr2IQE&|T!WIDTt-`1< z?E3ggMe#rRSn`xuMnt%$SdMddt8IT25?+PBtE~bbw*qOilALfee136VVB=QGIj?(= zPa^!CJ0x>f;3=V>O~0h!6)Zz84~;?A`-q zBlF9!2<~sjQp7IX?sRS0b^|&hm zmNh5vc=Zge3U{!s*Ftgt2pW$e8>!lkp5*eQXW*4O(*xa8GjFRjZ|vM9LT9TJ3MI6 z_>ZBc=KKnR73fF7%5Wbfb|~eUdJh4+s0fW`W?t{%P=I)u_g^sv3(Iu>Q&cqbK3p4y5vz;A1A;NNj+4O2vC6$w_5ER_Y>U(k{x zM-j92Qaxh3KPwWvG~vP=FENNtorB|o9?wqr>5*S0Qz|H^U?Xi!)%3!f%%Ex{xQ4o40VR3N50Q>A>SX0MU%5%bZ@+IN^nX%d*g-m4dP?_&Of^#+#J`)NL zQ%p=n(9%q!oi4n5tg?3;v>2Druz&r&bIVx*=?(|!AJ9#{Q7?t8CJ<4ePw9a-{7&j^ z12^Mzy@$s?B_(wm@ii1RVKXGFfW;j+dBUw;5c~rw z-bfB%(qVW~*fM~NIHnAodETkc0jL zmKmTH%SP|>9)VDmZSJnJ>F*M4C*hIzxmy*sMIxsm|Y#Is<*jGuzIS72}q*V}ubF5eHN_@WW$eVghF z{CX^MZcIu|_crQ)w(c;q*nmqX?ArvKKE_3wASk#jp5@G__}M_brE3`cBH^7aYfH%N zlyfXz-+-+Gwz(}t7`EK4i!@rZnjYiAO)bD#*+^QiQT4}I#v4R@WuFLL7tWirsXj)>gN}kK?6W9px`mr zysE~P9u=@hNgg%#?v?fyQl;Ir$?xuA0WDD9VfT}Iw+FzVimf!0KVU;6;U1U`&fSJ- zQBHlyi$LA?9>DRVbldaIsA_I4T!gxs%aJ<7d=?O>y%hQl70qaQOhsh($D{2`o7qS$ zMkaz;EeAWW1ZIT5TiXmFIzlM|X@qSIWwQ+x{$p^rba3|w z_4e{9)IcQY;Liq*F(SsNNRkKVLeY+c4Ex?Rr$;bQvA7_iCIlGW0u9*z@2opm;I3zU z`7Xl;;(78Y>CoeT4R=QpVc7+4L>^u~fNF+pEr*EGyH8MKPx2uQ*HdRB-ioe6t1rV< z?4U&q*rk0$mM~`oa|&}U-y0q)O7@(+f+XRj^NA7(P|G;%7h7JXx_L(!X0^h*NSJ`qS}Q9sPwYyf&F)S@5t$g z%kq>BNqV^sYD|pB?XOPh5&ryPSu1#z*~B(9MbCt@4c7(eGHw+o0_p7Z$0h2Y%E3!# zvV%;qwwF-F9h%kziO#$?^WUG!+%}o9)NY|N*WjXus-?x@IQJ`T0W|h5&-J%cfQ#_~ zq>aVATJV{X7oK)+k#;`)wb&W0WUy4G(#lz=p}hA*g2K5z=+4krtL4^u*9i(UDc4UT zOS-GcMweev?IEN-!N~$|EY*vJp=c7T-Ng@NtDEU0d<*i8tYXu}KRa3VsBX851znE^!A~UwTF&|X>P`o2DfctU(ka&!(7|E_o*gYce4#zne z->cWfi(d&nrgp;4*BnoNM5?yNOWU@n`A=P7m$zf}+g4+p2kz;^jlnz?;?Eh79tYB< zg07e~s^{%nptO?r;ziIy&hQG)RYfoEA_i+11UOCn#~%u<2s6(odg0ND5<9DgT<&Ff zznpO}v}UW?N=aHDGDOT$FU7W?0zX2XtVcvq#^$JwR-~TMOf(W1_Dq)MY)eLYQ?SSo zHKxJ>Uq1F%YxllW8%NCh;C*+9etc`uiFV6K2Y!qZV@w3s>(`f`C;vs&xD=V35^y2G zWC|)M8S##sS2i2MOS%zyiUY0pMhvih<_inPZ+6>;8(4{nh5oOWL$?uHj9lRbTwg)Y zZ^O);FXOTH|HA6zNSIfK_6Xv;tV39HAaE9%I?_>lPXE2GsW>3?l%rZnQt?E8m;Z?A zh0lw@T3>p*5C^1And0ZDD$QB*oH15`Te4Tfw^cxP?ScDc94z(V-?Tjra^GLZi+HpX zt4LDM^oKQL42acP7jK@S^PGaY@Hm%Si~Xp03)>nfC8K_P``je2NUM^4DZm>VW|+|# z?Jm=i;JtK$r3u{NL=A+gWHD)sR{o&|-RRK0-TZZE;Wh7?s1G=)M~3Y0LgkWz9Xyiq zsB%A0Oy&+0)7UYcrkj%QQPfvuaD;tsI=iU_z5eT!sd(F|)EjKecvhMvX|D}k3^cn# z_asLL;gs-Q9|Z^IR;nLxi)E|^G4h;f` zBYP`f+wE8>PTnbis`YF!%Xihd9%=V)(U{oC+;M9$n>*iKca4rnJ!}|7Q=yX*#;L`W zd|`yEGA|)!ZkgN5SH+n*DJ$9oj#;T~Gsk*iq>#o03jji~!y9`V+yjgwK6|Ri&tQ z>MrA74DM8d6g*BH{0|Bhsf94KkN*u;R`!fFJJQZC(M#(Id2h`|{bFqZ z4HPl>pY~ym>t@ogstDDhHkl|k@eiVAMvWtKq)Cvh@THV2XYmlK|pq!7dqpZ zrLa@=mXeT?NEi>SvI+-iJOBUTd>Oo^Y5o zji`*J52t6h?TJ+sUfqHW6zuX*%cDwk+{m>Bpc8S65PszVT4r?%JUjQQa+gc?J!8cQ zx$rx<5;G7sg&(vNM24i3!H>nb8vY>5oL}-yPE5ce%Fxh{rQY+*{U_S{GB*?=Z ziLRlt1j-!0VE@3tDiy4_lO{41NdPse<};!}n<*G4J@WAHUxFAeGOD){Y3;CB)I#d= zNI6K+H|2c)RQ2y}DCHb$1Ruy+{{Cnpnz!xM#NzilB8U%)Y~iC*a9Qb=C{y~eSbPju zA*%@M+3>$pTmL^?RFo?(hK5%cn^OT{_8tH2vfVb`x+N5ZxL^|w0N~Qe+1-Wu2}I(P zfNDY)HxDAsG)Qs~-UKJ-8ql-TP`JWP@cy6q(W^5svjq~BUjTyi&-y$Atb4Y>%dHuj zG4Sh4DM3a%8pixCK;;c65_Y|yz8~gs0=J(Ui0rHGKnTZ0sAND94tbsig94LBfiP?Y znmj;s`QdQJntN8_E2id2unR|EAG`q9NbP+!47BRx2MS4!;x!^FOq8&>(JIzorb6 zR9d0gWZiQ1<|j{m4ZsLCa)pn36M!Kg`0qMgYxo1bKPa6gG+*$d0xAz@%bWKkR(YGfp@DmQ zfBR|xR%Y#nh?Q)B4h^zbh00iA-@xvf(woJCy3Z`IBSE4gWkOsXh`HVEcF^}hQV$J+ za1jh9FM>YD&=-@ucq{WV^z>h_p~|bi!DQ~uRSTH6m`THzPlAwMdb0$jFehrz`c6FD z9`z!A8=uIt02sJ>Q-722p+*pqC?LVmscJ8nwK~12m%ODpS7q+|V#d9RoDy#){2YKr z^rh=SBq2wD6;B;om*1pvZiX4={h>kCPD3XE>p#2pl*2XdN_So1IM8J=IB-7o45~su z3%_jESp`$L@Z+EFd!V)I4|43j08PHqy>B;=(UaX(qznisP40p}^g(CbwWVDa;WAbggaW`;@{=))}84kh*)O&0!MoN%WK{250wXGVx1UpnMBCQqH38H!h zGFM<^ap#g7K;>x>>E9;NJo9#S3RF+=lCG~O^c8GiPS*|`D*&O}yAfPp0=X#WrlF3ZDO^A{qm}ToLN@?^5^1X}Oxy&s zLuXb#z2FCmd8rI)`eT3&0kPJjfQ|hgDni2)Bo~iJ6wWHulAQl(8%#CsAy(T1XFzp_ z7FVRSQ~-iCk@W4Z1F6G7t1xJ7IS=zZH9jX zE;1WoK#bq$dTkCEnwTp3dZm~sO||2n0h@*!C1e!`IJBEQ3_vrj}>kos2XgG{&W zt62|k?f5O=}>x#dIr!VHl(GJgl387urn5vUKpLS6h{Jh zhP~(-wvBqXRXT|F=cBs)KIKoZp-~!4~K4q78XUn9d zhJ-Hs^mB5lisNRmi!KF{O`3bOL_T@{BmTz{(vSy+afjh*e$J8t26Ntl<8{G~$a`bb z%x{7axFQpF9QSsg@xW1_r=ClawP&uT>W)Xwj2c?ODHPD8jLi4^-p0B0I@oY}C%l9B z7cWgpNKE1$_jL%y>oYE;jtW0GngpN`0PD~L-h@o0`|f2Qm%r77zG|NbPGG9O?W7_F z3f$df59>LqOYNPn8T>-JkWALu0SWzC-BRiA!DX1JX3ODpkSKHlRv$j2_jJ+pHd#HB z)MRP&^ebhDl=(Ix6^yx@in5t8&oEOFj;s&Ws`qe`mHHeq znQ)+09>4L4g{*9)>FfB3^~Y#i`>!da+CLEty$1_4V-2|{fNXrx&(V3C=yk!U;M|kY zH&#l7ClM>s(O>KjJ-LB`jA0hpw#`HhfPpF)U?zbafkAt=G<^ zT8dhRlLv`|U~u8?U}~6(A!qQ#)H@|&W=B_prv+DSP$-8g1&ks@Uj4`xl_Nr$wdQP| z&go`>%Z`W2NJkqLU`wJU?w?qGHv{E#Mk4J@0i?5iY@+ub^`Oi5lh#b^EAOtr#zY{3 zA73njUo>E(jXHqXkJ zi}Hy_Q#l4BXl+>NqVYm?Fduk}-=nW$O6-Kv^KEt304-NLq7b_FqA&BgSRnEqXd{+X zn{m*B+!w;y1EH*?RzwjRb_2<BajxGrfyebAXBe>`LJk?1!d-5+OV>=--e zdOk_Dz?G1ruSnl_58slXb@-Yt1JswzL0Y{?GbJ++gzi)8;`}A9i_f&tVwEP7H zv!KjI+jMxB_H02g5TnVJf z@yW8uTM2Ao7&Z8!Z2F$LSoDCQ*ozAMb$q3rl!VwqXboJ9QHn=_voxLurU%{7Kdi+?L@1 zLEKgCKwRIyMy~2vT4*yE-ziVc;wUDsIg@P|MU9I*kXRS6Q7#uK|c}e;_@;0ynv9 zvQ;^uky1s=u}#^5aNnNLz$o^5K_2)xZ69 z#p7K@uUfBouaOzI#w@}0dZBIkc?Us!dNV2wwh#*zy4= zz1g}gq+AYg+?bU+Mb6LUw{sKp5+kT~nOE=duW|X|ziPi>-}dM8tQy1p*Wm&WuR6b% z4#i|z1Hhs4D4MSZ&}Io*sSLB<7hINQstMZUWGpbj3$_&kQ7%j4sj*8e*^dYGG(Y{TjI=&ERv5Iy8lrzaOyb{(u^M8wlp(kgL#Dd+9WEh`3w- zrhcjj(PSV~!3u5=W2mPlG1vW0uswU%&85>7&c88b_)FOTwKNhlgusCABOhWNgny@i znHaQC;MVRVc?ug8`maxMX+3t6AdQt(=WlhEAj^133UpwT03IK4+BiD#8!mfU@NHGm9;%-Kcgfu{no zz|(zN0~sDBEs)_3;|5JbaI#OgP=nB$Acgy5b#sVrlESswQb@*C;)@_+BrpcB z!%}J4-8XR8mFsj2tf~;Tg+!C27DJRL&^gerKWV9fiVW*FUyAZ1&!o?8^eu^{mu8IT z7tvJ5WJTTi{Pqn*5<6SWdpJ7GT|@1he98o$DaiSB;j#!57|fDhgGX+m1pDGw!!P#B zId@eosu^yom?Dxx_OoDdU0*#dWt=02JbY`XSKuOLyuD;&6OiC|d7$jBbx zzBhqq?wSWTG~fv?NpG=@Nx6P9S39V+p!WJQbbp9|T!Q~PevFmpuDa?OZKoNweQCtLT zIY)b~Q7WWU1VDccbK^C?jTGV;JOI79EiRTUhSf1FoC3rU(#u1|?YdbuslYN8eZpPP@BFKPVO6-2ahNI)`D&E~ z!sHi0$OqB~K7(oavdUKzz-v;r->TQ>)E znh zymMMz@qona2?a=F-T=99A6Yq8Al8jP$l;{>^4$l0TFEQaO*}SQQQ){PL7%THTofdm z2X$BT>bpIyeJ}`JhY?cVzaT>3>Qj|1DQ&B8dJ-|8{!3WE;JVO-RF&WpApDJFZ?X(N z&8o4ZeDbx>1bbnU56f-!J9c+2UZ8a-riKwXef5=W5Vv3i)*3($F+n0AE(A2-{>rAT zt;GtUY`I!@o7x)5IM4Lc+Nh{`hdZxeaGfL{dkqKma`g^NI!Y8`fy01IWy6Q~Nv+`O zVVFayO?KXZ@%DY7)3l);3jS6Z1hwl^SkSz)AP2ngwV^DJsm@DEK}Vj_ohK#W56V3Q z^Lo!fb`CkJIIA1gILYjCv1k2X(E}C;+%23+gAk+ZH{=x9&-WyN7RLkWdcg#=fgUVv ziqJ?bL5h#Mr6(7CJHN470<#&~HbKlJO=eOejhNFcuSC~v9!|y|?nKlu@Me!mP5fY^So`?-n-q14*1 z3ED>IGomguf5#iVqBjYuaOUO`;Juh7-R?+s5dPA0x{)Q%nQjhC1Su}pGg;9ArOyHB zTz&_~8|zIG-4v=tl(sNiOrfpLZH9enpP0@kBRX*59j(_JC@sFV*Y`>tJ9$@awYNCH z++;9lKb`XZ=(0onX%9~qd*NB4gb(>eirrN=VlR#5t$o1=!Ne3r1|MvUZJ2R_#h~095pO01GP?qNtKwDJBwbC1 z85mOih|8hFfpW(1F^JFBHl21|APx&9kUd;U7v%(?FcZt2JFDie6(y9iw5V?@;9)*m z*?*Bm7A{CVp-T`x>KAxwW2lCg(&|r6YJnWtstCL=4x!d zSSU)<+Qvx)8VY~;TJOem>l22^nfl_pktM1cIj2^;tZ&JbZG?6fLX9anA%l2^L2OB( zwjU*a^cM72fL79Q1XHK>y|_Tuh2+;V@XknyPFAjg=sRzM>Yqrs)`V!gO0?K{K(hH) z430;*B4b+k^Wrx{`{yc)C&3weBN-3PzItCoSoC)Hqd}F9D;n3S)VfEOOpvYCl+IHb zRLK<)8vW5?wIDmG<@2WSz+7sL8F19gFkn7yCo5;aZCFN5J+OI6Nd?D>$l!zT%}<59 zs?Fh!5QI}{sM|Ak8!q1l6FLspt)Z!R;7DHEQKMZnSQO3bla8OQR9lmG=61Jf4xUpI z!(}dp#Z8_Hfn#UYKHWZx)yJ6Vzjgn0zxfn63u5WeVYREs(Dvj)6HMI%QvkVX zOaq3G^ETn35veAT`uU%o8T`IQlE+PklyO8KBf8b{M?VDp1w)zdd68COjHtx-lJ?A5tV^V6h%18+fx~YA3X)fA`;B7)_ zmCPve8s1L2ShVe1v@^ru%;TMl`@2g+S;yUVNuilRm{g9n74gHd;|9;2exhPW+%vh{ zc*PHU6=m7VPDS@`xp=TJ2F+AjDwXpPv_i)jnrYlZNBbhTh7Cv9l`l$VXpnX{u~pFJv*Lb}+{zY69zmdUX*pr1r22YKwNA=|+;_m#Ztj9!O9R|{g zpQM{T+5(k}ggA=uee*FB%53vjQ?p_Xf0u)1%u~a@A&)8D*evdWE{Cq7^x#Ut_7j?R z8|}`8EWXQzvKwKTvQVpKR>2_v^r!5<7Zi|u>BnWFZSm1crq`dy`DBS=8ie^TT{Hz* zo@VcgwZY%;Ozk=Q_{Q*`NQYC$HwkssAJaDkl^*inUwicwO*_0~%-+Ro@@aM!B5&Bk$jBIT&1{DKAJn~dRMu_REqbdD z28xIR3W_Kw5(3gvq6nyjw1l+MrF3JVh#&^t-6h>%U{KPCltqImDZS^sr@ZbImzdME=j_2;a4YQ`2#e^qM7gYFt)4;D!t$(mob> zv(JU;S{D9$uNc0k?Y%8}_!wO~^K30#(Zg~Ij2rH+G*`ZnlHC#BGAX*8z`y6yX}j%* zJ8#kznH{{qEoJp&k9^Gnr?*$r?VY(iGxS+!OLtwF+ukbX!@IOi$S`DDzyP>M>hC6j zZ_MXT9okypHSTP@XGtn=_qB5)Nyo=1uWWBgHW#1yp;_?G^1q*~`Kxpfd*ehzoPYWo zUfWI2V{}i;jo(?B?|8VW4|rvNY3D-XWZ^Sm&l*?sni?eV{JwNx5`*3Dy}!}Q`8S~^ zX-EIU4U6F>kXr|4`r~*#mmTVL(1b+-P3B4yJjnk^;-nFB$MAo&+IPk zZ344)M_hPhPhR`IyI{vF2(X<;DfVC&o#fXKQSXC2O^LorC-P~e?$2n1uYIl>`oDNJ z0XcjVakn2c9S^q;hYmu$!St2zG^SCk5}19;zezSTjbpcMq>iO@qv5*y>z&W)_Ct4h zN5~VRORoprejd@wyVT(01;_u=tOIfy^z{1@RX%PLri2O5+(m&|gDPFp?!IE{gAcyE zX)&xE(T48@1=>!GZVjMk8cG+1F>s-ib2MuaQvRbxX!m!(E#2d^y%V`ir`4vnorRls z(Q3qey1dUYvGITgvx2TLv)0sHicHuIa!e%s_{sfrsAbFTzu|iIamt?294!cgiBFINmN9P#JH~H{?ysK5+eX-nbAtDc& zigvk``3cj6>mk#>O+vrUita(n`~E+>ylB5l$_M!TT6tAvR(nT0y^2-FW=l=xBJ=s;)G=$v{a=^NBOVCz1z5Rw*Mxzk;b#SJsP2p zIn&s_zY4c5+!np3`x)rb_l$~-8Vu@MOOITYVdR=eN8vZdLXgSIY`#3D!aaFCpW^ow z8A)sOipK|Qyjksi8`7kEPm9E-sNEJWL$;qKmOpUa$7*-w+xM1vHfm2a)dU?VIUTld zl>1*30%c2GEZHxy+zUNTzh9~|Y_C?^%t{5D zpx2-d&8XuQuO9e@AGEEcGBOJF@hiURppv1Mj^-c}Gta(%4mUOLX}Y~QDfAHCXUh9^ zKI{dnRWD8+rayDk<}wp4Pk|v#6AcHP^$u>Ct=M^5^Pr$rp1#ojeQ4-g*jN~U-sM24 zId|k#FM8)Si(=bea^2U>+e3NGPbE=?!m1wBfmhbymWOyAF>j|+eg*ve-pp|C1cUuG z(>dA(fYi&!`+R$fT$V1}sC>bD^j;|s`1qTfS-H%?H$W5AqAtw5JtbVDIpV@U*Pa5{ z^jZJuc(>p{3it!Khbz#KLo4}Vz?Yh;I=kD5)OUjx3WCK_%+@^~DLm zbnlFKxjm>M*?LvWsNsRl@n2p7XZy)QGAk(QWNYL!UQ(-q?R4T zenRGLgh&X*X&6vKC;i zHh4hz&22(D)P!vQD)bDJLQ(YHW#{g0B?ck8{J3;!yH69?AYk1=m{{Opdv6fA4tOIM zYryS<>_X&`A^H6`MD=_pfJ|~GscTpVs7@5GvENds?3PqLM3*U5G{w=oeBfvcn$q^0 z=l?n#9#!_T*!Z$zxZCYz(F4sMbhM)3js^Gu6b^Q9PkhtRPuf4#fek^-(xoT>#(=`v z0kt@=Ex-`x_dmULKE~zSLlE>ala9<$>JA!lS{rLW+X>|vv|K-wz&X(wyOalg5rhwD zjtn{ZMqK+h7VyLaz~qI6Pl~rAD&9$eX}B96lZ2Tkj9yh`=K=MFb>}j_fP{oEh^F<_ z?98#bPVhIBZ2nn;b6~#r#ySeUn@kl4oL)Y$PU9|QUCu-O1?hzJvzr)l%10PVur)Rc zjwnQ(<{<-kCg~e_n2@{po`%8v{ z-FrdxPmaEZQy!k0*DF~RXTMFFkLJux7CP?URETRwg*4mXB6#oBMfs%W)f=4eA_kQU zJ!%ByeocW=7cw2XI5{^gC{0aZEsY$N%0>ek>tTZlg&q>2*S+E!Dk(kT1~;G|!mXKS zfX)}yZMXekacwue1gVXbugexAWjokbjc3JgmOkYENRkF?3*PEz)ZvgbJv3_FlB;0;mX(xUjEEV!F%AfZxp0|=!bIO z+apBto<_}&N)9rRZx9RL&9WPy^|fNYosCA)=ILDX7wt(mlS>TWdQHA~w>=a9WD|M# zgTrpy=Y7MhS`+H<+rZ%&TBQ|`zS!i-RdGOYc$-i5>F;k{bgQux2V*9t9hY#pRBz|K zIIl?{BN=b!^JoV}l6Nz%jUN&X;HYXsw+p_|t9c01?!P^ezai@uLrXjqA$C!uym?5x z$SYlkAg6BiBeYhfF-mXb@ev$vGnxE%{P|||Q9~B2J)sbzlKD4Z^C;8*TN2~be{02k zxBt(9?f=IwEEV+xK$_bbVE@gmB!P+=y4i*SoJFZAXmR$#}0xGUii{~ zXC;XsNLHY2f)j$vXDI6V0Hu($}=O@f%~7K^ zB8cNSr&TV$jr^$DAny^xqcKt<#v1MT1TX45bmORQobNv_ANC6@Il}A&yHn71064)K zh@5+EzU|Ri2FZ8
+HxOO)G!7-$IW-(B?bRyQn$WIM&H&nrPr%e&W50Z(3e24b( zPUwaf0Mrt+?n@AVNwGjH^`is)Gr4fyAwYDWY`U`Oe3EwIbS3Kg1GZMlgGv+0bWq;G z5AZKv!F6PxgOxncKM(f9^D?$AJN`J~h`<~+U8{g4T*-t`7nPO~58L;N*~ZJCUv}J| zuOIvT9H<^4j`=E8x{F5@{Lr}!vAuZFnit7viOOl!agXqgfVjwcKLFo5oej0~s z5NPCC1rWv_*Z8KPYcji^;_lWJ{Ps(yZ2C?g3*iuF_gynPmV{wbMXXzd(FA$RwO9B^ z`GD(4G~}+4uf%Xq0>zPJ{!Oz3%(_27w2!jIx(?Gh!5jwHz+ak4zW^o>!W9%&nSpCg zJZaIwi0LU%GISyewLU;-!~lBx_-b~XBjDw#TkHMV6G=&0xb)6R|NYEkf1MtvH1^PA z5DSCdBNElQqjQP~m$7c+YWd&5q z1FKbIoP!DvS1~VC{5He=Yv(|JBQ;yUvKy`yvV#vh)Mg%e#IlrwNW z6-98_e!_vldxu3Z^!jIMU$Q>ucqS zk3vF(rwCQap1cgbDiLK~YBiV*!S)~^2Bwg^^r^y(6mfo0wE2C(>E6yr^dKZvhz}E= z3ZJc=$`2$Xh3!+Kzak<)f9U{NHgP#>EPEsPED|a+bjNV*{ygc=cUz8|?mni!Vq(?4 z6}x1u3MbANqRh3QZ1-P-jFIIf3a3^8mh454`c)%cDi{ic9`Uf4%M+vL&!&i*{qIen zxY>j3iqLuaCKXnWYJwa+`V5I8_WThmBm@w@j}R$Ov`kw4<=P7)0gpVSYLB$(EX8#- zUe<7ngq+W(wgXqk_r>0`VskC$5qgf+cVmG-n2Ky|VpZkMz zTdyNqbo6cTNvJNR>6Qp-{rt#lQ+|8q0@3 zXU~$UT{MoNaxajiZoC`tti7NTaC7{8QbhjsEas@{+HP|CskM zJ1pef+~;mOZ1>SM_TzP=ngY%9&V4s8fd2gL`**yvG<~EfXfO8RL@i8obF;CTN=KK#IiIl@DYYZMQF_Onf9WP3lui(luz5DCejiRx>vRQ0*>T)$bz6t1;oMC=~ z52TLT255&4`ia2eEesrYLi`|9|*2M3q6fkQ!OJ#}s z(@zC#En^X!Xb)H&9p2PxW?RHcn9M!hMF8GsL>eFB{96aNjn? zO6aSfo~6Bz$~?_kKqA#0WL)1ws@+f`! z;sf(yT;j+#Evsg$m({6k@G)d9Gm6bc{p>+&?^F=#cW8$x%C&3PI7uXrfB11X@$j^M z?0Wk2>5d&cXh|flhmRjGr1qBcGF(woQj(FmKkepI-qE49lceOo)MH3DUhKlbOd>4> zS+nm`p(7p1{rguEsnp+KD+ym*%we1G$E}9#Cd6OA@ZO;#{`}wk&v&vj)xr2e$CaLn ziig-S1Zf`0IG0vF;bF62clPXA>@yHgJ>b@}lXwe$DVaiHY!2C)FO1-|_iDdKseLeoT z<()HNFiez;pl2(<6b4hLnukEw2LT15(Z|sL?%2KivYgy(`8C?tn@KlzXc&-4tIq1) zf8L>qsZL(;{jsQ;h$zxP=w5M`PUsB4N*;ch>Q_~d$tfrl5uY?}JGoh3=}jY}gxG9m zX6Ai+_nr(*V+o0g`B@S5Jnq)5wsYrbNwpMoT3RFPiwwhLOd=wmV{IDxcVnF-L{FaY zztZZ$k@!{KC{*5B#1hkx1(r=$64QoiO)IPD&C3 zqJRt**wP04Y@)1Te@#tAbrYTpfCPlQ2^t%yC%%D!`lUDGy}-8t+r=IQdoMf?l9H0( zaA5Pf@Pl*UrcIl$F2}XBv_g7`_vE6wD-#FO;LC=){s#gGI3)b7GX*pd;GZ*~h)(Jt zM5zbAha2EwYU(8YYeq(mGLn*#-@bhl1ZfLGmO+D*x2W~;RImZwx4a1-5+u&U0CnT! z5MH^+x>rz4j2L0+Sgf2qS+3cr{BJDa#k<2Rpr1!^{4WKg3ekWaXj5Zz_T31sG`(HJ zvxRg-2po`bad#(#fu|sO?nds~y*~W5A6fuxh zBAXs?Xjh;meX+0;s0?gn!g2n>!NS5KBH~7v?KFhIX~7*16E)yCX6FLgfvHB}lgXPS z(a!j;l+@H-Xa7f!R1FTISk>0nHZb^-XZSj32uC#|lKMko^rugn*k+8IuqSDJ8q)9&9rrHJ-aMIyHJwvE$I%;^ud9!_v;QF+ z{W*xZYxmrGu6~W>KL#X3W)fGX{w(d9=_Zm+{>m5aHPbRbgsEG2BVOJk+wic(D=Z7W zLh%32-b@kx{IMolrAEdQx6)qS2b?8TU}@aw85vw{-%O&m_4u=mI1K-*k*p#`f=l8E z$B|<=#-&G3HRP9W-9q{jYq{#)n7*9ukk9~w%vy`Q43&+ za`bBk^v=K<=YIsVbgOBaJ+?yE)VuFHjj9y?qXs$tu{yxRnVqzON79xzr`n??tJAUjK3Y`&@blF$3Bkao{dBPQa z#W)Kz$(Jv?|D2HDOsv-x8vF*xFYus&d0F@5*OgTw*+&(Vio=(joSc$^gN-Qj582=A zsCBSrlM<84%X+&@7ccs54x6^Sd}niZTRgG;rObToB9S`37}u#77^W0UhHM^?;oUHV z#M@-DR;_&uf5R1CPrO)M+81{`Jrb$bn7Ehm9O;l?;7dp5)E45;P~zeKZ~JF0L?!79 zKJ4lzsj0%lyL#3xWEtfGNn>MU!<|}&L&VCvEMCLLT9n|Ekm&hP|4eC(k9xBR#G!mqlM6~yR|Nf2qydxL2H>4tH zTR5Ug5E>Lzh_wZk6DiLs_Z_-hOw30L{CyIR7Zj@hbNs}KkC;^;BBHHaUsjE450k`e z*2|fCaPOpEj*dXi1<=;?6jH_v+3V1(VV|_+^6>E7zP+)OdHczeC#9vO5UT7tbhV6Zjr{ht{{3R{c850X?*gN;2mI_)yw6P03HsC_{BR9;>l;Zu{P zsjSS-%uI|LbYJ}GSe&krJAkV)5`0YTbOcm8P&m-OJ0;Zb@;w0epzw_^x2CH%Z^l_U z`x_5U>3ULflGU19x(^c*Uwi$hY!7qAN%{uU%Z&-=Nor6(}C4?_a>6_>WC`c^PSa3w10QA`o+IV*V(hR z$1|^zTcxL`(~{y2%oe=n?XX4%1&=oP>C>-@eE%`bA$Z~I|5VoOQO7lWW8G z25UPxG&eVgM|=_x6cSPFcM^$Ofr#&y<>luJ9Is!$E-86u=o!KT68dr6n}wm+)&#=} zo-DWyj0^nzcbTs2J$i1aJ_ykTF(!ZSW&J#IZ3l-!iT;oZ2E3-1lji;M>;z?=eRX-5 zMd$%H=jI`3qIWZ%5yL5CVlsl%OYfyTv0rX+Z2lWtQo!Z@r@6s(C}XT zh7p4P1H3U3Nm`DU$n0+o{<&jL{Lp`Gi2Bd*kL)U<Y5sA#!;< zaq7lbuU{h#t+FZq{q|s{TKYG6OTmC0d-v_*L{^Oho|Q@Z`y1^4Lt3mSgksGSI6RC0uOksZ6YPoraT>u7{pF&mor5EMjh7J*hsh) zo<4nAT)e-`D;3B5?C!t+bShVED{{Tx|0Rj-PV@~`?CkPmv(KMD&#hBxKtr=>)9s_j z6iGKOg#-u37}FG@=tOSA)Mco31@Ywpq6(&v-{a|Q-+RqMpTfXiNp+NiN-qWW)5t&Dl zy*}=h&K+5*Hpu@zYlGE#!T&(KJ}31xk!>*j@`i~Q3NPGa<-nDo7VXM)+M>G#)GK=) zO`dX?F``qG^ZZJv+lW?%eLi^hA`=}c%84=ycG}tz_WHok0F4hIl7K8(YSKi%rsnN5Cm@XCsV6 zH`NM6hdd>Jwl-5^V+LV6DJ?k(#}uNP7jCkP(v?*bSp(3UDuCzQMwixDS7?*eUSK?q8J zfy6AHLLVaxq851vUW&)hKH-2XMOkiH%`y%4T|K6u_El+BY@uwaz(fbK)OO)k8-<3q zD#SWsv(iH=G<8kt&mKgUkf43FajZO5F1-2m`}Y^^%PIhsqbRi$ZsyA7hTT1vK8wG( z-lZwPw(-US7cX7niNr*sI4&E+%th^DwrA#&M|pUZ3k>{R%YS7gndmYQXR}oCF6FFK zE`oc83=iC@fI)WmfyEbs>{2`CvyXjM%5yZG)5-cR0zCzKR$uMphDNH_vM0sRnNZ#@ zOs2D0P&^fQkGYqD0>|=|{LOW);;zCgm`_HAgrvN%g%m?AKW}|;(k`d_gDN`kWbT}d z5?*3lT+M}A=w1LqR&$-zE`}cQ!ufGjd;(UVlxlJS9pLt$^%H8M80*l670L$#q_Gl6 z0N_~5`K~9KlR)mrbOztwlEdA^~!0LhZjK?@|W1b@rHN zWkurs&@eD$la07_9+EtNULmAC+v1w*_+b=g(>M>*AEFwX8O*J+Z!Vz8Eil@8?*&u% zEH}TBl{JI{(7MuZ;wYzt`&u%Q#%Mn-Xt4A|kX5rjwP@X%;3Lx(GBVl5CWQV$Kv*>K zN}6JT@PJ^O-3>`K8lK;8#(E$4Uhrf742@uUjaS$j9JPz1t*^1mk(PEQbF;9#t+afB zBwu!Y_$TfY%bxsXTDA4Ir5;=<=p+hhANJ-txh}GMe;cIJCWRNi4L{g&GS@(To^=9C) zJzZ(i=1WS=4@oR!Pfn28RoSt;NSN=byrR5X>Y*^=Fwtd=ZXJNQ&X>5-5`iN^kK0@A zDcVm-(wWWa3~bm&(z%wi7TDliy~7bQK^%u}4|!C}78V(RH;@`A$@DGi{~HTvb?3N- z@C4H1Pmg^W8K(bAv9zSQ!{{n4sqpQ!0(YlT@#E`%b6KN=VEOF%^EE35dU`itWSOIB zS4)1PpvRDiB3LhQLc$&o(g8>)_Qf}I7%?IPF`uwF#w`Yftj8LO@F1=Spbi++aym4+ z7tNVQh-bR9uXbZIK+8nstO_|wUn3J=qEEkC(e)lvE7A!<`3QPjLGHNa&t0&iXaVeGPzw&ef^>z1wzc5;t zfF94p=b{PpI@>@%1Q`UAY)9Z~?fa`m0YLa8Kfks+_G>7j2VvR zbzw}N{{lpVisn>JKSOC1+

4?Vsd!LdB$AVA07QeY<`Yk zh*~=?_o<&y)o1Eem8Tr#_l#R&i}r(Y3xO7R=luSp1d{_AMofUcnyoFMp2Kg`qNKjB zy!;ZPno-f0B0aXt$)^HL3fHk1CZb3rJV)an$Ukf>n@kdA|K9I0+ob(CAfVGf|7{B% zrNZ;kLlgf(?D@4j^T%I^9OV3{Wg4s7JfELKzrz;Apr5^4r_Xt0g;(EMD;^wz| zA)QB#AFs2r_)>ejymV@?W)K(6#grXBZ{A2X-C-LZdzJYt0*7|hDPS8P#hJ$5Dx^Xg z^n`?X5C)L;?a0-LQrXpFfAf6AWxX=HWLk$z@rBOov!G}yfyR+*9C^PV0Oe>;L262x zfK=C!6=*fOp1hajNx5l_o-RC^ac3b}!W@k8PSp>umIn(O%Jfbs;$EVjEb;(jy|nan zy9pvKcuLAcO_n`ZKwP9qM-Q=ql>~zfh8@B`PJByYth78!%`g{PYf) zhHhTfb?#3V$qr{zO4k?Qv3Hi{7T{}u9I0%)Fv4~Q3n&Bo2fy;DjAtnYYiMf=jqN6p z{*2L6ngBU$FO_oSqu}QM(~{==#+QnoV}i)G_aD=4Z{A@^?TF zMbUJS<{inmzd(Q0GD;j8DhfA#!>o(Lt2-&EucD_l)n8E}Q|tE=iH54V5khFl_sJRU zK};M`KiKFA+&BJd;gE2JLwx|i6K0dYuS5%j3uj2_KF1BEhq5r$cq)T`7?ICwwy8M9Mk-ue#P?Lfd(X zd2a!g`@gQIW9NiVY!0J78$CF&A zWi=F?VF!6D#m=4WdZ_T8L_|b%i!k5>`UVeOG$4HF}4GaTXzBO@d8@$oS>HZ~tb zL$`F5YafZk?WcmLvYnDr!`2!3FbYkMX<~JKvd@ZSt^UI|rRlD|NQB^>q+72HQx11_ zr?unS_otnq&m_K2qjtzy^U9g^xALli4t7}5OX8208RpN7-(9WQ53uJ8d4v7~;~&Cd z_|2^!@K8>*ccx?gx&E#{HzPyH(-H%db-Remcz#3aiy!LiiRXWf*g02cIt+Ve4mAdC zVEy`@goF%xpGqV)P2ImXO-j$ZzwLwn$#%13IB?+oHi|8zzJxRHKXNnu{V`&%yaV0$j^dpAsRe#|Igqf3QEb8SA?n;6`Mgh`8R3OuC^ z{HG5qeokjxfxzRHm6ahSMK@VGfNKm_{;c*A+=NzVV>jls(=*qgAg#t74_NqfxtLyb z2QG$8^~{uLuw`Ge%ai3Z&krN(@dgtxpJ-du8gtC~^}s3G0iA=SloS9kyho3oYaK#v zlb1tH_6j*zv1NBN`#mIqv^+wXZw=s*S{Xdv!LfMb;}ivzXn$p%qOmxk{;6Zj z_?NqPcY5C6Px5>CswxKKHOj%C9RC`0RqeQHs};~fED+imN)Mb-(+{Afhf@#o+9V+& zhGj(qpSUa9?kEwFde5T2NA&2Xy%vGWrnMI)bfukpb^(AAeED+2O5j2b_}8M%24F`0 zZM06j@S0BG)WJ0L!eQ4pGsjq`oY5rq^V#_GPJ+KD9GKRbKY`w>iL++7)cet8Br=%> zFruwcIsF7x8QqPVwJF=r+90kUK;e*#x=lTb)_`b-q6npo<7G*1P4^n`?qg}>1gZ)1qI#VyaM6zISqDQgSDDRo(C3gFL~?h1~?ukcvV@M zwN{jzX&jg2%SSs|0;^yt&a63E*_+rSSRj3l7g|J1VL$-r0m1s6N=OGHBU|u+4Li({(;H zYg4W8f+o15Dx;!R*a<@U(Z?{vdZ?X|m1;cifI6`A&bT`;n+f0|jGl9?I!-@&2aS7$ zGrCm|W`|37uOrLu4Amm}m3Yi;zQ46G;+wLDk?zfMBMyf@?fFR4=nh{oWbgCS0;Dph z%KglBI>GHHsIM7vry+euix_FXE~-99!SNigmEJ)?i;j!dg(*OO{m#2Z6zX|~ElbXe zx3;#Do21-!+g?XZpwGDNRsEigV7Kf%#wzaCV<UEVDFm2zrJ5C8xW85L=78FA-U-BRdQ*OkcTxEz`fplzu~lg}aIegV1G*DH-SmA_>G z^cmy|566=^ZdqO$&&T*5tS`92Xj_;t(Y~V(M>4Y_tc(&qTWAqaq?ECXZk1u#C7Q<(;P3IT1fF z0bZy^VCK&S7!0;$Wt4Nxw^vUQG^GbA3~Xn|mCYoTXG6;qcb^TE?p3b3Jo||K#Q5#e z+Zz}{Onb|+1GX}06TumDNVD5>=?$S@I(yz??X^BQpoo`_#x3(l#l>~^Q|KlVp8haX z4|wG`wT#7Ze9lEpPhlBi0qz>Tyyc$Ua|J9(YB}Ay2m>Of?w#U6P7aw|yyb;>pz8Zo ztPPRnM}+V`NZOfLZxlh+o*l^vw$2{XC8=|2x(?=@NA%`t=hKOm?ygx^ul8`I@h-d| zy_=%XaY@C!$#~pv5rvs$D*elej5LlsqlcachHKRr+B|6v&)rbF*h)4yDYq@FD)L!S zJO!|0>28$?_7H20WV!Myu(;NGsq*nI3k4q^U(=%%*q5=s(u!dUi#hb%PCy|2vAlTs zxUxxg+pqDY$@$OGvYNJ-nvSI|i#tORe1zp?0J#&ohh4t-e&%h*lEfDtG3QT(=;%x0^RE`hrz6W|74)0;PX) zwWhr~Aa8($h1^Z{Y{>wKQ5D;DdI^cv?H4&zkJzOq;(RwPy7ekIUs z6D!kCkFk-ygN+_P+;Rm7Xg)a$_W(PuW{j555Qk4P{4eCi@0m>U4`0cR<0+K2z}SIe zdZ>qWJ?}2>Q203-VGoiE@KKGc=$KC@dwoMf(Ip@i{eUYF@PtFL#unY&7_VjXwT>Uj zcAs&HID8Pe{j%k9r@lo2y}ATMH}xj;eH>xtRl=_@%H#68l$(`WiLv9{G{>kn??0&- z=bOx*kQqLEs_`~k!06`n#hMLW&#&MI(^NU=4=DyaQy+`6rB}AK7W5KmNX}GImzC(G z8%yrE6GOMA=29@7c`IA^bi0g1Q)aH|*fUW9&hOv9cZ!bE2)SOm#xPr~PUgPGKN~iF zy+~z?ZdEQE<>sJ5Q2ow>)wQE|v@Z}g>y7 z9@%z>oMky(`?()`xjJW4Zw*oiHO$)$NELX~?IX9JXP!gv;G%UdM`VuQ@SPO)@cNVN zMRj#`uDZV+INLU%*E*^l5f(POT1=^+9qw#t8HmHl$1yb#lOm|DsVWgVB^1XR8@4k3 zH7(pzf9m2U8AYGi>ERmLA8qPmx6^jEC}6Uk@^PhBwN00LKD{~Bp!&Tntad<+%bOz1 zTH{;iz`8o~77BV`XDP;0%DL5g`u5M4{=!2C1wRSaYbv|9XG#5}mp z724c7Q|a>zWSkR@9OaT3C-O}P8~LBOZEJE0Kl6>erP|?K7n!VQ6#22TMnk%sLrI~p zC`>YTSk%TZ6+>rr;h|P1XEN{5p7|T}rhPYFQ$(EYl(#^GtMd<(lfUf?#vS=xvz9cP z{+ZBFn24BX=Urd==xx(i`l4TjAz3#H4s#d^gsY$P_zMC|N8+<7Te#zT@=ysxO3)5!aHoSNrRZO=*Q01yTb61mkV!Q)?Y-gdb4Yih^Ae!D^LNd4H#zjY5G&9qXCHrNCPLYuI@*%I zyfwSK)(sw!^oxW7&TmsQr)5ZeWqfUqp8Y+iE$!S+qb)BhSDF>sg9!i`H$AzZ$%-IN zms+M)u~m4zL-XsrP#P+A@2f=?!H0tw4)}2$?~N*F7cUvVLd=sP(o46>lp$cN0%M{E zwWQ>Tn$;n0J;P6-8YhgGUOF(9^gBhm@z37(sA5Wao7}yO0zmD8Np-A@!|>=Ebhx5K zL=|XO3Ow=9Qk#`OsGdk2t{ZguPBCj90-!>V>_tO-Lt>8FI0NPrB-l%2F^*sXGHqKbs$5;?AEz35L!}CNZ;M+=t4XENd*nNlY!=Su3Eh|2g2Y zNXm^rc>M6I-W2eqx6u6B`k7*o7?uTCSWLdY@(CGXo^goopr`P4tW#3N>_M08Rx()u z9{0$I-gY;IO>RYt<%iI5QCHusdn}q`O0aro$jI!D;3BkdjJ?TxJu^kJF<1!h<(HTp z3);@?YCpRq6)MHh2-@~9#aD&?UZB@cj=RBY987^ z61%%j#=gV08)_TSQ*tLe8og`$J#_+Vla-R1wWgWtry6EcJjFN0Bjv9pPk67-#-cA( z$A3Pe;pL}>>NVZV(cfWQ5f`J@p0&R5Qt+0gUh#g~fRIjZvw<5Z5>3U_J{^GFY?*qf zz!2l?xqPxw^o#3{1d1<&r3g4?cllVNcOB8GvfpI+r0x*^;&tz>wm-L(p|ja!x9rt# z{F1ySD0^>GY_Sut+uqT|JSgXYjd z52ge*Q&;5&7O5wjevN75nj1v6CF)(Q&`@ntrLn=V2RE9|-te)hsN$#-ZeNBUZ8}M+ z^XKx6VBk-mM7GjbX#Sk#S5z#u)j+_o$AhqM;nNp(^ma22F5Dm94$XT4dq`W{Pze&P z^!7GKMRTd(JrqBqOEN+d`g`_Y;;3kdJFNS)s!E|M1z3Ky=t=!`XLXD8%Wu5yR&C6* z2D7q7qsOo$Icnbh{Mve6PfJfHgzCx0O~*3cr|(JTJEvY`B^2iFJ<=(x*Zgz~>lIuG z9|*c-H@PYIaBn(ITb-0=$g4Y5#%mc2!tSgE`rf=z9zP3DC>U(Upl)~FELfN@GvN71Psr2n=o2hrmBnV7O7Gk3;|&qiCZ>n_aaiP`dd&Nc~gO8i-H7Iu7`#7inu&AznXPtM80 z$ARbfc!5XFqsGbZ*t?aRZGV?a(iCh~gTWt#S=96hn_7MSFmtPp_ZJeWA;3@TX$so_ zWYxc+8P@pCx47fnvM7mFQ^UM2izF=Zdx>1>iXdYcD40(Z*#`O#hgx0o-bVV#K1aDh zpkAaOkN*?H$qzNn?%nY34Vhj(XjA$!HaEEQzolau4!TS#!FsoLZ18ZoraU1USRIm z$kTEE`H^5wKW47D?QF6y<#p?TI*e6#&5gE#tye(jW43F@PILPYL5hy`b_3MHD@0)Z82WW})-+?$# zpJLbNPE5?`tkX5s&ejMlLGt4>d*h>yit7_7!f_h1 z#p(yE{6b*x&T853wK4^P!uR6`C;SfH)AScz~lT z*M|P{B^K558*0%t<%D^a7z9$jRLoHvy3;WPa4|UD#gT+n1RpzA$W{*>ijyR1sjpiL z($doOH1XXaT)X4)MKPDPHbuR#_j+1pxzzUlZ&RoWz69{mc z(p$1$?ytpVjnOo%NIri2_<=F2tyVfhL~E%0=cl_3z!2WOd2`(u1jMyPw8e_V99Lmo zXW5<6Hd_5rY}SAJ9FQT^S#Q9?VAgS0{*nao(4n#X8XVpPZkB?&56C5*F+;zmn*QCp zcQtTBGB-EZyNXpU@tjWPoOm1`ACGp)I;_Jx%I@KWdK<{et z%!*wfEcD$tL8)FuQ_1{Sd@wGHw; z_r#3*Vapq!!Tf>fgdb={?q{5v`1C%s(#L|^f+B6)Kobp;ul$+J)P_z!|8304pkp{& ze1Vws+Qc()z(EW>nsGr zK*Si&1+50Y2ihRxf$<2)mGD(Nk<=Xl(c?H-+vInjo%13?8 zlE~KFsZoKM^3e61g*jjS3A8R+MuVZqa#ARXCntkT^z(UN)q^_ zOno#;zEyA!3KH#BnE~$&&)Ka3ICkR>GvmXn!X_&&QTl8@-9N89E+tXlf3>!a_RyiM zP36O%_*0-bx$PQSkh9t3j)De-?!Hc$@7lUtjsj{m$qhNnB;C`S#reI>tFJ)aczl1! zc+|hK0CEz&P6Q_h=O}&r#0;>+Mq|m|GKeujI7n`2=Pct~8y96|Rrm^2qH_+v zf*{qArYqNE^cRtLVJbidT^zdc(rbl}FIdm{Hdf6^!pFrHn5~;}kwOzirfRnv#r6)% zar3z<4Jg{ER|~boZG{k;v5dc^<(zo(;RwCW zYZ-@?O`i_AvK3iLa7DXA5*sWc$B;c80d)^+5~McLuc@b6`!o{wG5kQ%ZW*I45!f>k z)RlMM)TjKabRX7i%%fOrnoaiNMWq&XtFC_@lwazXqa;AT?lQvx#c%VJ_0FQzi`rFg zMUNF0J9P`=xcU}dJT?qX-0#&2oM*V5^C048dtLo1jY4t1A6GDa$_6SC!7W`@-P8j$ zEV;WI>T}%32dY-UNX%ABI$W>Q=OLPCRWjKrt~;aWJ`ja36}r7`h1VSjA{c)cVyYPK6j6!UpEjrL~$`PonspqB-L*#sRX;AVY$Ugq!aXup1_ zVAe(c;9&cJ^Bc0bCpr(~gSqxabhdXyb7``l){JRLNZBQz*O4s6A#j0Em1cLVJz=@m zpn_iF-ORg>1^dhC@jD$oqQ%(s}Ed7rDA{{49Bve@x^(5k)S@0x47 z6lg`MxGlv?c~Kj~rZ3xCwcNz0)U)m8;|HR=X`&lK9ABz$+7-UvVP}xpy8K=EL!Uq2_Rft9{vUV~2#f z%oS@0_2L9%JIeQ1zE%F1v3!Cva=#?yh4=<5sfG-Di<1dsA{5I7RiV4(cU+;moLh8% zXU~~E43QZ(BDAmQhjmMEckyL;)kR&b=AX~J_6G{=!YKDr^_z;SJGAp+&yLUs9Gvr2 zt83elI3TAwshxW=W9;gbuHu~`p$P#49JI#0_lZ4C_0lRj@v z;q9oa6f}&2nZcud;lXu5D-Qs7TgTN9m&c34iL|N5`_bP0fDC;k?PmR%A z%dngV38T}n?(N%Q{_Zh3#b~_<6!f-zrOM%5^*O5Tv+17Z4I+mf<+T%k_GVuF&B<=k zEAw>~Wwx@|m5(+<$&vPLnsR~4?-E_b>vTCPT7#Uv3VhqFbTLQj6d$0e#PZTefpV9g zZYjPo;}U*re&)1>CtVZl;ZuT++}15KI(kj?x`_h3T^&J7orm9R#|Wuu3ae@p#$G8t zm_&ZGQ^>)5!NOo`QTcszp6pNsHm5g1qsjF$!k(8>f9*lKZrGM)woSwNxjLggX*y;3 zT3ys}KZC8+WhVqS=TW}R;#W-UiZnkZJ{w^QUy|1Z_u*6e7*7!~re^J}+{%t_fzJc^ zDHjE}S_O3b^TsGWrjsWss+>8tWs4J2Kl*j4-q%0s)^q1fb-t(_^D3^}KBxPY*!5h> z+EZInqPo&vl%E&ppUrHho$sMH*V^!8;eHTv-2Tb%*xqQ$=2W@YJDbXyBZU-C450L# z43`^`U>%(L;gH*Byh7P-i7jBQ{HD#(7Q-O_svXe~p%FGltL1?bk-%ZfFVnHq+dUNF9xSf%a zda<{|NCmgVqD?uaxTd{N$J?fY5m8T**(oW9^-*aEVeLVHgabj^mmj_9J2QIb+3dH^+%KlhWDPql{}nWz6%IcLqU9GUR1LA;lb%lB zpP}Uu?AaKbpcS*YXeoJ-=I~0ux?YF;bJvKr0tw5Du4nmF%Gu$kptJwN#)W{4fzfI} zbnWWkn#-b8e=wPr*pzX#(?5tyT{zec4`E$pDp=h?EaFic?VM_;pg60zy!e1s$&>c# zsH+Rfx*mV@GmMBbqyE2WDr(>{CtExAkDych&V}p@eps zkAWtMj^*QwkNEUm71rE7jOn>0+bs>;vDz@)+Uf=4k(BIitZC1*wo-pzQ_<_GHvnWhaGd#^v3=9&1>pSfUhx$XzH@J1&odDk8Sg#o)7f5%tH+0K$TEgSFRr*}-5}E=x zMq7^RUWb1S#5Tt!Y7$kfFGRlt0FyyYt^_cGB0o+B|fiahv3I zIn}VwD@1)I-llkF=}nf}yV)Pl6*IUYjw~K>6|GKLlrQC-ChNR)4RTz!&yE+}bor0G zclz!|1-#v9rHZS1aT&qCg90`hG#Jz(HcY}jatF+nD^>;#^{w9h0}VL9TMq_ZfK%+n460dBjbzX-^UU;bToHV zRTH6ss=c+z?RLi`GV1<(V>$HkFoeu_NGHzzaG+N6m|n#EscC%y#`Ui&ij5PDDFuaL z?$rPIF&XdG^_Fk?fhN?jwXGgv0G@6iV(;B4YA8OSk9|uCN`}}&kebAPc_VKiwK%u-dQe1IXjgz zJRP?4ZnL2{e??-a(n-dZwo}rSZ#KM7hhEy2zVGcB9|xHw;ju3O6NaW4S(MV?VX(H; z9ooISVWzT4esN~%TvWfGd7`(1)2-jS&$Cz8EP036-1pcTcN@Ao(tUpVSgLz~C{7&3 z&|rSt+|1g`Tv&0U_xSWsL&?sD;;U@OaBIICm{$wy5Of5UQXhA{e8y+d+Osk=I9lSacrN65-imXm)tjlVPoSr zK5X(_m+|VUc5HGzH}_nkd!o|hS(M;Uk0cZ{ys_^LR>1QV5D1?3Tzc*NUC%orc=Br} zOQ`dX;a7`#3)TEppTW-?QHZ$Mbu76#{9Nw>t<`uxI|#JVH{=tws@+kQM6qPq;bVfXs$<^~?v8_k}v z*44J>o*4;ddB`94ajdFd_yy(bGp8h`(_Tp;q3%r0x>~}O`%(U@dSecq+oYRr2)x??H5eLfBMI07_`TA9GK5VoSg6Nj=x|<`GG^mHT@1 zDs;E3e$%`lDQWTxgABypFss2u$c4)60i1}V1dn<8lu$(0J;RJ7wTEdSZ4zV@AbCZ4 z(#5E9K;?p6kOsVyzQYZI-=OvuzEDS?!p{Q*2oOz;4D>yI&;8Wk&>f*cOpk`zq71cL z`4}LK^>-5M@hg~}2J=jPL1xFgqy0RboC?+$@2m%l^SHPpvkJcf3c=zpaJOP_$J4YQ zGdQQiyMFo!UD*5ZBCzsw`1k>%ndYD9zAW#w>P5 zjBeCBIt~_;h7;x~bO4h{-Bkd36v9pbqeL}u3<$;Qr7hwrRu@YC8tj1|75um(4&Tt_ zUE^s6u!z99zbQe4sNpVf`o@=9FseJ`cgQ~d=&|>ft_08YDQ|e1{L;Pyb%SLB6W)xDY+6lx0JL z|B9FWEL$s#xM=ET1+f@W7d}?J8TXki^yGW1otT*K6!P?sw)bggDwba2y**W+Q#n_l zDESe}72WWa%d6Wk)y+S2RSR@ki05s*qz+x#$iQ$1ZCm#d49V2tUBH3~|2xo0xqERn z1jA8)`kcxbf(c%8W)2K5Xjo}Q$@#!s2#1xD@Lr92^)ST;3RzjCKcOg$IsZ3+Z3ZiX zhnXj>hCk#^xIl(wZUpC4n#bGeC-lv5U?N-g%GO7!*U8(RjGw~;d4oe8W_X_@mWMnH z=d46C#sZT1>7_Qr><)2p2)JziGgmbJa-O*LDN=|9V1O17nNsPm=l{AE`>ZRH(SD+S zI!4{smcXH@uOG_6{VnNanskZ$Z*uJWlbhBK?As;CRWb|l7UhdmmgCSQvgiaYCpX@J zVeR`a$*z#Kz&3WKfa7C-EPk|hK)+&DW5+B1QLn#Za8JH9?%nQWMGXo~Y|@15Mun>woqF!a zb5}IrFF3gAA7XfU$J58)bxVO;uDVYXmu0V(S$z$?xocCdkG4xIwgAA4*CG*vQNgQ8 z^a!on3ub{Y%~KERbfQN!6}UQ9Og>!9aDSyH@whHy-;b{=RkWP6^1avpm(XDTbE&da zUB%je9^K(uw7Nm0fBqZ|i-oV`a98Lzv zN2T}|Uy)XayazGs335f$nnCd?1GF366gvbPS_C0u|uMU1Lx&!OFDeT`U z3Tx}?{(Su!Fg-A37Bxf+6--+TSA)fAhMtWQqi9U-%a<=9Q0BsT4h_+Z7cZifycd?c zQsozSw5JjKg?8?gZEZ&;H*o-nKmTZ^m7EeLrEGS`(BADq^JFL`pG(~S*wV+Lp);6r z0-L8-hS$&hV_wcBxN#zVT#;V2IG}Ba#i`?nsOC;6w8f8Y{*$2lxQfw?w>O6VlH6#e z))v{&{~@K-oW`oss1gIhauu<@cZAixA8LDeh0$7!_2~(Oj>%gThu}x&-xL~#mfN;( zKWN_*(GhJM^jZo2q1xKoIY>>~v9m={m_!O#YoyiG)oU@^hSZWQm`S)7&4dXh@!QvG zLTNOvu{8l=$)B|4QOmO-xfCY5M7KTGc8S@NtSzVR=Y9Kjjlek;eDU;^3VNzNq>nz7 zH+sU3zi_&3m~?&bMu*f#;MIq6j2cb&Q(BZ8xGN>0X8TF|FphN?FbK}j;=*g4r}@Wv z%H4*af4N7Lk_RqIPZ7@Qt~qlfMI+pE^ROLO2=oN{m)bK}Kn3K?&Q5-9NkaPYl8uQ( z6ffdqzop5hK1Tzki%i3}1=626emhGxyDEZc_JBX2IM=G_7>xjR+%uo7$&ASPw;b z5uUTuSrP_(HXqb_`(%(S$K2<1;OP+C0E>&x2f7L%H_?G7D!+w%w$O7uPOxB3fpWfx z6YuV~J2poK9%5dab>WkA=O+JmxaY&~-;d}S!I6-4ia&Dh+&TD<*?f@YvOg*XR10w& z1vg}{$6qProH$Ia2ktS4S%)GQR6*HAdmO_Uv&GDXo=;SDF!gClj~4Fm4-QW74t@P_ zg)G>||Op{=90SyFYu>q~B`Qy*_-!qXF{I-~Bk8b8WC;mz% zL#L9eu?R_4PiljD z8CvjA!E;Y664I(3h{bfAJz;|p3&VXQsOpa$KdzB_`_`@Jh*TKmICTp7Ck!zCY?vyR zolRhfy5E(xIYv06HZeaLQUr<;*%#RSG58T|E}YPbQVYngQcSM8@^4jM(^jZ(;-5IR zUeUfZXl_95lsez&_cSibo2fdh8h9N(bei2XiEsUEQ_j%gt-7jBsaFWPf9asy4(AX;Mn z_@UAO4`;p5x%)O52vdLKOO24NPq^AvET8evAl6g;w9~Qx2kSzJ6B_D@s~iZzH*ekq zGC^05T$gYna5`~ry0d5z5fOO9!4SG}_wHSgOHW!8B-(v0X445%ms7BPwu>Kgg+pX(!Q(B>)MW7#5%zl4P96DyFM z!SfV%`SIY|U!r^|0Wf*UHv+kZa>r-G{#c~!(6&z`Ut?rqTCq`N4~1d@H^7AAqV)`0 z^!^dh0gldbKd+hI?3A-N3HIJ36~FgM_}T(~;dSmj69kS-$JNpV6M1$UlO^PN{%hT) z|LnPuGox+5&%(;OKsGzDEc?US_3N<+?oDQ9=KJ^Wu?FkzM$sXtfk!Pg0A|7Y2N?nD zHian^MJqriv>`D9s)~hSF7f=MjY*c8tKH5i=zSi00x#FzqFH0uZ)neBo4J;(k&#}6oG~4g*jXRk9WU=m6Z%_ zIkYlVL$3>WurFmk8N`?M&|My|p^n-QD@`Bo4k4)$_A}_NSh0ePly@ryb}Tmqx&%ZY zQyqZ;D7Kl+5gD+HZVb@~65>^W!eBDCxBbU!$+ZmnYbve9jM>N5vwp1%Mc z1psbD!iM_BSV#WmzX%ifpC^X+o(p!KzX!{Wq#%t!f^@kzCIBArF-!0SKuCc9U7v1> zm^W_F%e!UEOW28a@bZ?~k_+d`%zH?^2V1fhFP?9AQ2~7*2Q@eH0r&|8q2vdh8IkhV z?b`yZ097D#tXaF3h}4E#pH*3E9653Xir-&+xbkudyRN$nfVinhW7B;B@Jzu2g!6$M zvjzP?kTP(~X1OnrlB%Tx-$KHQyl!A)xuBpR+gFg17Gl7K9D>|X57cSQw1lO6SZ|FA zwnhQ-P_&=0*9Z}7Wc#jNawci0fc#*!LF!cCkcyNN{a#sFS=P{5_{j&R2BQ@gho+h&RN$jO?5{`#V95W6?P2l)Pz%OekYplnc(tJ%UcB*9 z$VaH08cc-i4u?5TB8b)mg8{sujJ*7`wbq$M!r8zjbx7i&-j7eL|7$Pc$8f^$rr1y# zUNQDfmrxNA2oQPnNYR27CdFtmWx66Kf)Q>`kd+}tB+nNVfq4UVVAa;FiF(+Cbq|_L z%d(H4^M~nzA;KPzOAsWXxseE1DONzFkAWYbPMS%Vi3L3-Ej-B=9Z3~>DdEiGuV56Vrb!Ei*-ammL( znyw57r9ZdYYdrF=sPG=QlEmn~8B|E{eWIy_kgbo$`?t)%2(vF_2HfXhpH9S^URJ*F zo{L`^0Iv_iVdO7PFjqn#TvmuA5m8C``0@Vs!a~JUU9reF;K#dhE{2EPm)*u>m&6b_n(bm@x9~yN5bL%MpXWPz7f)m%Os&) z!7qXX9Nv^pM_Uajgq^!~VT9I)ZWQs#e?d+-fH3^a+_|kX%NR7vte&2p46`cnfN?tv zV>WEqP-+0?Hqh;4{cFn)YsRi zrt;_JH=&gdL45YiVKr!%^0f3%3V5k{-(*XuWLnwgp=$dKJ9K&ZeGiXnQ@+qa??_2+O45ftdXds^6m z_E;IUDa$EN6f_)YAKCCDJ{(h0%4E>zrO;beU3Oz+$*%rGYpcSEEM{prSI>o%QsA@$u}=l(9K(laz%(Ll z{n=n0k!Xa(Al|FaLa$I%6FsO?i9mN@Qk0z7vWv#=JZRNol_SZz0CiCBY(a1cob zz@-495g*wffBcc;*@ZqsPns$SMK9gZNS{UL8s&{DhDu^~U3q{4uI7AU4q!EkY26Dc z=(J8C`jMW{sZ49?#*8q`F!RvU_+?waRLjBYD0E;23rix7H@HK5RjIAg0*z2xBcxsG zt=+kQza5ZtkdXk|>DAauXlBL1;<@Ovn-! z!6(kJzCvz8Gg!K0$w}^Fw%&~NbmDbMig_YlM_{3$rXZW`FV;RA}kS~6dkeOq0OQunE|XQOjIv_Y4Pd^ZbZn8pBZHR*tTYl?;F z(GwUYR;P2|8Vpos-E=^8-810NT=-hqN`xZs2L{IL@E1uKT&AjLq^uw{#)r69H=S~u zaH6PD@R59j!Q9dsg-jW_6;%db=NQ(NU^rFpX`S`9Aox7-ierQ{n03A`H5aH`oP&C{ ztQ+q*&U(O0&+&&2cd<|kpHaaKzJd@*dg|UkLncqrIXUlgNvx^=D_PIyng7_A>oDiK zb!II|xR2pnPRMiOYsHC(+UkJTBoQ_lA%jc9RvVV0kggl{7JNGCZ_cJ7!{-9|3r8HZ z8TBiL{VBPcal-oVoXiRnltS%tx;usfEB+Jq-}r(wbnkH^i6lQ8L)bv z<=-f9`i+hIXb3bmK=JCYX-w9OyLh|Ll*8~Go-AGbppU-;Xi`z6j_c5;2Om9mF22O} z1=VVgO4}ERYd7Qe+!kcloU`30S-3JB%RrbUgZZ^~*v)w{c@qLD;SOv!mYqbQMiUai z_`mGpC8uN?%mWixy9!XcRv&8_2{jP!l=J0GUY$X{66+mgB?r$GFNd~HvpjEZ5o?jE zt7~3u4O$M=opt`H7>hE016wO~lbJS7vqSG4RI9_fSQtlpSeWC2QVESO8l!)qy*t*; znBt|YH0dr7(jG@O1#!D_%>LQF>F`DC7<q|L)V=_V)}^cIMd%6XgKiF6zs={MtS|^TyiH(Z+Oq`$+#h~+ z+pZ}5_Eq7kK0Q<+>784kYt8iE^V-@#HX|||v}aI!7rWc-XXxO#l?jb4r7-o z)Fzc~!r(&hPCb!>Cd4^Qe5{;Uj_RJJdbfAeIm{F>u7!!`ARHe^%fm>+zJ=9(ydMMF z#$iC!9_v~6e$3B3p=;0=W=5+lf!C%tNrSKvo6AtkSGpnLA%#AW9zkdqlHyHcRu&{; zYBf{>NIt#wfw_5J{NvYBiG)9KV zluOFXLy@~+VFN_Wk>eB1o%(c-7ytUJNy+^L%C`kx!2jhnMSx~@K5uW}j-~xbMd5%e zh-Mpe@!E}q;OYit4%c&RiH!7d#fua#bD=0AHxO>6p|U!P9b}4UD@}H{scjgBX0Fi| zt29jggm2)3Fb|gLK8a97y4E!bkEy;5&F7pd#is$`T$d%$W_n zjHYyD4F<{sxU^m=Z8-M@^`@rf9WjGRd!9O{^&2*r?{y^X5U>FVd2<}Ilj_d0>jD+L zv{{T^Gn1ImTQ$`{5m;Yjg?2t-wydBs_l{k=EY>OssH`*euSdz!@yt>p(EQt}k3$Y) zJ!%cy=%&p;iZqH} zVcF!Z=}gT;j&MT;d^#UX5Mm3_|MPJJ=~`$aS(_fNxVCXV9Dq7;Cosz+jp?0V-uk{R zDhd(-59bNCQ&MaU%d&Ag&K+yR{eTf?UlV%UfNH;&{vMmNHB1*qNr4SXm{$YMkG$a- zSaob2!!P>xHkel^T~(Y-xE++s>r4yMnS)S^8rR_Z2p{17Mff81otpkaVS}WORd7UH zRB5`*k7PWocxOV0`Q^x0H|6pPB=Mwo6v%p+JT#}mYWs0!W_hf_XMFFAzyBg$u< zhe$vU*1p&MGC;b{)ltjN0OV$LF5*3`BbLY%g%@erbD@u{xCVV>l07v21gJ2+#dwJ9 zJ@tv^t7?wXu!e_=L*O9!N~xDotbo&7&BUitRNlX@T6H-p7T(Q1o5avXGFXP1D}|vw zN`577Ic&_)4k*C3pV5>eC{HL;A-PzCMn^rPt2%YO+YFR8^r^e_FAW{Hqe8j_)m5FQ zcpbf9m|;nGOC z8vHT+@UcI02AC@>f7DsA7ZIEn-}k>TLz`ZdaXkq1dEa_*@8ST>Ez3&9R=Slj z(XQN4EWddXgX;Pp3{VAT}_YwxQYs)|6-!Fv3@CS_ek7MK)qhd&N-gU2Jj86{9RjqW3cAIh4uIUrXAs9 z0V{_2#`E<0>2nh`^Tu)r(iddI#rNoiAD;Pw23o0yy#oV~VYBy~xyFyowR@2V(%BjH zry!_bXoDlwRnf+#h|@~D<37r}-JhR)3jCM;sj2?dyT-i3n*P`_cpd<*3Ght4&@IVp z?8Uc>M#T&hWsjMZ)RmuMuJ*RSP-7h<2lw!Lyec12IvRe^aqpY$LPA3P_lJ)zW?)MN zw+`Ix*x&i()By#3?Cr(AT;fm1eg{Vx0K2lI~UkjTQwVGN5QP3MCjJ{%$NF&NvmI0;x4 zM4NHE#-^2_K)RCs=1EC=61R-uMF%&JcpIUdNs?B5-%6ITs4nkJKD`!A`3YtitEA78VxBsSAI6X0aMYfB#vlC`wGQs{mX;Nz?&}Ze|7C zaeZ9#KW?tIp9Fp^8ag9TiJ)Nl2nm7}S9oLp5N44rWQiwAi*{~#*oDuC51 z+6~JIw%vbP(=~&<*IRr?CO&bA_(y{EoyCuQ7Vk;oW&h)6*ptRIgi~;2LtHwCP|W>w zni$Y2Tk~;j6b)`PQxa5Zwvn;p|=;CGQPkWWz$7pX-);vxaH4DGxrX{t$acJ2<0iHpPSJB**- z5*IIq6F-VK+5}cX-rkDu{#+hscHV{X!itlvQRuDs=Qo@KByO7Xpx8w z9#DU4KIcUv032F+Unp`_bo?Gvc?abI2L+Iv8a(aTOhA7cbXycuqZIWZgTT@f7*fYG z(01{dNS(>Xh!%h5wZ;yk4_J2vc@HwqQ5caeM1SCUjEw@oMmkxL6YNpBB{+Pw+&qLl zPxQelXDCt44{VWAqRgPwp-D7Xz&&_aG#3RU>IJr1*I~;()2?n{g_ z#9Si%j#u5T1x5c2lfx&ND3R{D=(A}zh!m=*&=QT69=rsL&<^#`D zv_@)XyaEezCtb(#{D1Yqpx#3W-Ivo0ffVmMlj(c%Ut$6bu_pB=K6flkL!6GD&>h8R zAFQ{1+qSXcJ~R=ZY8!s z<#IfKoDkT#FL}xOHdwd7qM5x`b)rtSzpX;MNO`O!!%@e4x!ZAc45a<$iK> z8q%JhC@=wj{%PYe9ilyrz(mOC81P_r+Jbp2t`>~`qaHY6(1uSHTr#(;UNndBsUj9f z`_IpG9|qXC232FjMm9df%byrc3c(sDt1k5d0e=(`#$|ya6By@$0sE2H$tC;Vj*sMt*3oem5OTNd94{IjRV4fApY2^ByWY3%)GZ(#B z*wqO164x2-GK>kh`S`3bSj4ZX5m>W^eZvOY+xYGEzkvbd<7}VsoYW$DF_j8S4@0W7 zQ!mR2ukTn7n~B46-Fb)s2^za!p)NzxzDL!c3iX8Feg6kz61@DD@$@5eh)C?^n(s9w z631Sc!)*4S&~Ofsu<|GI-YeoNeEzRS0dLm~brW^>VjgxC};40(8%v;hDQfGLnw zYS6P#21&>6s!YzY=A%W=2$UsCD`s`ssOS-0F?j7dV{nm;qr=w}tGPEr^fD5Kc>Ac+ znn);;H)W#yMc`>Umi-Lx-2T0LPfWdWt)MS@I3>g6#j8=Qxo)z&hLcI=mxEuZvZwHyKaoF%G69F#Kk;?gn+0p|W=tf2x2Ebyf-(G&#@{RTbbkzsJFb*^$% zs(Ola)UQlyCMhszEM|x4<9)mRH!P7#sww+xFCb%Hb(e_8K9V`Hu~BW zPMnFo4*c2hx~;S=v5Em}H9gB{KH0_-DQ7Yyo-jww>W2T8Z2=Qda9`Xkr(}^xR?Ri_ z;8LE7mim+5M^<|LxF$US_K@*CnG(gfNmndPCZo|EJR$$IFSMDi_+VDjj|s`2h0~!M zs!JnsWw73BvX4=5M0JC*-n$r?uChu~%K3zoj$*WOGW`O3_!Hh{ZDDg%ePV9Wo=@+L zwq7CEOs$W!m+h_m+R9~$yS!N~=nTe<+%ECV(b7SkH2old?+ugpS&>SwZ-2Bm(MQPEZt{tQRz%YkQ`f&4fy{B$Yyh?qWZh{Y?4k-I%- zWRubI-tP&u{&l}a$Hf(b;`?HFyRU>ggml6a&Vp`Fgd|V9RFGhT*N5(=e zR^qoiw9@+yCB3(X_91S(tGRJ|t=yjZz#Xy1mxBKKy9MkM6tsnTMo|%SL%2%cLR-?B zq_bbv4;n`O?Q_Va8z=Y1NpM2#ibT&9_8w4o^KkX~tx2D~ZfB+FA5YvsX|Jw7Lhj&; z@lWJ&$L8l3qZE(o(mub(HHuA2%Q`n~zCCs8n15TkvPnSn0a0iZsWEIm@O9@v)$G_? zYXaEdgtk)WoPayRC05S|w)C&2?HI>P+cG%F=oN!sODR~$9C5L-uwp-3Jc6&EaTcSb zH(np7wzf%yI%Vg|@j97C5;j%BY@3k0h3bjusVZs?)Z4|I4p2}q#W&S@WnLdorT}*4#)ssJDN80&szSv^Ax_iNuEu|JZE_tiu(<@K$=r-d z+>?}%rATD1B&vSA;ma)TeRiYBd1T{4?K!b*0&uVZifov7`8${|WL*ptQsar%&iJ|6 z82s8Z^UysaHZ?d}5_X)6(bWzA-j5!zzbg^SAy~TXGt^TR-pG1`=g@V;AtMaPi)LYAb2yP<(_-P z>f*O=E!Ux&o&DtOU=Qr60fIN~eZuIgPI0q8vk62u#9WkbIodGtSNFtS@oeH(g20qq z4e(FU65M`#x+(qx?dza@8c$G3A-}N|7Z?*Kjvw92z&8^g4?GplR&RwYL1w!54e0Ia zD7ykV@z=jN&cTL$XwfxdX_bV8zDhuFZ`?f%7H~rHns_aT+gv3cRK7&_w9HAQPgr=j zZ}?NZ;g%UK4;o=|_|Pw0%NyTq1vQFz?$}tLP-9tg(tjeWfgDvS(Z>VJ&tB79> zg6B9h^~r`$PGL_CS2kK^VEba_QcJ__$xSNBH9Fs60Pj(#QwSWq=5WZuY%$&AHi0C^ zs@c2ub%R`;<#?)!- z={$P00YqCG+SnEy4x!9G&Q~fmwSejMQFQpqA&Knrp_2=@QXQeqc84c*Cp$TmaQqzyePS(O}1e_&e~!vepQ#4 zb&rAvcI%iN9XwHWt+-m;-Krbil*Tido!k?|A@j+uu!~Wqy&}+5ZYppq$@R)D57M8l z_&DM!d_8;<9U!<^p%0EJY3-r3?E=t8-s^dZB{PZFqh5X)=b-t#_T_xQ>G?3@X(i5A zym$QRE1Vt34ldxjXxDDcaqJY`9M{wwZzfkUR*v&?i!;_U#qc1#>D?js*%>6=nTLdh zvoNdl7I2keuh6SSlO7PLtpAL@Lqt?#wr>SJy^f`@%T%J zU0=tH!+cwsU-qZj0HRnG84SCtq8)cZ)m^vB_Z!J2NBCf}AJDJ;F_&9CWPgdrpNE-p zq0g`Js8&P+3)$4&k&XsYY+4wFp>#2aV}f7w^QZ441b&i;5E(wzctN=#QTOSw2 z80#%+MD==Y;olm(2>O+gu)cy0>W1%=e2%?$iq3a8A8eEm_u5PUGB)eNFKNj7mfD!w z6yF#<;+3$bOXEP;TY7glF^N2;ejrxZu)!LOxUSW-vB>47_hh~}@SS5x;B}T<;&XR= zNBowkJr;9eb=XJ85_S8Gg$$dH)@eK#vYvVSpgs;8+= zNdzqIckups=w!F4e^A)U#t~PCsJUf5M|v3>zGSWBiCfz5T7PLV{cXxo^%r$#ug>|z zuaDn6b!dOV@W*rHOLf?7JoBY8z-t|phWjgAGBtzFKV zUU9O0u5t7OEj{jC^dfB>BFzV>67A*5Pg-tm(6P8BhN4k&3A>-rl1cFsChR^YVH@k8 zYaNm8Wb|=Vwu(e65c8HHOd55aHy0LvMh~mjIi?v)6$9l5$DcX1M9tnR-+eaWs@yk4 zWyL)$VNFM?PwmsoR$3~7J~Q3MLbO>bi({wIY{F6Xr>dvk?2w}yguIW~tP@_7ecSVf zleztSRB%)Vjh)ppzaQnP5Nv!w%K8eR%d6mVX!JW4Ef)J|_YwM915#>y?+&lW&#*e8la)_ql!7&vTN7H3=Q<44O>1njr*H)a2Bvm}v z`&wCF;S{G{Lpv1hj<$p=(c5k}U9`BI;Q12Gl=vN{BJs9Y6_+}{E8WWRJfzU7j(Z-+ zPt8@>p%6CuL$-)_bhbw^LziX*f~+HE_#rPs%P~Phb?+?G(u8r)9+k?e1|Sc&*ViSU zNkNPp-5O8udu~{-X?|e_79JPL8Hg+_uFpaht`pu8aj$d=;$G!y8=P-rD`@PR zlDy4At@#JIlA282xRT~{c#iaX)t%~pv)`kvd)%1%Lw)1)jxno_@%MaHLHuy@1GBYl z(yX`=LjIh=`wF3^h~26owK|cW7JJ(p$S$Gz&Tf}-O$F8K7u>T# z<>GbE_)ZAjujsYxs8`=^(64UWEFGTxCHDv;d&4Q{Ps3WZtrQn)bu;B>(*@$27&K`L zzM^saexDR~XMNT1K8LZLwT~gR-74;A?cPw!ylkhExOeVbnAhhFv^dHe+8a9u25j<} zZr(G=$M){ov{RUm*!R(8<=H`os**kdc9~9Ysg!1#6!EQhecUtocBp&*nNb9VQQYv5 zVcK9(!x6lGrL)to-OM9m(+cc-v=5c=q$|d>$m5+0QP$=hyHzDVZ!3EuYpTjUzO^|D$^XBSCrow9qx734vg$qFdR0-6_+`sgK5Z+GY6 z-?>Wm2$-Ua1`%c**}?j&fy%uJT4M}TsmvPuEX>U6%q)%bFScSE1v>i=lte>f;Vi{q z4u3Jeam|#6C3mC=8D@F6+miJVx`kPnjl3KC9i4wxRtT!1pN1@a)!&g}V-R0@X5Bu+ u@x|%-!PFMW-)_qz`*3e(C?n)>nq5fkV%m?_WU1{!y;$^ literal 0 HcmV?d00001 From 8de7f12ef8022b74ed9c941fb74491ea8fb7cb38 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 02:57:34 +0800 Subject: [PATCH 17/51] Add bounded Runtime history export seam --- .../runtime-observability-design.md | 62 +++- services/agents-api/cmd/server/main.go | 5 + .../internal/runtimeobs/exporter.go | 196 ++++++++++++ .../agents-api/internal/runtimeobs/service.go | 73 +++-- .../internal/runtimeobs/service_test.go | 283 ++++++++++++++++++ 5 files changed, 599 insertions(+), 20 deletions(-) create mode 100644 services/agents-api/internal/runtimeobs/exporter.go diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index be3153ada..26e6d5718 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -2,9 +2,11 @@ Status: provider abstraction with Docker and microsandbox sampling, the current-snapshot API/client contract, and the initial Core Web current-snapshot -Dashboard are implemented. The history backend and other provider sources are not -implemented. Microsandbox idle suspension is a separate durable lifecycle feature; -it does not consume this telemetry as authority. +Dashboard are implemented. Phase 4 backend qualification and the bounded, +sanitized exporter seam are implemented, but no concrete export transport, +history backend, or history API is configured. Other provider sources are not +implemented. Microsandbox idle suspension is a separate durable lifecycle +feature; it does not consume this telemetry as authority. ## 1. Problem statement @@ -238,6 +240,51 @@ Provider type, Runtime mode, and coarse status are safe low-cardinality labels. High-cardinality identities require tenant-scoped access and retention policies; they are not global Prometheus labels by default. +### 10.1 Qualified public implementation reference + +The first Phase 4 qualification uses E2B Runtime commit +`ccf2a64ee40472645209b92525a5459d413bce76` as implementation evidence, not as +an API contract to copy. Its sandbox observer samples every five seconds, exports +provider metrics through OTLP, and attaches sandbox and team identity. The +OpenTelemetry Collector sends ordinary operational metrics to Mimir but routes +the high-cardinality `e2b.*` sandbox series to ClickHouse. Its authenticated API +derives team identity from the caller, queries with both `team_id` and +`sandbox_id`, validates the requested time range, calculates a bounded step, and +retains the specialized sandbox table for seven days. Relevant public files are: + +- [`packages/orchestrator/pkg/metrics/sandboxes.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/orchestrator/pkg/metrics/sandboxes.go) + for bounded collection and identity attributes; +- [`packages/local-dev/otel-collector.yaml`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/local-dev/otel-collector.yaml) + for OTLP fan-out to Mimir and ClickHouse; +- [`packages/clickhouse/migrations/20250717135224_sandbox_metrics.sql`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/clickhouse/migrations/20250717135224_sandbox_metrics.sql) + for the high-cardinality history schema and retention; and +- [`packages/api/internal/clusters/resources_local.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/api/internal/clusters/resources_local.go) + plus [`packages/clickhouse/pkg/sandbox.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/clickhouse/pkg/sandbox.go) + for tenant-scoped, downsampled reads. + +Dify commit `a068c47ea993ccc0f943131c274b7830b16de9f4` independently demonstrates an +optional OTLP exporter that becomes a no-op when disabled, but it does not +provide a Runtime-incarnation history query boundary. It supports the exporter +choice, not the history adapter design. + +For Core, the qualified topology is therefore: + +1. a bounded, best-effort provider-neutral handoff after validated current + observations; +2. an optional operator-managed OTLP Collector; +3. a high-cardinality history store behind a separate server-side adapter; and +4. authenticated Core history routes that resolve tenant and Session ownership + before issuing a backend query. + +Mimir remains suitable for low-cardinality service health. The initial Runtime +history qualification does not treat a shared Prometheus label filter as a +tenant security boundary and does not let Web query Mimir, ClickHouse, or the +Collector directly. ClickHouse is the first reference backend because its query +shape can require tenant, Session, allocation, and incarnation predicates, but +the public API and `runtimehistory` interface must remain backend-neutral. The +backend, Collector, and exporter are disabled by default and are not required for +Session execution or current observations. + ## 11. Dashboard information architecture ### 11.1 Overview @@ -386,6 +433,15 @@ Implemented for the browser-local current-snapshot live window. ### Phase 4: optional history +- Implemented: bounded asynchronous handoff of sanitized, validated current + observation results. It is disabled by default, drops on queue saturation, and + cannot fail the current-observation request path. +- Qualified: optional OTLP Collector fan-out with a separate high-cardinality + history store and server-side tenant-scoped query adapter. ClickHouse is the + first reference backend; no backend is a Core execution dependency. +- Not implemented: a concrete OTLP exporter, `runtimehistory` query adapter, + public history extension, durable Web ranges, retention configuration, and + exporter coverage telemetry. - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. - Replace or extend the ephemeral live window with explicitly advertised durable diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 59cdefd9e..1f0f7a396 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -109,6 +109,11 @@ func run() error { if err != nil { return err } + defer func() { + closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + defer cancel() + _ = observationService.Close(closeCtx) + }() var workerDone chan error var worker *execution.Worker options := []api.Option{api.WithSubagents(executionStore), api.WithSkills(executionStore), api.WithSourceFiles(executionStore), api.WithSessionArtifacts(executionStore), api.WithRuntimeObservations(observationService)} diff --git a/services/agents-api/internal/runtimeobs/exporter.go b/services/agents-api/internal/runtimeobs/exporter.go new file mode 100644 index 000000000..bbc87f41d --- /dev/null +++ b/services/agents-api/internal/runtimeobs/exporter.go @@ -0,0 +1,196 @@ +package runtimeobs + +import ( + "context" + "errors" + "sync" + "time" +) + +const ( + defaultExportQueueCapacity = 256 + defaultExportTimeout = 2 * time.Second +) + +// ExportRecord is the provider-neutral, sanitized history handoff produced +// after Runtime identity and a current observation result have been validated. +// It intentionally excludes provider receipts, native identifiers, and raw +// provider errors. +type ExportRecord struct { + TenantID string + SessionID string + EnvironmentID string + AllocationID string + + Mode Mode + ProviderType string + Status Status + Reason string + + ResolvedAt time.Time + SourceDuration time.Duration + Sample *Sample +} + +// Exporter persists or forwards sanitized Runtime observation records. Core +// invokes exporters only through a bounded asynchronous dispatcher, so an +// exporter outage cannot fail or block the current-observation request path. +type Exporter interface { + Export(context.Context, ExportRecord) error +} + +type ExportOptions struct { + QueueCapacity int + Timeout time.Duration +} + +type ServiceOption func(*serviceOptions) error + +type serviceOptions struct { + exporter Exporter + exportOptions ExportOptions +} + +// WithExporter enables best-effort history export. Queue saturation drops the +// newest handoff; it never changes the Runtime observation result. +func WithExporter(exporter Exporter, options ExportOptions) ServiceOption { + return func(config *serviceOptions) error { + if exporter == nil { + return errors.New("Runtime observation exporter is required") + } + if options.QueueCapacity < 0 { + return errors.New("Runtime observation export queue capacity cannot be negative") + } + if options.Timeout < 0 { + return errors.New("Runtime observation export timeout cannot be negative") + } + config.exporter = exporter + config.exportOptions = options + return nil + } +} + +type exportDispatcher struct { + exporter Exporter + timeout time.Duration + queue chan ExportRecord + done chan struct{} + + mu sync.Mutex + closed bool + cancel context.CancelFunc +} + +func newExportDispatcher(exporter Exporter, options ExportOptions) *exportDispatcher { + capacity := options.QueueCapacity + if capacity == 0 { + capacity = defaultExportQueueCapacity + } + timeout := options.Timeout + if timeout == 0 { + timeout = defaultExportTimeout + } + ctx, cancel := context.WithCancel(context.Background()) + dispatcher := &exportDispatcher{ + exporter: exporter, + timeout: timeout, + queue: make(chan ExportRecord, capacity), + done: make(chan struct{}), + cancel: cancel, + } + go dispatcher.run(ctx) + return dispatcher +} + +func (d *exportDispatcher) enqueue(record ExportRecord) { + d.mu.Lock() + defer d.mu.Unlock() + if d.closed { + return + } + select { + case d.queue <- record: + default: + } +} + +func (d *exportDispatcher) run(ctx context.Context) { + defer close(d.done) + for record := range d.queue { + exportCtx, cancel := context.WithTimeout(ctx, d.timeout) + exportSafely(exportCtx, d.exporter, record) + cancel() + if ctx.Err() != nil { + return + } + } +} + +func exportSafely(ctx context.Context, exporter Exporter, record ExportRecord) { + defer func() { + _ = recover() + }() + _ = exporter.Export(ctx, record) +} + +func (d *exportDispatcher) close(ctx context.Context) error { + d.mu.Lock() + if !d.closed { + d.closed = true + close(d.queue) + } + d.mu.Unlock() + + select { + case <-d.done: + d.cancel() + return nil + case <-ctx.Done(): + d.cancel() + return ctx.Err() + } +} + +func exportRecord(observation Observation) ExportRecord { + return ExportRecord{ + TenantID: observation.Target.TenantID, + SessionID: observation.Target.SessionID, + EnvironmentID: observation.Target.EnvironmentID, + AllocationID: observation.Target.Instance.AllocationID, + Mode: observation.Target.Mode, + ProviderType: observation.ProviderType, + Status: observation.Status, + Reason: observation.Reason, + ResolvedAt: observation.ResolvedAt, + SourceDuration: observation.SourceDuration, + Sample: cloneSample(observation.Sample), + } +} + +func cloneSample(sample *Sample) *Sample { + if sample == nil { + return nil + } + cloned := *sample + if sample.StartedAt != nil { + value := *sample.StartedAt + cloned.StartedAt = &value + } + if sample.CPUUsageSecondsTotal != nil { + value := *sample.CPUUsageSecondsTotal + cloned.CPUUsageSecondsTotal = &value + } + if sample.CPUCapacityCores != nil { + value := *sample.CPUCapacityCores + cloned.CPUCapacityCores = &value + } + if sample.MemoryUsageBytes != nil { + value := *sample.MemoryUsageBytes + cloned.MemoryUsageBytes = &value + } + if sample.MemoryLimitBytes != nil { + value := *sample.MemoryLimitBytes + cloned.MemoryLimitBytes = &value + } + return &cloned +} diff --git a/services/agents-api/internal/runtimeobs/service.go b/services/agents-api/internal/runtimeobs/service.go index e56d17348..34cf222a8 100644 --- a/services/agents-api/internal/runtimeobs/service.go +++ b/services/agents-api/internal/runtimeobs/service.go @@ -4,28 +4,42 @@ import ( "context" "errors" "fmt" + "regexp" "time" ) +var providerTypePattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) + type Observation struct { - Target Target - Status Status - Sample *Sample - Reason string - ProviderType string - ResolvedAt time.Time + Target Target + Status Status + Sample *Sample + Reason string + ProviderType string + ResolvedAt time.Time + SourceDuration time.Duration } type Service struct { resolver TargetResolver sources map[string]Source now func() time.Time + exports *exportDispatcher } -func NewService(resolver TargetResolver, sources map[string]Source) (*Service, error) { +func NewService(resolver TargetResolver, sources map[string]Source, options ...ServiceOption) (*Service, error) { if resolver == nil { return nil, errors.New("Runtime observation resolver is required") } + config := serviceOptions{} + for _, option := range options { + if option == nil { + return nil, errors.New("invalid Runtime observation service option") + } + if err := option(&config); err != nil { + return nil, err + } + } copySources := make(map[string]Source, len(sources)) for key, source := range sources { if key == "" || source == nil { @@ -33,7 +47,27 @@ func NewService(resolver TargetResolver, sources map[string]Source) (*Service, e } copySources[key] = source } - return &Service{resolver: resolver, sources: copySources, now: time.Now}, nil + service := &Service{resolver: resolver, sources: copySources, now: time.Now} + if config.exporter != nil { + service.exports = newExportDispatcher(config.exporter, config.exportOptions) + } + return service, nil +} + +// Close drains pending history handoffs within ctx. Current-observation callers +// may keep using a Service without an exporter; Close is then a no-op. +func (s *Service) Close(ctx context.Context) error { + if s.exports == nil { + return nil + } + return s.exports.close(ctx) +} + +func (s *Service) finish(observation Observation) Observation { + if s.exports != nil { + s.exports.enqueue(exportRecord(observation)) + } + return observation } func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string) (Observation, error) { @@ -43,7 +77,7 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string if target.TenantID != tenantID || target.SessionID != sessionID || target.Mode != ModeManaged || target.EnvironmentID == "" { return Observation{}, errors.New("Runtime observation resolver returned invalid pending allocation identity") } - return Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}), nil } if err != nil { return Observation{}, err @@ -56,7 +90,7 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string return Observation{}, errors.New("Runtime observation resolver returned mismatched Environment identity") } if target.Mode == ModeNone || target.Mode == ModeSelfHosted { - return Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: resolvedAt}, nil + return s.finish(Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: resolvedAt}), nil } if target.Mode != ModeManaged || target.Instance.AllocationID == "" || target.Instance.ProviderKey == "" { return Observation{}, errors.New("invalid managed Runtime observation target") @@ -67,30 +101,35 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string } switch target.Instance.AllocationState { case "creating": - return Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}), nil case "cleanup_pending", "released": - return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ResolvedAt: resolvedAt}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ResolvedAt: resolvedAt}), nil case "running": default: return Observation{}, errors.New("invalid managed Runtime allocation state") } source, ok := s.sources[target.Instance.ProviderKey] if !ok { - return Observation{Target: target, Status: StatusUnavailable, Reason: "source_not_configured", ResolvedAt: resolvedAt}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "source_not_configured", ResolvedAt: resolvedAt}), nil } providerType := "" if typed, ok := source.(interface{ ObservationProviderType() string }); ok { providerType = typed.ObservationProviderType() + if providerType != "" && !providerTypePattern.MatchString(providerType) { + return Observation{}, errors.New("invalid Runtime observation provider type") + } } + sourceStarted := time.Now() sample, err := source.Observe(ctx, target) + sourceDuration := time.Since(sourceStarted) if errors.Is(err, context.DeadlineExceeded) { - return Observation{Target: target, Status: StatusUnavailable, Reason: "sample_timeout", ProviderType: providerType, ResolvedAt: s.now()}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "sample_timeout", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil } if errors.Is(err, ErrNotRunning) { - return Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ProviderType: providerType, ResolvedAt: s.now()}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil } if errors.Is(err, ErrUnavailable) { - return Observation{Target: target, Status: StatusUnavailable, Reason: "sample_unavailable", ProviderType: providerType, ResolvedAt: s.now()}, nil + return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "sample_unavailable", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil } if err != nil { return Observation{}, fmt.Errorf("observe Runtime: %w", err) @@ -98,5 +137,5 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string if err := sample.validate(s.now()); err != nil { return Observation{}, err } - return Observation{Target: target, Status: StatusObserved, Sample: &sample, ProviderType: providerType, ResolvedAt: s.now()}, nil + return s.finish(Observation{Target: target, Status: StatusObserved, Sample: &sample, ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil } diff --git a/services/agents-api/internal/runtimeobs/service_test.go b/services/agents-api/internal/runtimeobs/service_test.go index 6cb76682f..ab9ab3cbe 100644 --- a/services/agents-api/internal/runtimeobs/service_test.go +++ b/services/agents-api/internal/runtimeobs/service_test.go @@ -2,8 +2,10 @@ package runtimeobs import ( "context" + "encoding/json" "errors" "math" + "sync" "testing" "time" ) @@ -51,6 +53,70 @@ func (blockingSource) Observe(ctx context.Context, _ Target) (Sample, error) { func (blockingSource) ObservationProviderType() string { return "docker" } +type channelExporter struct { + records chan ExportRecord + err error +} + +func (e channelExporter) Export(ctx context.Context, record ExportRecord) error { + select { + case e.records <- record: + return e.err + case <-ctx.Done(): + return ctx.Err() + } +} + +type gatedExporter struct { + started chan struct{} + release chan struct{} + + mu sync.Mutex + calls int +} + +type stubbornExporter struct { + started chan struct{} + release chan struct{} +} + +type panicExporter struct { + called chan struct{} +} + +func (e panicExporter) Export(context.Context, ExportRecord) error { + e.called <- struct{}{} + panic("exporter panic must remain isolated") +} + +func (e stubbornExporter) Export(context.Context, ExportRecord) error { + close(e.started) + <-e.release + return nil +} + +func (e *gatedExporter) Export(ctx context.Context, _ ExportRecord) error { + e.mu.Lock() + e.calls++ + e.mu.Unlock() + select { + case e.started <- struct{}{}: + default: + } + select { + case <-e.release: + return nil + case <-ctx.Done(): + return ctx.Err() + } +} + +func (e *gatedExporter) callCount() int { + e.mu.Lock() + defer e.mu.Unlock() + return e.calls +} + func TestServiceDoesNotCallSourcesForUnsupportedModes(t *testing.T) { for _, mode := range []Mode{ModeNone, ModeSelfHosted} { source := &fixedSource{} @@ -95,6 +161,223 @@ func TestServicePreservesUnavailableAndObservedZero(t *testing.T) { } } +func TestServiceExportsOnlySanitizedValidatedRecords(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + startedAt := now.Add(-time.Minute) + cpuSeconds := 12.5 + cpuCapacity := 2.0 + memoryUsage := uint64(1024) + memoryLimit := uint64(2048) + target := Target{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{ + AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running", + ProviderState: json.RawMessage(`{"native_id":"must-not-export"}`), + }, + } + source := typedSource{fixedSource: &fixedSource{sample: Sample{ + ObservedAt: now, StartedAt: &startedAt, + CPUUsageSecondsTotal: &cpuSeconds, CPUCapacityCores: &cpuCapacity, + MemoryUsageBytes: &memoryUsage, MemoryLimitBytes: &memoryLimit, + }}, providerType: "docker"} + records := make(chan ExportRecord, 1) + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": source}, WithExporter(channelExporter{records: records}, ExportOptions{QueueCapacity: 1, Timeout: time.Second})) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusObserved { + t.Fatalf("observation failed: %+v %v", observation, err) + } + cpuSeconds = 99 + memoryUsage = 99 + + select { + case record := <-records: + if record.TenantID != "tenant" || record.SessionID != "session" || record.EnvironmentID != "environment" || record.AllocationID != "allocation" { + t.Fatalf("exported identity mismatch: %+v", record) + } + if record.ProviderType != "docker" || record.Mode != ModeManaged || record.Status != StatusObserved || record.Reason != "" { + t.Fatalf("exported classification mismatch: %+v", record) + } + if record.Sample == nil || record.Sample.CPUUsageSecondsTotal == nil || *record.Sample.CPUUsageSecondsTotal != 12.5 || record.Sample.MemoryUsageBytes == nil || *record.Sample.MemoryUsageBytes != 1024 { + t.Fatalf("exported sample was not independently copied: %+v", record.Sample) + } + case <-time.After(time.Second): + t.Fatal("timed out waiting for Runtime observation export") + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} + +func TestServiceRejectsUnsafeProviderTypeBeforeSamplingOrExport(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + source := typedSource{fixedSource: &fixedSource{sample: Sample{ObservedAt: now}}, providerType: "docker native_id=secret"} + records := make(chan ExportRecord, 1) + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": source}, + WithExporter(channelExporter{records: records}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err == nil || source.calls != 0 { + t.Fatalf("unsafe provider type reached sampling: err=%v calls=%d", err, source.calls) + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } + select { + case record := <-records: + t.Fatalf("unsafe provider type reached exporter: %+v", record) + default: + } +} + +func TestServiceExportQueueNeverBlocksOrChangesObservation(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + exporter := &gatedExporter{started: make(chan struct{}, 1), release: make(chan struct{})} + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now}}}, + WithExporter(exporter, ExportOptions{QueueCapacity: 1, Timeout: time.Second}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + + if observation, err := service.ObserveSession(t.Context(), "tenant", "session"); err != nil || observation.Status != StatusObserved { + t.Fatalf("first observation failed: %+v %v", observation, err) + } + select { + case <-exporter.started: + case <-time.After(time.Second): + t.Fatal("exporter did not start") + } + for range 2 { + completed := make(chan error, 1) + go func() { + observation, observeErr := service.ObserveSession(t.Context(), "tenant", "session") + if observeErr == nil && observation.Status != StatusObserved { + observeErr = errors.New("unexpected observation status") + } + completed <- observeErr + }() + select { + case observeErr := <-completed: + if observeErr != nil { + t.Fatal(observeErr) + } + case <-time.After(100 * time.Millisecond): + t.Fatal("Runtime observation blocked on history exporter") + } + } + close(exporter.release) + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } + if calls := exporter.callCount(); calls != 2 { + t.Fatalf("export queue should retain one pending record and drop overflow, got %d calls", calls) + } +} + +func TestServiceIgnoresExporterFailureAndValidatesOptions(t *testing.T) { + if _, err := NewService(fixedResolver{}, nil, WithExporter(nil, ExportOptions{})); err == nil { + t.Fatal("nil exporter was accepted") + } + if _, err := NewService(fixedResolver{}, nil, WithExporter(channelExporter{}, ExportOptions{QueueCapacity: -1})); err == nil { + t.Fatal("negative export queue capacity was accepted") + } + if _, err := NewService(fixedResolver{}, nil, WithExporter(channelExporter{}, ExportOptions{Timeout: -time.Second})); err == nil { + t.Fatal("negative export timeout was accepted") + } + + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + records := make(chan ExportRecord, 1) + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now}}}, + WithExporter(channelExporter{records: records, err: errors.New("backend unavailable")}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusObserved { + t.Fatalf("exporter failure changed observation: %+v %v", observation, err) + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} + +func TestServiceCloseHonorsItsDeadlineWhenExporterDoesNot(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + exporter := stubbornExporter{started: make(chan struct{}), release: make(chan struct{})} + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now}}}, + WithExporter(exporter, ExportOptions{QueueCapacity: 1, Timeout: time.Millisecond}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err != nil { + t.Fatal(err) + } + select { + case <-exporter.started: + case <-time.After(time.Second): + t.Fatal("exporter did not start") + } + + closeCtx, cancel := context.WithTimeout(t.Context(), 20*time.Millisecond) + defer cancel() + if err := service.Close(closeCtx); !errors.Is(err, context.DeadlineExceeded) { + t.Fatalf("Close did not preserve its deadline: %v", err) + } + close(exporter.release) +} + +func TestServiceIsolatesExporterPanics(t *testing.T) { + now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) + exporter := panicExporter{called: make(chan struct{}, 1)} + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now}}}, + WithExporter(exporter, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + observation, err := service.ObserveSession(t.Context(), "tenant", "session") + if err != nil || observation.Status != StatusObserved { + t.Fatalf("observation failed: %+v %v", observation, err) + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } + select { + case <-exporter.called: + default: + t.Fatal("panicking exporter was not invoked") + } +} + func TestServiceMapsOnlyDeclaredUnavailability(t *testing.T) { target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} for _, tc := range []struct { From fabcba4ff95f4cd655c8109ce7499b8a63fd8da6 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 03:33:13 +0800 Subject: [PATCH 18/51] Add optional OTLP Runtime history export --- .env.example | 3 + .../runtime-observability-design.md | 69 ++++- contracts/agents-api/runtime-observability.md | 19 +- go.mod | 16 +- go.sum | 22 +- services/agents-api/cmd/server/main.go | 18 +- .../agents-api/cmd/server/runtime_history.go | 115 ++++++++ .../cmd/server/runtime_history_test.go | 88 ++++++ .../runtimeobs/otlpexporter/exporter.go | 274 ++++++++++++++++++ .../runtimeobs/otlpexporter/exporter_test.go | 226 +++++++++++++++ 10 files changed, 834 insertions(+), 16 deletions(-) create mode 100644 services/agents-api/cmd/server/runtime_history.go create mode 100644 services/agents-api/cmd/server/runtime_history_test.go create mode 100644 services/agents-api/internal/runtimeobs/otlpexporter/exporter.go create mode 100644 services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go diff --git a/.env.example b/.env.example index b58a813f9..eb23fc2ea 100644 --- a/.env.example +++ b/.env.example @@ -5,6 +5,9 @@ AGENTS_API_PROXY_TARGET=http://127.0.0.1:8091 # Preferred: plaintext execution-tenant bearer file read only by the local # Vite server. If unset, Web also checks this conventional path automatically. AGENTS_API_PROXY_TOKEN_FILE=~/.parsar/agents-api/web-token +# Optional server-only OTLP/HTTP Runtime history exporter configuration. +# When unset, current observations and the Dashboard require no telemetry backend. +# AGENTS_API_RUNTIME_HISTORY_FILE=/absolute/path/runtime-history.json # Alternative server-side value. Never use a VITE_ prefix, and do not set this # together with AGENTS_API_PROXY_TOKEN_FILE. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 26e6d5718..88ae447d8 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -2,11 +2,11 @@ Status: provider abstraction with Docker and microsandbox sampling, the current-snapshot API/client contract, and the initial Core Web current-snapshot -Dashboard are implemented. Phase 4 backend qualification and the bounded, -sanitized exporter seam are implemented, but no concrete export transport, -history backend, or history API is configured. Other provider sources are not -implemented. Microsandbox idle suspension is a separate durable lifecycle -feature; it does not consume this telemetry as authority. +Dashboard are implemented. Phase 4 backend qualification, the bounded sanitized +exporter seam, and its optional OTLP/HTTP transport are implemented. No history +backend, history query API, or durable Web range is configured. Other provider +sources are not implemented. Microsandbox idle suspension is a separate durable +lifecycle feature; it does not consume this telemetry as authority. ## 1. Problem statement @@ -285,6 +285,56 @@ the public API and `runtimehistory` interface must remain backend-neutral. The backend, Collector, and exporter are disabled by default and are not required for Session execution or current observations. +### 10.2 OTLP transport configuration and instruments + +Core enables Runtime history export only when +`AGENTS_API_RUNTIME_HISTORY_FILE` points to a server-only JSON file. With the +variable unset, no exporter is created and no Collector or history store is +required. A minimal configuration is: + +```json +{ + "transport": "otlp_http", + "endpoint": "https://collector.example.com/v1/metrics", + "headers": {"Authorization": "Bearer operator-managed-secret"}, + "queue_capacity": 256, + "timeout_seconds": 2 +} +``` + +The file may contain transport credentials and must never be served to Web or +committed. Plain HTTP requires the explicit combination of an `http` endpoint +and `"insecure": true`; HTTPS rejects that flag. Endpoint userinfo, query +strings, fragments, invalid headers, reserved transport headers, queues above +4096 records, and timeouts above 30 seconds fail startup without echoing config +contents. Export is best effort through the bounded queue documented above. + +The OTLP request uses standard protobuf metrics and these instruments: + +| Instrument | OTLP aggregation | Source | +| --- | --- | --- | +| `agents.runtime.cpu.usage` | monotonic cumulative sum, seconds | provider cumulative CPU counter | +| `agents.runtime.cpu.capacity` | gauge, cores | configured provider capacity | +| `agents.runtime.memory.usage` | gauge, bytes | provider memory usage | +| `agents.runtime.memory.limit` | gauge, bytes | configured provider limit | +| `agents.runtime.sample` | monotonic delta sum | one validated result, including unavailable/unsupported | +| `agents.runtime.sample.duration` | delta histogram, seconds | bounded provider read duration | + +Core-owned tenant, Session, Environment, allocation, mode, provider type, +status, and safe reason are metric attributes. CPU, capacity, and memory points +are exported only when the sample also carries the compute `started_at` fence; +that fence is included as an attribute on every such point. Provider keys, +provider receipts, native container/pod/instance +identifiers, raw errors, paths, and credentials are not attributes. Missing +measurements produce no value point; they are represented only by the explicit +sample status and reason. + +The initial transport exports observations produced by the current read path; it +does not yet run a deployment-wide background sampler. Consequently, a future +history API must expose actual sample coverage and Core Web must not advertise a +durable range until the operator backend and a qualified collection cadence are +both configured. + ## 11. Dashboard information architecture ### 11.1 Overview @@ -436,12 +486,15 @@ Implemented for the browser-local current-snapshot live window. - Implemented: bounded asynchronous handoff of sanitized, validated current observation results. It is disabled by default, drops on queue saturation, and cannot fail the current-observation request path. +- Implemented: optional server-only OTLP/HTTP protobuf transport for the six + documented Runtime instruments, including allocation and compute-incarnation + fencing attributes. Configuration is strict and secrets never reach Web. - Qualified: optional OTLP Collector fan-out with a separate high-cardinality history store and server-side tenant-scoped query adapter. ClickHouse is the first reference backend; no backend is a Core execution dependency. -- Not implemented: a concrete OTLP exporter, `runtimehistory` query adapter, - public history extension, durable Web ranges, retention configuration, and - exporter coverage telemetry. +- Not implemented: deployment-wide sampling cadence, `runtimehistory` query + adapter, public history extension, durable Web ranges, retention configuration, + and exporter queue/drop/error coverage telemetry. - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. - Replace or extend the ephemeral live window with explicitly advertised durable diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 73490e9c8..a8836f074 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -78,6 +78,23 @@ an activity revision and timestamps such as `idle_since` and `shutdown_requested_at`. Metrics, an in-memory cache, or a monitoring backend must not become the lifecycle authority. +## Optional history export boundary + +An optional server-only OTLP/HTTP exporter can forward validated observation +results to an operator Collector when `AGENTS_API_RUNTIME_HISTORY_FILE` is set. +It is disabled by default and does not change the current API result, execution +ownership, or lifecycle authority. The file may contain endpoint authorization +headers and is never returned to Web. Provider receipts, native identifiers, raw +errors, paths, and credentials are excluded from metric attributes. CPU, +capacity, and memory points require the provider-qualified compute `started_at` +fence; unfenced observations export coverage and read duration only. + +This transport currently exports observations produced by the current read path; +it is not a deployment-wide background sampler. The Collector, +high-cardinality history backend, server-side history query adapter, collection +cadence, and durable Web ranges remain separate optional capabilities in the +[full design](runtime-observability-design.md). + ## First-phase boundary The current-snapshot implementation adds no migration, metrics backend, token @@ -87,5 +104,5 @@ attribution or the existing sandbox lifecycle interface. The current API and browser-local Web live window are documented in the [full design](runtime-observability-design.md) and the -[public extension](runtime-observability-api.md). Durable history, additional +[public extension](runtime-observability-api.md). History queries, additional providers, self-hosted telemetry, and idle-policy authority remain later phases. diff --git a/go.mod b/go.mod index 68c42d207..026a7c9c3 100644 --- a/go.mod +++ b/go.mod @@ -3,7 +3,6 @@ module github.com/MiniMax-AI-Dev/parsar go 1.25.13 require ( - connectrpc.com/connect v1.18.1 github.com/BurntSushi/toml v1.6.0 github.com/containerd/errdefs v1.0.0 github.com/go-chi/chi/v5 v5.3.2 @@ -14,7 +13,13 @@ require ( github.com/moby/moby/client v0.6.0 github.com/openai/openai-go/v3 v3.61.0 github.com/pressly/goose/v3 v3.27.3 + go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 + go.opentelemetry.io/otel/sdk v1.44.0 + go.opentelemetry.io/otel/sdk/metric v1.44.0 + go.opentelemetry.io/proto/otlp v1.10.0 golang.org/x/crypto v0.55.0 + golang.org/x/net v0.58.0 + golang.org/x/sys v0.47.0 google.golang.org/protobuf v1.36.12 gopkg.in/yaml.v3 v3.0.1 ) @@ -38,7 +43,7 @@ require ( github.com/tidwall/sjson v1.2.5 // indirect go.opentelemetry.io/auto/sdk v1.2.1 // indirect go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect - go.opentelemetry.io/otel v1.44.0 // indirect + go.opentelemetry.io/otel v1.44.0 go.opentelemetry.io/otel/metric v1.44.0 // indirect go.opentelemetry.io/otel/trace v1.44.0 // indirect ) @@ -53,6 +58,8 @@ require ( ) require ( + github.com/cenkalti/backoff/v5 v5.0.3 // indirect + github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 // indirect github.com/jackc/pgpassfile v1.0.0 // indirect github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect github.com/jackc/puddle/v2 v2.2.2 // indirect @@ -60,9 +67,10 @@ require ( github.com/rogpeppe/go-internal v1.15.0 // indirect github.com/sethvargo/go-retry v0.4.0 // indirect go.uber.org/multierr v1.11.0 // indirect - golang.org/x/net v0.58.0 // indirect golang.org/x/sync v0.22.0 // indirect - golang.org/x/sys v0.47.0 // indirect golang.org/x/text v0.41.0 // indirect golang.org/x/time v0.15.0 // indirect + google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa // indirect + google.golang.org/genproto/googleapis/rpc v0.0.0-20260720211330-0afa2a65878a // indirect + google.golang.org/grpc v1.82.1 // indirect ) diff --git a/go.sum b/go.sum index 864415fa8..d0481a593 100644 --- a/go.sum +++ b/go.sum @@ -1,9 +1,9 @@ -connectrpc.com/connect v1.18.1 h1:PAg7CjSAGvscaf6YZKUefjoih5Z/qYkyaTrBW8xvYPw= -connectrpc.com/connect v1.18.1/go.mod h1:0292hj1rnx8oFrStN7cB4jjVBeqs+Yx5yDIC2prWDO8= github.com/BurntSushi/toml v1.6.0 h1:dRaEfpa2VI55EwlIW72hMRHdWouJeRF7TPYhI+AUQjk= github.com/BurntSushi/toml v1.6.0/go.mod h1:ukJfTF/6rtPPRCnwkur4qwRxa8vTRFBF0uk2lLoLwho= github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERoyfY= github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU= +github.com/cenkalti/backoff/v5 v5.0.3 h1:ZN+IMa753KfX5hd8vVaMixjnqRZ3y8CuJKRKj1xcsSM= +github.com/cenkalti/backoff/v5 v5.0.3/go.mod h1:rkhZdG3JZukswDf7f0cwqPNk4K0sa+F97BxZthm/crw= github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs= github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs= github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI= @@ -32,6 +32,8 @@ github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag= github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE= github.com/golang-jwt/jwt/v5 v5.3.1 h1:kYf81DTWFe7t+1VvL7eS+jKFVWaUnK9cB1qbwn63YCY= github.com/golang-jwt/jwt/v5 v5.3.1/go.mod h1:fxCRLWMO43lRc8nhHWY6LGqRcf+1gQWArsqaEUEa5bE= +github.com/golang/protobuf v1.5.4 h1:i7eJL8qZTpSEXOPTxNKhASYpMn+8e5Q6AdndVa1dWek= +github.com/golang/protobuf v1.5.4/go.mod h1:lnTiLA8Wa4RWRcIUkrtSVa5nRhsEGBg48fD6rSs7xps= github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8= github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU= github.com/google/jsonschema-go v0.4.3 h1:/DBOLZTfDow7pe2GmaJNhltueGTtDKICi8V8p+DQPd0= @@ -40,6 +42,8 @@ github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0= github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo= github.com/gorilla/websocket v1.5.3 h1:saDtZ6Pbx/0u+bgYQ3q96pZgCzfhKXGPqt7kZ72aNNg= github.com/gorilla/websocket v1.5.3/go.mod h1:YR8l580nyteQvAITg2hZ9XVh4b55+EU/adAjf1fMHhE= +github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 h1:5VipnvEpbqr2gA2VbM+nYVbkIF28c5ZQfqCBQ5g2xfk= +github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0/go.mod h1:Hyl3n6Twe1hvtd9XUXDec4pTvgMSEixRuQKPTMH2bNs= github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM= github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg= github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo= @@ -109,14 +113,20 @@ go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4 go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI= go.opentelemetry.io/otel v1.44.0 h1:JjwHmHpA4iZ3wBxluu2fbbE7j4kqlE8jXyAyPXH7HqU= go.opentelemetry.io/otel v1.44.0/go.mod h1:BMgjTHL9WPRlRjL2oZCBTL4whCGtXch2H4BhOPIAyYc= +go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA= +go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo= go.opentelemetry.io/otel/metric v1.44.0 h1:1w0gILTcHdr3YI+ixLyjemwrVnsMURbTZFrSYCdDdmc= go.opentelemetry.io/otel/metric v1.44.0/go.mod h1:8O7hanEPBNgEMmybD3s2VBKcgWOCsA6tzHBPODAiquo= +go.opentelemetry.io/otel/metric/x v0.66.0 h1:YkCrx1zLOChi9ZcZ6euupOcsgzbVlec7D/xoEU1+cTA= +go.opentelemetry.io/otel/metric/x v0.66.0/go.mod h1:d1+BDj9t96do0/1LoU1ayfCv79ZgNE41qbhBvnMOBZk= go.opentelemetry.io/otel/sdk v1.44.0 h1:nHYwb9lK+fJPU/dnT6s7W7Z8itMWyqrnVfbheVYrZ58= go.opentelemetry.io/otel/sdk v1.44.0/go.mod h1:Osuydd3Se74nqjAKxid74N5eC+jfEqfTegHRnq58oK0= go.opentelemetry.io/otel/sdk/metric v1.44.0 h1:3LlKgI+VjbVsjNRFZJZAJ30WjXC5VkNRks6si09iEfI= go.opentelemetry.io/otel/sdk/metric v1.44.0/go.mod h1:5B5pMARnXxKhltooO4xUuCBorl65a4EpnTalObqOigA= go.opentelemetry.io/otel/trace v1.44.0 h1:jxF5CsGYCe74MCRx2X4g7WsY/VBKRqqpNvXlX/6gtIk= go.opentelemetry.io/otel/trace v1.44.0/go.mod h1:oLl1jrMQAVo6v3GAggN+1VH9VIz9iUSvW53sW1Q8PIE= +go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g= +go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk= go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= go.uber.org/multierr v1.11.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y= golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M= @@ -135,6 +145,14 @@ golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U= golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno= golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE= golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk= +gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4= +gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E= +google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa h1:Kjn0N0tCrDgiAFW+lGO4JZ3ck44CehvJQMAwj9QF0G8= +google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:q4lMZS6kskjT5HvCPrnnypcDPVJqT/f4nfxmkE7gryY= +google.golang.org/genproto/googleapis/rpc v0.0.0-20260720211330-0afa2a65878a h1:qI/YMH1ep2qQtqcp00gMQyoU7mjvbhg88GJKCvfoLj0= +google.golang.org/genproto/googleapis/rpc v0.0.0-20260720211330-0afa2a65878a/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8= +google.golang.org/grpc v1.82.1 h1:NnAxzGRA0677vCa4BUkOAnO5+FfQqVl9iUXeD0IqcGE= +google.golang.org/grpc v1.82.1/go.mod h1:yzTZ1TB1Z3SG+LIYaI+WiE8D5+PZ3ArnrSp8zF3+/ZA= google.golang.org/protobuf v1.36.12 h1:pJOKDDOyeXErUroCihFAd5LQuwXBSpVnKGrj5o/fwxc= google.golang.org/protobuf v1.36.12/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 1f0f7a396..5cc190472 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -105,14 +105,30 @@ func run() error { if err != nil { return err } - observationService, err := runtimeobs.NewService(resolver, observationSources) + historyOption, historyExporter, err := runtimeHistory(ctx) if err != nil { return err } + observationOptions := []runtimeobs.ServiceOption{} + if historyOption != nil { + observationOptions = append(observationOptions, historyOption) + } + observationService, err := runtimeobs.NewService(resolver, observationSources, observationOptions...) + if err != nil { + if historyExporter != nil { + closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + defer cancel() + closeRuntimeHistory(closeCtx, historyExporter) + } + return err + } defer func() { closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() _ = observationService.Close(closeCtx) + if historyExporter != nil { + closeRuntimeHistory(closeCtx, historyExporter) + } }() var workerDone chan error var worker *execution.Worker diff --git a/services/agents-api/cmd/server/runtime_history.go b/services/agents-api/cmd/server/runtime_history.go new file mode 100644 index 000000000..0b03adf4e --- /dev/null +++ b/services/agents-api/cmd/server/runtime_history.go @@ -0,0 +1,115 @@ +package main + +import ( + "bytes" + "context" + "encoding/json" + "errors" + "io" + "net/url" + "os" + "strings" + "time" + + "golang.org/x/net/http/httpguts" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" +) + +const ( + defaultRuntimeHistoryQueueCapacity = 256 + defaultRuntimeHistoryTimeoutSeconds = 2 + maxRuntimeHistoryQueueCapacity = 4096 + maxRuntimeHistoryTimeoutSeconds = 30 +) + +type runtimeHistoryConfig struct { + Transport string `json:"transport"` + Endpoint string `json:"endpoint"` + Insecure bool `json:"insecure"` + Headers map[string]string `json:"headers,omitempty"` + QueueCapacity int `json:"queue_capacity,omitempty"` + TimeoutSeconds int `json:"timeout_seconds,omitempty"` +} + +type runtimeHistoryExporter interface { + runtimeobs.Exporter + Close(context.Context) error +} + +func runtimeHistory(ctx context.Context) (runtimeobs.ServiceOption, runtimeHistoryExporter, error) { + file := os.Getenv("AGENTS_API_RUNTIME_HISTORY_FILE") + if file == "" { + return nil, nil, nil + } + raw, err := os.ReadFile(file) + if err != nil { + return nil, nil, errors.New("cannot read AGENTS_API_RUNTIME_HISTORY_FILE") + } + var config runtimeHistoryConfig + decoder := json.NewDecoder(bytes.NewReader(raw)) + decoder.DisallowUnknownFields() + if decoder.Decode(&config) != nil || decoder.Decode(new(any)) != io.EOF { + return nil, nil, errors.New("invalid Runtime history configuration") + } + if err := validateRuntimeHistoryConfig(config); err != nil { + return nil, nil, err + } + if config.QueueCapacity == 0 { + config.QueueCapacity = defaultRuntimeHistoryQueueCapacity + } + if config.TimeoutSeconds == 0 { + config.TimeoutSeconds = defaultRuntimeHistoryTimeoutSeconds + } + timeout := time.Duration(config.TimeoutSeconds) * time.Second + exporter, err := otlpexporter.New(ctx, otlpexporter.Config{ + Endpoint: config.Endpoint, Headers: config.Headers, Insecure: config.Insecure, RequestTimeout: timeout, + }) + if err != nil { + return nil, nil, err + } + return runtimeobs.WithExporter(exporter, runtimeobs.ExportOptions{QueueCapacity: config.QueueCapacity, Timeout: timeout}), exporter, nil +} + +func closeRuntimeHistory(ctx context.Context, exporter runtimeHistoryExporter) { + defer func() { + _ = recover() + }() + _ = exporter.Close(ctx) +} + +func validateRuntimeHistoryConfig(config runtimeHistoryConfig) error { + if config.Transport != "otlp_http" { + return errors.New("Runtime history transport must be otlp_http") + } + endpoint, err := url.Parse(config.Endpoint) + if err != nil || endpoint.Host == "" || endpoint.User != nil || endpoint.RawQuery != "" || endpoint.Fragment != "" || endpoint.RawPath != "" || endpoint.Path == "" || endpoint.String() != config.Endpoint { + return errors.New("Runtime history endpoint must be a canonical absolute OTLP metrics URL") + } + switch endpoint.Scheme { + case "https": + if config.Insecure { + return errors.New("Runtime history insecure transport requires an http endpoint") + } + case "http": + if !config.Insecure { + return errors.New("Runtime history http endpoint requires insecure=true") + } + default: + return errors.New("Runtime history endpoint scheme must be https or explicit insecure http") + } + if config.QueueCapacity < 0 || config.QueueCapacity > maxRuntimeHistoryQueueCapacity { + return errors.New("Runtime history queue_capacity is out of range") + } + if config.TimeoutSeconds < 0 || config.TimeoutSeconds > maxRuntimeHistoryTimeoutSeconds { + return errors.New("Runtime history timeout_seconds is out of range") + } + for key, value := range config.Headers { + lower := strings.ToLower(key) + if !httpguts.ValidHeaderFieldName(key) || !httpguts.ValidHeaderFieldValue(value) || lower == "host" || lower == "content-length" || lower == "content-type" || lower == "content-encoding" { + return errors.New("Runtime history headers contain an invalid or reserved entry") + } + } + return nil +} diff --git a/services/agents-api/cmd/server/runtime_history_test.go b/services/agents-api/cmd/server/runtime_history_test.go new file mode 100644 index 000000000..81ac79dea --- /dev/null +++ b/services/agents-api/cmd/server/runtime_history_test.go @@ -0,0 +1,88 @@ +package main + +import ( + "context" + "os" + "path/filepath" + "strings" + "testing" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" +) + +type panicHistoryExporter struct{} + +func (panicHistoryExporter) Export(context.Context, runtimeobs.ExportRecord) error { return nil } +func (panicHistoryExporter) Close(context.Context) error { panic("close") } + +func TestRuntimeHistoryIsDisabledByDefault(t *testing.T) { + t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", "") + option, exporter, err := runtimeHistory(t.Context()) + if err != nil || option != nil || exporter != nil { + t.Fatalf("disabled history created dependencies: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + } +} + +func TestRuntimeHistoryLoadsStrictServerOnlyConfig(t *testing.T) { + file := filepath.Join(t.TempDir(), "runtime-history.json") + config := `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Authorization":"Bearer test-only"},"queue_capacity":12,"timeout_seconds":3}` + if err := os.WriteFile(file, []byte(config), 0600); err != nil { + t.Fatal(err) + } + t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) + option, exporter, err := runtimeHistory(t.Context()) + if err != nil || option == nil || exporter == nil { + t.Fatalf("valid history config was rejected: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + } + closeCtx, cancel := context.WithTimeout(t.Context(), time.Second) + defer cancel() + if err := exporter.Close(closeCtx); err != nil { + t.Fatal(err) + } +} + +func TestRuntimeHistoryConfigFailsClosedWithoutLeakingSecrets(t *testing.T) { + tests := []struct { + name string + config string + }{ + {name: "unknown field", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","secret":"must-not-leak"}`}, + {name: "implicit insecure", config: `{"transport":"otlp_http","endpoint":"http://collector.example.test/v1/metrics"}`}, + {name: "userinfo", config: `{"transport":"otlp_http","endpoint":"https://user:must-not-leak@collector.example.test/v1/metrics"}`}, + {name: "header newline", config: "{\"transport\":\"otlp_http\",\"endpoint\":\"https://collector.example.test/v1/metrics\",\"headers\":{\"Authorization\":\"Bearer must-not-leak\\n\"}}"}, + {name: "reserved header", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Host":"must-not-leak"}}`}, + {name: "oversized queue", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","queue_capacity":4097}`}, + {name: "oversized timeout", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","timeout_seconds":31}`}, + } + for _, test := range tests { + t.Run(test.name, func(t *testing.T) { + file := filepath.Join(t.TempDir(), "runtime-history.json") + if err := os.WriteFile(file, []byte(test.config), 0600); err != nil { + t.Fatal(err) + } + t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) + option, exporter, err := runtimeHistory(t.Context()) + if err == nil || option != nil || exporter != nil { + t.Fatalf("unsafe history config was accepted: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + } + if strings.Contains(err.Error(), "must-not-leak") { + t.Fatalf("history error leaked config content: %v", err) + } + }) + } +} + +func TestRuntimeHistoryAllowsExplicitLocalHTTPCollector(t *testing.T) { + config := runtimeHistoryConfig{ + Transport: "otlp_http", Endpoint: "http://127.0.0.1:4318/v1/metrics", Insecure: true, + Headers: map[string]string{"X-Scope-OrgID": "operator-history"}, + } + if err := validateRuntimeHistoryConfig(config); err != nil { + t.Fatal(err) + } +} + +func TestRuntimeHistoryClosePanicIsIsolated(t *testing.T) { + closeRuntimeHistory(t.Context(), panicHistoryExporter{}) +} diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go new file mode 100644 index 000000000..ac32492a6 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go @@ -0,0 +1,274 @@ +// Package otlpexporter sends sanitized Runtime observation records to an +// operator-configured OTLP/HTTP metrics endpoint. It is optional history +// transport, not Runtime execution or lifecycle authority. +package otlpexporter + +import ( + "context" + "errors" + "math" + "regexp" + "time" + + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp" + "go.opentelemetry.io/otel/sdk/instrumentation" + sdkmetric "go.opentelemetry.io/otel/sdk/metric" + "go.opentelemetry.io/otel/sdk/metric/metricdata" + "go.opentelemetry.io/otel/sdk/resource" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" +) + +const scopeName = "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" + +const ( + CPUUsageName = "agents.runtime.cpu.usage" + CPUCapacityName = "agents.runtime.cpu.capacity" + MemoryUsageName = "agents.runtime.memory.usage" + MemoryLimitName = "agents.runtime.memory.limit" + SampleName = "agents.runtime.sample" + SampleDurationName = "agents.runtime.sample.duration" +) + +var safeLabelPattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,63}$`) + +type Config struct { + Endpoint string + Headers map[string]string + Insecure bool + RequestTimeout time.Duration +} + +type metricClient interface { + Export(context.Context, *metricdata.ResourceMetrics) error + Shutdown(context.Context) error +} + +type Exporter struct { + client metricClient + resource *resource.Resource +} + +func New(ctx context.Context, config Config) (*Exporter, error) { + if config.Endpoint == "" { + return nil, errors.New("Runtime history OTLP endpoint is required") + } + options := []otlpmetrichttp.Option{ + otlpmetrichttp.WithEndpointURL(config.Endpoint), + otlpmetrichttp.WithHeaders(config.Headers), + otlpmetrichttp.WithRetry(otlpmetrichttp.RetryConfig{Enabled: true}), + } + if config.Insecure { + options = append(options, otlpmetrichttp.WithInsecure()) + } + if config.RequestTimeout > 0 { + options = append(options, otlpmetrichttp.WithTimeout(config.RequestTimeout)) + } + client, err := otlpmetrichttp.New(ctx, options...) + if err != nil { + return nil, errors.New("cannot initialize Runtime history OTLP exporter") + } + return newWithClient(client), nil +} + +func newWithClient(client metricClient) *Exporter { + return &Exporter{ + client: client, + resource: resource.NewSchemaless( + attribute.String("service.name", "parsar-agents-api"), + attribute.String("service.namespace", "parsar-core"), + ), + } +} + +func (e *Exporter) Export(ctx context.Context, record runtimeobs.ExportRecord) error { + metrics, err := recordMetrics(record) + if err != nil { + return err + } + return e.client.Export(ctx, &metricdata.ResourceMetrics{ + Resource: e.resource, + ScopeMetrics: []metricdata.ScopeMetrics{{ + Scope: instrumentation.Scope{Name: scopeName}, + Metrics: metrics, + }}, + }) +} + +func (e *Exporter) Close(ctx context.Context) error { + return e.client.Shutdown(ctx) +} + +func recordMetrics(record runtimeobs.ExportRecord) ([]metricdata.Metrics, error) { + if err := validateRecord(record); err != nil { + return nil, err + } + attributes := recordAttributes(record) + result := []metricdata.Metrics{deltaCountMetric(SampleName, "Validated Runtime observation results", record.ResolvedAt, attributes)} + if record.SourceDuration > 0 { + result = append(result, durationMetric(record, attributes)) + } + if record.Sample == nil { + return result, nil + } + + sample := record.Sample + // Resource values are meaningful only within a provider-qualified compute + // incarnation. Keep the coverage result, but never let an unfenced sample + // collapse CPU or memory points across Runtime restarts. + if sample.StartedAt == nil { + return result, nil + } + if sample.CPUUsageSecondsTotal != nil { + result = append(result, cumulativeMetric( + CPUUsageName, "Cumulative CPU time consumed by the Runtime incarnation", "s", + *sample.CPUUsageSecondsTotal, sample.StartedAt, sample.ObservedAt, attributes, + )) + } + if sample.CPUCapacityCores != nil { + result = append(result, gaugeMetric(CPUCapacityName, "Configured Runtime CPU capacity", "{core}", *sample.CPUCapacityCores, sample.ObservedAt, attributes)) + } + if sample.MemoryUsageBytes != nil { + result = append(result, gaugeMetric(MemoryUsageName, "Current Runtime memory usage", "By", float64(*sample.MemoryUsageBytes), sample.ObservedAt, attributes)) + } + if sample.MemoryLimitBytes != nil { + result = append(result, gaugeMetric(MemoryLimitName, "Configured Runtime memory limit", "By", float64(*sample.MemoryLimitBytes), sample.ObservedAt, attributes)) + } + return result, nil +} + +func validateRecord(record runtimeobs.ExportRecord) error { + if record.TenantID == "" || record.SessionID == "" || record.ResolvedAt.IsZero() || record.ResolvedAt.Unix() < 0 { + return errors.New("invalid Runtime history export identity") + } + if record.ProviderType != "" && !safeLabelPattern.MatchString(record.ProviderType) { + return errors.New("invalid Runtime history provider type") + } + if record.Reason != "" && !safeLabelPattern.MatchString(record.Reason) { + return errors.New("invalid Runtime history result reason") + } + switch record.Mode { + case runtimeobs.ModeNone, runtimeobs.ModeSelfHosted, runtimeobs.ModeManaged: + default: + return errors.New("invalid Runtime history mode") + } + switch record.Status { + case runtimeobs.StatusObserved: + if record.Sample == nil { + return errors.New("observed Runtime history record has no sample") + } + case runtimeobs.StatusUnavailable, runtimeobs.StatusUnsupported: + if record.Sample != nil { + return errors.New("unobserved Runtime history record has a sample") + } + default: + return errors.New("invalid Runtime history status") + } + if record.SourceDuration < 0 { + return errors.New("invalid Runtime history source duration") + } + if record.SourceDuration > 0 && record.ResolvedAt.Add(-record.SourceDuration).Unix() < 0 { + return errors.New("invalid Runtime history source start time") + } + if record.Sample == nil { + return nil + } + return validateSample(record.Sample, record.ResolvedAt) +} + +func validateSample(sample *runtimeobs.Sample, resolvedAt time.Time) error { + if sample.ObservedAt.IsZero() || sample.ObservedAt.Unix() < 0 || sample.ObservedAt.After(resolvedAt) { + return errors.New("invalid Runtime history sample time") + } + if sample.StartedAt != nil && (sample.StartedAt.IsZero() || sample.StartedAt.Unix() < 0 || sample.StartedAt.After(sample.ObservedAt)) { + return errors.New("invalid Runtime history start time") + } + if sample.CPUUsageSecondsTotal != nil && (*sample.CPUUsageSecondsTotal < 0 || math.IsNaN(*sample.CPUUsageSecondsTotal) || math.IsInf(*sample.CPUUsageSecondsTotal, 0)) { + return errors.New("invalid Runtime history CPU usage") + } + if sample.CPUCapacityCores != nil && (*sample.CPUCapacityCores <= 0 || math.IsNaN(*sample.CPUCapacityCores) || math.IsInf(*sample.CPUCapacityCores, 0)) { + return errors.New("invalid Runtime history CPU capacity") + } + const maxSafeInteger = uint64(1<<53 - 1) + if sample.MemoryUsageBytes != nil && *sample.MemoryUsageBytes > maxSafeInteger { + return errors.New("invalid Runtime history memory usage") + } + if sample.MemoryLimitBytes != nil && (*sample.MemoryLimitBytes == 0 || *sample.MemoryLimitBytes > maxSafeInteger) { + return errors.New("invalid Runtime history memory limit") + } + return nil +} + +func recordAttributes(record runtimeobs.ExportRecord) attribute.Set { + values := []attribute.KeyValue{ + attribute.String("agents.tenant.id", record.TenantID), + attribute.String("agents.session.id", record.SessionID), + attribute.String("agents.runtime.mode", string(record.Mode)), + attribute.String("agents.runtime.status", string(record.Status)), + } + if record.EnvironmentID != "" { + values = append(values, attribute.String("agents.environment.id", record.EnvironmentID)) + } + if record.AllocationID != "" { + values = append(values, attribute.String("agents.runtime.allocation.id", record.AllocationID)) + } + if record.ProviderType != "" { + values = append(values, attribute.String("agents.runtime.provider.type", record.ProviderType)) + } + if record.Reason != "" { + values = append(values, attribute.String("agents.runtime.reason", record.Reason)) + } + if record.Sample != nil && record.Sample.StartedAt != nil { + values = append(values, attribute.Int64("agents.runtime.compute.started_at_unix_nano", record.Sample.StartedAt.UnixNano())) + } + return attribute.NewSet(values...) +} + +func deltaCountMetric(name, description string, observedAt time.Time, attributes attribute.Set) metricdata.Metrics { + return metricdata.Metrics{Name: name, Description: description, Unit: "{result}", Data: metricdata.Sum[int64]{ + DataPoints: []metricdata.DataPoint[int64]{{Attributes: attributes, StartTime: observedAt, Time: observedAt, Value: 1}}, + Temporality: metricdata.DeltaTemporality, IsMonotonic: true, + }} +} + +func durationMetric(record runtimeobs.ExportRecord, attributes attribute.Set) metricdata.Metrics { + duration := record.SourceDuration.Seconds() + bounds := []float64{0.01, 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10} + buckets := make([]uint64, len(bounds)+1) + index := len(bounds) + for i, bound := range bounds { + if duration <= bound { + index = i + break + } + } + buckets[index] = 1 + start := record.ResolvedAt.Add(-record.SourceDuration) + return metricdata.Metrics{Name: SampleDurationName, Description: "Provider Runtime sample duration", Unit: "s", Data: metricdata.Histogram[float64]{ + DataPoints: []metricdata.HistogramDataPoint[float64]{{ + Attributes: attributes, StartTime: start, Time: record.ResolvedAt, Count: 1, + Bounds: bounds, BucketCounts: buckets, Min: metricdata.NewExtrema(duration), Max: metricdata.NewExtrema(duration), Sum: duration, + }}, + Temporality: metricdata.DeltaTemporality, + }} +} + +func cumulativeMetric(name, description, unit string, value float64, startedAt *time.Time, observedAt time.Time, attributes attribute.Set) metricdata.Metrics { + point := metricdata.DataPoint[float64]{Attributes: attributes, Time: observedAt, Value: value} + if startedAt != nil { + point.StartTime = *startedAt + } + return metricdata.Metrics{Name: name, Description: description, Unit: unit, Data: metricdata.Sum[float64]{ + DataPoints: []metricdata.DataPoint[float64]{point}, Temporality: metricdata.CumulativeTemporality, IsMonotonic: true, + }} +} + +func gaugeMetric(name, description, unit string, value float64, observedAt time.Time, attributes attribute.Set) metricdata.Metrics { + return metricdata.Metrics{Name: name, Description: description, Unit: unit, Data: metricdata.Gauge[float64]{ + DataPoints: []metricdata.DataPoint[float64]{{Attributes: attributes, Time: observedAt, Value: value}}, + }} +} + +var _ sdkmetric.Exporter = (*otlpmetrichttp.Exporter)(nil) +var _ runtimeobs.Exporter = (*Exporter)(nil) diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go new file mode 100644 index 000000000..640f5cb8d --- /dev/null +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go @@ -0,0 +1,226 @@ +package otlpexporter + +import ( + "context" + "io" + "math" + "net/http" + "net/http/httptest" + "testing" + "time" + + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/sdk/metric/metricdata" + collectorv1 "go.opentelemetry.io/proto/otlp/collector/metrics/v1" + "google.golang.org/protobuf/proto" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" +) + +type captureClient struct { + metrics *metricdata.ResourceMetrics + calls int + closed bool +} + +func (c *captureClient) Export(_ context.Context, metrics *metricdata.ResourceMetrics) error { + c.calls++ + c.metrics = metrics + return nil +} + +func (c *captureClient) Shutdown(context.Context) error { + c.closed = true + return nil +} + +func observedRecord() runtimeobs.ExportRecord { + observedAt := time.Date(2026, 9, 23, 3, 0, 0, 0, time.UTC) + startedAt := observedAt.Add(-2 * time.Minute) + cpuUsage, capacity := 45.5, 2.0 + memoryUsage, memoryLimit := uint64(1024), uint64(2048) + return runtimeobs.ExportRecord{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", AllocationID: "allocation", + Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusObserved, + ResolvedAt: observedAt, SourceDuration: 50 * time.Millisecond, + Sample: &runtimeobs.Sample{ + ObservedAt: observedAt, StartedAt: &startedAt, + CPUUsageSecondsTotal: &cpuUsage, CPUCapacityCores: &capacity, + MemoryUsageBytes: &memoryUsage, MemoryLimitBytes: &memoryLimit, + }, + } +} + +func TestExporterBuildsFencedProviderNeutralMetrics(t *testing.T) { + client := &captureClient{} + exporter := newWithClient(client) + if err := exporter.Export(t.Context(), observedRecord()); err != nil { + t.Fatal(err) + } + if client.calls != 1 || client.metrics == nil || len(client.metrics.ScopeMetrics) != 1 { + t.Fatalf("unexpected export call: calls=%d metrics=%+v", client.calls, client.metrics) + } + serviceName, ok := client.metrics.Resource.Set().Value("service.name") + if !ok || serviceName.AsString() != "parsar-agents-api" { + t.Fatalf("unexpected service resource: %v %v", serviceName, ok) + } + + byName := map[string]metricdata.Metrics{} + for _, metric := range client.metrics.ScopeMetrics[0].Metrics { + byName[metric.Name] = metric + } + for _, name := range []string{SampleName, SampleDurationName, CPUUsageName, CPUCapacityName, MemoryUsageName, MemoryLimitName} { + if _, ok := byName[name]; !ok { + t.Fatalf("missing metric %q in %#v", name, byName) + } + } + cpu, ok := byName[CPUUsageName].Data.(metricdata.Sum[float64]) + if !ok || cpu.Temporality != metricdata.CumulativeTemporality || !cpu.IsMonotonic || len(cpu.DataPoints) != 1 || cpu.DataPoints[0].Value != 45.5 { + t.Fatalf("unexpected cumulative CPU metric: %#v", byName[CPUUsageName].Data) + } + attrs := attributeStrings(cpu.DataPoints[0].Attributes.ToSlice()) + for key, want := range map[string]string{ + "agents.tenant.id": "tenant", "agents.session.id": "session", "agents.environment.id": "environment", + "agents.runtime.allocation.id": "allocation", "agents.runtime.mode": "openai_hosted", + "agents.runtime.provider.type": "docker", "agents.runtime.status": "observed", + } { + if attrs[key] != want { + t.Fatalf("attribute %q = %q, want %q", key, attrs[key], want) + } + } + if _, ok := attrs["agents.runtime.compute.started_at_unix_nano"]; !ok { + t.Fatal("compute incarnation fence is missing") + } + if err := exporter.Close(t.Context()); err != nil || !client.closed { + t.Fatalf("exporter did not close: %v closed=%v", err, client.closed) + } +} + +func TestExporterPreservesUnavailableWithoutInventingResourceValues(t *testing.T) { + client := &captureClient{} + exporter := newWithClient(client) + record := runtimeobs.ExportRecord{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", AllocationID: "allocation", + Mode: runtimeobs.ModeManaged, Status: runtimeobs.StatusUnavailable, Reason: "sample_timeout", + ResolvedAt: time.Date(2026, 9, 23, 3, 0, 0, 0, time.UTC), + } + if err := exporter.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + metrics := client.metrics.ScopeMetrics[0].Metrics + if len(metrics) != 1 || metrics[0].Name != SampleName { + t.Fatalf("unavailable result invented measurements: %#v", metrics) + } +} + +func TestExporterPreservesObservedCoverageWithoutUnfencedResourceValues(t *testing.T) { + client := &captureClient{} + exporter := newWithClient(client) + record := observedRecord() + record.Sample.StartedAt = nil + + if err := exporter.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + metrics := client.metrics.ScopeMetrics[0].Metrics + if len(metrics) != 2 || metrics[0].Name != SampleName || metrics[1].Name != SampleDurationName { + t.Fatalf("unfenced sample exported resource measurements: %#v", metrics) + } + for _, metric := range metrics { + attributes := metricAttributes(metric) + if _, ok := attributes.Value("agents.runtime.compute.started_at_unix_nano"); ok { + t.Fatalf("unfenced sample invented a compute incarnation attribute: %#v", metric) + } + } +} + +func TestExporterRejectsUnsafeRecordsBeforeTransport(t *testing.T) { + client := &captureClient{} + exporter := newWithClient(client) + record := observedRecord() + record.ProviderType = "docker native=secret" + if err := exporter.Export(t.Context(), record); err == nil || client.calls != 0 { + t.Fatalf("unsafe record reached transport: err=%v calls=%d", err, client.calls) + } + record = observedRecord() + record.Sample = nil + if err := exporter.Export(t.Context(), record); err == nil || client.calls != 0 { + t.Fatalf("invalid observed record reached transport: err=%v calls=%d", err, client.calls) + } + record = observedRecord() + record.SourceDuration = time.Duration(math.MaxInt64) + if err := exporter.Export(t.Context(), record); err == nil || client.calls != 0 { + t.Fatalf("invalid source start reached transport: err=%v calls=%d", err, client.calls) + } +} + +func TestOTLPHTTPTransportSendsProtobufAndServerOnlyHeaders(t *testing.T) { + requests := make(chan *collectorv1.ExportMetricsServiceRequest, 1) + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, request *http.Request) { + if request.URL.Path != "/v1/metrics" { + t.Errorf("unexpected OTLP path %q", request.URL.Path) + } + if request.Header.Get("Authorization") != "Bearer test-only" { + t.Error("configured authorization header was not sent") + } + body, err := io.ReadAll(request.Body) + if err != nil { + t.Error(err) + w.WriteHeader(http.StatusBadRequest) + return + } + var decoded collectorv1.ExportMetricsServiceRequest + if err := proto.Unmarshal(body, &decoded); err != nil { + t.Error(err) + w.WriteHeader(http.StatusBadRequest) + return + } + requests <- &decoded + response, _ := proto.Marshal(&collectorv1.ExportMetricsServiceResponse{}) + w.Header().Set("Content-Type", "application/x-protobuf") + _, _ = w.Write(response) + })) + defer server.Close() + + exporter, err := New(t.Context(), Config{ + Endpoint: server.URL + "/v1/metrics", Insecure: true, + Headers: map[string]string{"Authorization": "Bearer test-only"}, RequestTimeout: time.Second, + }) + if err != nil { + t.Fatal(err) + } + defer func() { _ = exporter.Close(t.Context()) }() + if err := exporter.Export(t.Context(), observedRecord()); err != nil { + t.Fatal(err) + } + select { + case request := <-requests: + if len(request.ResourceMetrics) != 1 || len(request.ResourceMetrics[0].ScopeMetrics) != 1 { + t.Fatalf("unexpected OTLP request: %+v", request) + } + if got := len(request.ResourceMetrics[0].ScopeMetrics[0].Metrics); got != 6 { + t.Fatalf("expected 6 OTLP metrics, got %d", got) + } + case <-time.After(time.Second): + t.Fatal("OTLP request was not received") + } +} + +func attributeStrings(values []attribute.KeyValue) map[string]string { + result := make(map[string]string, len(values)) + for _, value := range values { + result[string(value.Key)] = value.Value.Emit() + } + return result +} + +func metricAttributes(metric metricdata.Metrics) attribute.Set { + switch data := metric.Data.(type) { + case metricdata.Sum[int64]: + return data.DataPoints[0].Attributes + case metricdata.Histogram[float64]: + return data.DataPoints[0].Attributes + default: + return attribute.NewSet() + } +} From a1513ffd7bb8968329188a50851040dfd39adfe3 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 04:30:27 +0800 Subject: [PATCH 19/51] Add singleton Runtime history sampling --- .env.example | 2 + .../runtime-observability-design.md | 51 ++- contracts/agents-api/runtime-observability.md | 15 +- services/agents-api/cmd/server/main.go | 40 ++- .../agents-api/cmd/server/runtime_history.go | 50 ++- .../cmd/server/runtime_history_test.go | 24 +- .../db/queries/runtime_allocations.sql | 12 + .../db/sqlc/runtime_allocations.sql.go | 43 +++ .../internal/runtimeobs/exporter.go | 34 +- .../runtimeobs/otlpexporter/exporter.go | 6 + .../runtimeobs/otlpexporter/exporter_test.go | 7 +- .../agents-api/internal/runtimeobs/sampler.go | 215 +++++++++++++ .../internal/runtimeobs/sampler_test.go | 304 ++++++++++++++++++ .../agents-api/internal/runtimeobs/service.go | 71 +++- .../internal/runtimeobs/service_test.go | 101 +++++- .../runtimeobs/storeresolver/resolver.go | 13 + .../runtimeobs/storeresolver/resolver_test.go | 19 ++ .../internal/store/runtime_allocations.go | 45 +++ .../store/runtime_observation_scan_test.go | 70 ++++ 19 files changed, 1035 insertions(+), 87 deletions(-) create mode 100644 services/agents-api/internal/runtimeobs/sampler.go create mode 100644 services/agents-api/internal/runtimeobs/sampler_test.go create mode 100644 services/agents-api/internal/store/runtime_observation_scan_test.go diff --git a/.env.example b/.env.example index eb23fc2ea..9661b8355 100644 --- a/.env.example +++ b/.env.example @@ -8,6 +8,8 @@ AGENTS_API_PROXY_TOKEN_FILE=~/.parsar/agents-api/web-token # Optional server-only OTLP/HTTP Runtime history exporter configuration. # When unset, current observations and the Dashboard require no telemetry backend. # AGENTS_API_RUNTIME_HISTORY_FILE=/absolute/path/runtime-history.json +# Set sample_interval_seconds (5..300) in that server-only file to enable the +# execution-owner singleton sampler. Omit it for on-read export only. # Alternative server-side value. Never use a VITE_ prefix, and do not set this # together with AGENTS_API_PROXY_TOKEN_FILE. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 88ae447d8..cc6e455f8 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -3,10 +3,11 @@ Status: provider abstraction with Docker and microsandbox sampling, the current-snapshot API/client contract, and the initial Core Web current-snapshot Dashboard are implemented. Phase 4 backend qualification, the bounded sanitized -exporter seam, and its optional OTLP/HTTP transport are implemented. No history -backend, history query API, or durable Web range is configured. Other provider -sources are not implemented. Microsandbox idle suspension is a separate durable -lifecycle feature; it does not consume this telemetry as authority. +exporter seam, its optional OTLP/HTTP transport, and execution-owner singleton +background sampling are implemented. No history backend, history query API, or +durable Web range is configured. Other provider sources are not implemented. +Microsandbox idle suspension is a separate durable lifecycle feature; it does +not consume this telemetry as authority. ## 1. Problem statement @@ -298,7 +299,8 @@ required. A minimal configuration is: "endpoint": "https://collector.example.com/v1/metrics", "headers": {"Authorization": "Bearer operator-managed-secret"}, "queue_capacity": 256, - "timeout_seconds": 2 + "timeout_seconds": 2, + "sample_interval_seconds": 30 } ``` @@ -307,7 +309,11 @@ committed. Plain HTTP requires the explicit combination of an `http` endpoint and `"insecure": true`; HTTPS rejects that flag. Endpoint userinfo, query strings, fragments, invalid headers, reserved transport headers, queues above 4096 records, and timeouts above 30 seconds fail startup without echoing config -contents. Export is best effort through the bounded queue documented above. +contents. `sample_interval_seconds` is optional; values from 5 through 300 enable +the deployment sampler, while omission retains on-read export only. Periodic +sampling requires the execution Worker because its database lease is the +deployment singleton boundary. Export is best effort through the bounded queue +documented above. The OTLP request uses standard protobuf metrics and these instruments: @@ -321,7 +327,8 @@ The OTLP request uses standard protobuf metrics and these instruments: | `agents.runtime.sample.duration` | delta histogram, seconds | bounded provider read duration | Core-owned tenant, Session, Environment, allocation, mode, provider type, -status, and safe reason are metric attributes. CPU, capacity, and memory points +status, safe reason, and collection source (`on_read` or `periodic`) are metric +attributes. CPU, capacity, and memory points are exported only when the sample also carries the compute `started_at` fence; that fence is included as an attribute on every such point. Provider keys, provider receipts, native container/pod/instance @@ -329,11 +336,22 @@ identifiers, raw errors, paths, and credentials are not attributes. Missing measurements produce no value point; they are represented only by the explicit sample status and reason. -The initial transport exports observations produced by the current read path; it -does not yet run a deployment-wide background sampler. Consequently, a future -history API must expose actual sample coverage and Core Web must not advertise a -durable range until the operator backend and a qualified collection cadence are -both configured. +When periodic sampling is enabled, the execution-owner service performs one +immediate, non-overlapping full keyset scan and repeats it after the configured +interval. The read-only scan covers nondeleted managed Sessions across tenants, +uses bounded pages and provider concurrency, gives each source an independent +deadline, and reuses the same resolver and observation service as current reads. +It never keeps, wakes, pauses, stops, or otherwise mutates compute. Failed rows +do not prevent later rows from being attempted, and a failed sweep is retried on +the next interval. The sampler monitors the execution lease during a sweep, +cancels in-flight provider reads on detected ownership loss, and rechecks the +lease before every periodic export handoff. `on_read` remains distinct from +`periodic`, so ad hoc API +traffic cannot be counted as qualified cadence coverage. + +A future history API must expose actual sample coverage. Core Web must not +advertise a durable range until the operator backend, query adapter, and a +qualified periodic collection cadence are all configured. ## 11. Dashboard information architecture @@ -489,12 +507,15 @@ Implemented for the browser-local current-snapshot live window. - Implemented: optional server-only OTLP/HTTP protobuf transport for the six documented Runtime instruments, including allocation and compute-incarnation fencing attributes. Configuration is strict and secrets never reach Web. +- Implemented: optional execution-owner singleton sampling across all nondeleted + managed Sessions. Keyset scans, provider concurrency, source deadlines, and + non-overlapping sweeps are bounded; collection source is exported explicitly. - Qualified: optional OTLP Collector fan-out with a separate high-cardinality history store and server-side tenant-scoped query adapter. ClickHouse is the first reference backend; no backend is a Core execution dependency. -- Not implemented: deployment-wide sampling cadence, `runtimehistory` query - adapter, public history extension, durable Web ranges, retention configuration, - and exporter queue/drop/error coverage telemetry. +- Not implemented: `runtimehistory` query adapter, public history extension, + durable Web ranges, retention configuration, and exporter queue/drop/error + coverage telemetry. - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. - Replace or extend the ephemeral live window with explicitly advertised durable diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index a8836f074..22c200d3f 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -89,11 +89,16 @@ errors, paths, and credentials are excluded from metric attributes. CPU, capacity, and memory points require the provider-qualified compute `started_at` fence; unfenced observations export coverage and read duration only. -This transport currently exports observations produced by the current read path; -it is not a deployment-wide background sampler. The Collector, -high-cardinality history backend, server-side history query adapter, collection -cadence, and durable Web ranges remain separate optional capabilities in the -[full design](runtime-observability-design.md). +The server-only file may also enable a bounded periodic cadence. Only the Core +service holding the execution database lease runs that deployment-wide sampler; +it scans nondeleted managed Sessions with bounded pages and concurrency, gives +each provider read an independent deadline, monitors lease ownership throughout +the sweep, rechecks ownership before export, and never mutates Runtime state. +Exports distinguish `periodic` samples from `on_read` samples. + +The Collector, high-cardinality history backend, server-side history query +adapter, retention policy, and durable Web ranges remain separate optional +capabilities in the [full design](runtime-observability-design.md). ## First-phase boundary diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 5cc190472..ae690646a 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -105,20 +105,20 @@ func run() error { if err != nil { return err } - historyOption, historyExporter, err := runtimeHistory(ctx) + history, err := runtimeHistory(ctx) if err != nil { return err } observationOptions := []runtimeobs.ServiceOption{} - if historyOption != nil { - observationOptions = append(observationOptions, historyOption) + if history.Option != nil { + observationOptions = append(observationOptions, history.Option) } observationService, err := runtimeobs.NewService(resolver, observationSources, observationOptions...) if err != nil { - if historyExporter != nil { + if history.Exporter != nil { closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() - closeRuntimeHistory(closeCtx, historyExporter) + closeRuntimeHistory(closeCtx, history.Exporter) } return err } @@ -126,8 +126,8 @@ func run() error { closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() _ = observationService.Close(closeCtx) - if historyExporter != nil { - closeRuntimeHistory(closeCtx, historyExporter) + if history.Exporter != nil { + closeRuntimeHistory(closeCtx, history.Exporter) } }() var workerDone chan error @@ -169,6 +169,32 @@ func run() error { options = append(options, api.WithHostedEnvironments()) } } + if history.SampleInterval > 0 { + if worker == nil { + return errors.New("Runtime history periodic sampling requires the execution worker") + } + sampler, err := runtimeobs.NewSampler(resolver, observationService, worker, runtimeobs.SamplerOptions{ + Interval: history.SampleInterval, + Report: func(result runtimeobs.SweepResult) { + fields := []any{"listed", result.Listed, "observed", result.Observed, "failed", result.Failed, "complete", result.Complete} + if result.Complete { + log.Bg().Debug("Runtime history sampling sweep complete", fields...) + } else { + log.Bg().Warn("Runtime history sampling sweep incomplete", fields...) + } + }, + }) + if err != nil { + return err + } + samplerCtx, cancelSampler := context.WithCancel(ctx) + samplerDone := make(chan error, 1) + go func() { samplerDone <- sampler.Run(samplerCtx) }() + defer func() { + cancelSampler() + <-samplerDone + }() + } handler, err := api.NewHandler(executionStore, auth, engine, options...) if err != nil { return err diff --git a/services/agents-api/cmd/server/runtime_history.go b/services/agents-api/cmd/server/runtime_history.go index 0b03adf4e..ebffe2766 100644 --- a/services/agents-api/cmd/server/runtime_history.go +++ b/services/agents-api/cmd/server/runtime_history.go @@ -18,19 +18,28 @@ import ( ) const ( - defaultRuntimeHistoryQueueCapacity = 256 - defaultRuntimeHistoryTimeoutSeconds = 2 - maxRuntimeHistoryQueueCapacity = 4096 - maxRuntimeHistoryTimeoutSeconds = 30 + defaultRuntimeHistoryQueueCapacity = 256 + defaultRuntimeHistoryTimeoutSeconds = 2 + maxRuntimeHistoryQueueCapacity = 4096 + maxRuntimeHistoryTimeoutSeconds = 30 + minRuntimeHistorySampleIntervalSeconds = 5 + maxRuntimeHistorySampleIntervalSeconds = 300 ) type runtimeHistoryConfig struct { - Transport string `json:"transport"` - Endpoint string `json:"endpoint"` - Insecure bool `json:"insecure"` - Headers map[string]string `json:"headers,omitempty"` - QueueCapacity int `json:"queue_capacity,omitempty"` - TimeoutSeconds int `json:"timeout_seconds,omitempty"` + Transport string `json:"transport"` + Endpoint string `json:"endpoint"` + Insecure bool `json:"insecure"` + Headers map[string]string `json:"headers,omitempty"` + QueueCapacity int `json:"queue_capacity,omitempty"` + TimeoutSeconds int `json:"timeout_seconds,omitempty"` + SampleIntervalSeconds int `json:"sample_interval_seconds,omitempty"` +} + +type runtimeHistorySetup struct { + Option runtimeobs.ServiceOption + Exporter runtimeHistoryExporter + SampleInterval time.Duration } type runtimeHistoryExporter interface { @@ -38,23 +47,23 @@ type runtimeHistoryExporter interface { Close(context.Context) error } -func runtimeHistory(ctx context.Context) (runtimeobs.ServiceOption, runtimeHistoryExporter, error) { +func runtimeHistory(ctx context.Context) (runtimeHistorySetup, error) { file := os.Getenv("AGENTS_API_RUNTIME_HISTORY_FILE") if file == "" { - return nil, nil, nil + return runtimeHistorySetup{}, nil } raw, err := os.ReadFile(file) if err != nil { - return nil, nil, errors.New("cannot read AGENTS_API_RUNTIME_HISTORY_FILE") + return runtimeHistorySetup{}, errors.New("cannot read AGENTS_API_RUNTIME_HISTORY_FILE") } var config runtimeHistoryConfig decoder := json.NewDecoder(bytes.NewReader(raw)) decoder.DisallowUnknownFields() if decoder.Decode(&config) != nil || decoder.Decode(new(any)) != io.EOF { - return nil, nil, errors.New("invalid Runtime history configuration") + return runtimeHistorySetup{}, errors.New("invalid Runtime history configuration") } if err := validateRuntimeHistoryConfig(config); err != nil { - return nil, nil, err + return runtimeHistorySetup{}, err } if config.QueueCapacity == 0 { config.QueueCapacity = defaultRuntimeHistoryQueueCapacity @@ -67,9 +76,13 @@ func runtimeHistory(ctx context.Context) (runtimeobs.ServiceOption, runtimeHisto Endpoint: config.Endpoint, Headers: config.Headers, Insecure: config.Insecure, RequestTimeout: timeout, }) if err != nil { - return nil, nil, err + return runtimeHistorySetup{}, err } - return runtimeobs.WithExporter(exporter, runtimeobs.ExportOptions{QueueCapacity: config.QueueCapacity, Timeout: timeout}), exporter, nil + return runtimeHistorySetup{ + Option: runtimeobs.WithExporter(exporter, runtimeobs.ExportOptions{QueueCapacity: config.QueueCapacity, Timeout: timeout}), + Exporter: exporter, + SampleInterval: time.Duration(config.SampleIntervalSeconds) * time.Second, + }, nil } func closeRuntimeHistory(ctx context.Context, exporter runtimeHistoryExporter) { @@ -105,6 +118,9 @@ func validateRuntimeHistoryConfig(config runtimeHistoryConfig) error { if config.TimeoutSeconds < 0 || config.TimeoutSeconds > maxRuntimeHistoryTimeoutSeconds { return errors.New("Runtime history timeout_seconds is out of range") } + if config.SampleIntervalSeconds != 0 && (config.SampleIntervalSeconds < minRuntimeHistorySampleIntervalSeconds || config.SampleIntervalSeconds > maxRuntimeHistorySampleIntervalSeconds) { + return errors.New("Runtime history sample_interval_seconds is out of range") + } for key, value := range config.Headers { lower := strings.ToLower(key) if !httpguts.ValidHeaderFieldName(key) || !httpguts.ValidHeaderFieldValue(value) || lower == "host" || lower == "content-length" || lower == "content-type" || lower == "content-encoding" { diff --git a/services/agents-api/cmd/server/runtime_history_test.go b/services/agents-api/cmd/server/runtime_history_test.go index 81ac79dea..e2382df5b 100644 --- a/services/agents-api/cmd/server/runtime_history_test.go +++ b/services/agents-api/cmd/server/runtime_history_test.go @@ -18,26 +18,26 @@ func (panicHistoryExporter) Close(context.Context) error func TestRuntimeHistoryIsDisabledByDefault(t *testing.T) { t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", "") - option, exporter, err := runtimeHistory(t.Context()) - if err != nil || option != nil || exporter != nil { - t.Fatalf("disabled history created dependencies: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + setup, err := runtimeHistory(t.Context()) + if err != nil || setup.Option != nil || setup.Exporter != nil || setup.SampleInterval != 0 { + t.Fatalf("disabled history created dependencies: option=%v exporter=%v interval=%v err=%v", setup.Option != nil, setup.Exporter != nil, setup.SampleInterval, err) } } func TestRuntimeHistoryLoadsStrictServerOnlyConfig(t *testing.T) { file := filepath.Join(t.TempDir(), "runtime-history.json") - config := `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Authorization":"Bearer test-only"},"queue_capacity":12,"timeout_seconds":3}` + config := `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Authorization":"Bearer test-only"},"queue_capacity":12,"timeout_seconds":3,"sample_interval_seconds":30}` if err := os.WriteFile(file, []byte(config), 0600); err != nil { t.Fatal(err) } t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) - option, exporter, err := runtimeHistory(t.Context()) - if err != nil || option == nil || exporter == nil { - t.Fatalf("valid history config was rejected: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + setup, err := runtimeHistory(t.Context()) + if err != nil || setup.Option == nil || setup.Exporter == nil || setup.SampleInterval != 30*time.Second { + t.Fatalf("valid history config was rejected: option=%v exporter=%v interval=%v err=%v", setup.Option != nil, setup.Exporter != nil, setup.SampleInterval, err) } closeCtx, cancel := context.WithTimeout(t.Context(), time.Second) defer cancel() - if err := exporter.Close(closeCtx); err != nil { + if err := setup.Exporter.Close(closeCtx); err != nil { t.Fatal(err) } } @@ -54,6 +54,8 @@ func TestRuntimeHistoryConfigFailsClosedWithoutLeakingSecrets(t *testing.T) { {name: "reserved header", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Host":"must-not-leak"}}`}, {name: "oversized queue", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","queue_capacity":4097}`}, {name: "oversized timeout", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","timeout_seconds":31}`}, + {name: "too frequent sampling", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","sample_interval_seconds":4}`}, + {name: "oversized sampling interval", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","sample_interval_seconds":301}`}, } for _, test := range tests { t.Run(test.name, func(t *testing.T) { @@ -62,9 +64,9 @@ func TestRuntimeHistoryConfigFailsClosedWithoutLeakingSecrets(t *testing.T) { t.Fatal(err) } t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) - option, exporter, err := runtimeHistory(t.Context()) - if err == nil || option != nil || exporter != nil { - t.Fatalf("unsafe history config was accepted: option=%v exporter=%v err=%v", option != nil, exporter != nil, err) + setup, err := runtimeHistory(t.Context()) + if err == nil || setup.Option != nil || setup.Exporter != nil { + t.Fatalf("unsafe history config was accepted: option=%v exporter=%v err=%v", setup.Option != nil, setup.Exporter != nil, err) } if strings.Contains(err.Error(), "must-not-leak") { t.Fatalf("history error leaked config content: %v", err) diff --git a/services/agents-api/internal/db/queries/runtime_allocations.sql b/services/agents-api/internal/db/queries/runtime_allocations.sql index 248be3e56..f66c0715f 100644 --- a/services/agents-api/internal/db/queries/runtime_allocations.sql +++ b/services/agents-api/internal/db/queries/runtime_allocations.sql @@ -17,6 +17,18 @@ JOIN sessions s ON s.id = e.session_id WHERE a.id > $1 AND a.state <> 'released' ORDER BY a.id LIMIT 32; +-- name: ListRuntimeObservationSessions :many +SELECT s.id, s.tenant_id +FROM sessions s +LEFT JOIN environments e ON e.session_id = s.id +LEFT JOIN runtime_allocations a ON a.environment_id = e.id +WHERE s.id > $1 + AND s.deleted_at IS NULL + AND s.configuration->'environment'->>'type' = 'openai_hosted' + AND (a.id IS NULL OR a.state <> 'released') +ORDER BY s.id +LIMIT $2; + -- name: ObserveRuntimeRunning :one UPDATE runtime_allocations SET state = 'running', create_settled = true WHERE id = $1 AND state IN ('creating', 'running') diff --git a/services/agents-api/internal/db/sqlc/runtime_allocations.sql.go b/services/agents-api/internal/db/sqlc/runtime_allocations.sql.go index b6ae080bf..2915eb6a3 100644 --- a/services/agents-api/internal/db/sqlc/runtime_allocations.sql.go +++ b/services/agents-api/internal/db/sqlc/runtime_allocations.sql.go @@ -190,6 +190,49 @@ func (q *Queries) ListRuntimeAllocations(ctx context.Context, id pgtype.UUID) ([ return items, nil } +const listRuntimeObservationSessions = `-- name: ListRuntimeObservationSessions :many +SELECT s.id, s.tenant_id +FROM sessions s +LEFT JOIN environments e ON e.session_id = s.id +LEFT JOIN runtime_allocations a ON a.environment_id = e.id +WHERE s.id > $1 + AND s.deleted_at IS NULL + AND s.configuration->'environment'->>'type' = 'openai_hosted' + AND (a.id IS NULL OR a.state <> 'released') +ORDER BY s.id +LIMIT $2 +` + +type ListRuntimeObservationSessionsParams struct { + ID pgtype.UUID `json:"id"` + Limit int32 `json:"limit"` +} + +type ListRuntimeObservationSessionsRow struct { + ID pgtype.UUID `json:"id"` + TenantID pgtype.UUID `json:"tenant_id"` +} + +func (q *Queries) ListRuntimeObservationSessions(ctx context.Context, arg ListRuntimeObservationSessionsParams) ([]ListRuntimeObservationSessionsRow, error) { + rows, err := q.db.Query(ctx, listRuntimeObservationSessions, arg.ID, arg.Limit) + if err != nil { + return nil, err + } + defer rows.Close() + items := []ListRuntimeObservationSessionsRow{} + for rows.Next() { + var i ListRuntimeObservationSessionsRow + if err := rows.Scan(&i.ID, &i.TenantID); err != nil { + return nil, err + } + items = append(items, i) + } + if err := rows.Err(); err != nil { + return nil, err + } + return items, nil +} + const listUnallocatedHostedEnvironments = `-- name: ListUnallocatedHostedEnvironments :many SELECT e.id, s.tenant_id FROM environments e JOIN sessions s ON s.id = e.session_id diff --git a/services/agents-api/internal/runtimeobs/exporter.go b/services/agents-api/internal/runtimeobs/exporter.go index bbc87f41d..b3aa54267 100644 --- a/services/agents-api/internal/runtimeobs/exporter.go +++ b/services/agents-api/internal/runtimeobs/exporter.go @@ -22,10 +22,11 @@ type ExportRecord struct { EnvironmentID string AllocationID string - Mode Mode - ProviderType string - Status Status - Reason string + Mode Mode + ProviderType string + Status Status + Reason string + CollectionSource CollectionSource ResolvedAt time.Time SourceDuration time.Duration @@ -151,19 +152,20 @@ func (d *exportDispatcher) close(ctx context.Context) error { } } -func exportRecord(observation Observation) ExportRecord { +func exportRecord(observation Observation, collectionSource CollectionSource) ExportRecord { return ExportRecord{ - TenantID: observation.Target.TenantID, - SessionID: observation.Target.SessionID, - EnvironmentID: observation.Target.EnvironmentID, - AllocationID: observation.Target.Instance.AllocationID, - Mode: observation.Target.Mode, - ProviderType: observation.ProviderType, - Status: observation.Status, - Reason: observation.Reason, - ResolvedAt: observation.ResolvedAt, - SourceDuration: observation.SourceDuration, - Sample: cloneSample(observation.Sample), + TenantID: observation.Target.TenantID, + SessionID: observation.Target.SessionID, + EnvironmentID: observation.Target.EnvironmentID, + AllocationID: observation.Target.Instance.AllocationID, + Mode: observation.Target.Mode, + ProviderType: observation.ProviderType, + Status: observation.Status, + Reason: observation.Reason, + CollectionSource: collectionSource, + ResolvedAt: observation.ResolvedAt, + SourceDuration: observation.SourceDuration, + Sample: cloneSample(observation.Sample), } } diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go index ac32492a6..5620471d2 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go @@ -165,6 +165,11 @@ func validateRecord(record runtimeobs.ExportRecord) error { default: return errors.New("invalid Runtime history status") } + switch record.CollectionSource { + case runtimeobs.CollectionSourceOnRead, runtimeobs.CollectionSourcePeriodic: + default: + return errors.New("invalid Runtime history collection source") + } if record.SourceDuration < 0 { return errors.New("invalid Runtime history source duration") } @@ -206,6 +211,7 @@ func recordAttributes(record runtimeobs.ExportRecord) attribute.Set { attribute.String("agents.session.id", record.SessionID), attribute.String("agents.runtime.mode", string(record.Mode)), attribute.String("agents.runtime.status", string(record.Status)), + attribute.String("agents.runtime.collection.source", string(record.CollectionSource)), } if record.EnvironmentID != "" { values = append(values, attribute.String("agents.environment.id", record.EnvironmentID)) diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go index 640f5cb8d..a346dabad 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go @@ -42,7 +42,8 @@ func observedRecord() runtimeobs.ExportRecord { return runtimeobs.ExportRecord{ TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", AllocationID: "allocation", Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusObserved, - ResolvedAt: observedAt, SourceDuration: 50 * time.Millisecond, + CollectionSource: runtimeobs.CollectionSourceOnRead, + ResolvedAt: observedAt, SourceDuration: 50 * time.Millisecond, Sample: &runtimeobs.Sample{ ObservedAt: observedAt, StartedAt: &startedAt, CPUUsageSecondsTotal: &cpuUsage, CPUCapacityCores: &capacity, @@ -83,6 +84,7 @@ func TestExporterBuildsFencedProviderNeutralMetrics(t *testing.T) { "agents.tenant.id": "tenant", "agents.session.id": "session", "agents.environment.id": "environment", "agents.runtime.allocation.id": "allocation", "agents.runtime.mode": "openai_hosted", "agents.runtime.provider.type": "docker", "agents.runtime.status": "observed", + "agents.runtime.collection.source": "on_read", } { if attrs[key] != want { t.Fatalf("attribute %q = %q, want %q", key, attrs[key], want) @@ -102,7 +104,8 @@ func TestExporterPreservesUnavailableWithoutInventingResourceValues(t *testing.T record := runtimeobs.ExportRecord{ TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", AllocationID: "allocation", Mode: runtimeobs.ModeManaged, Status: runtimeobs.StatusUnavailable, Reason: "sample_timeout", - ResolvedAt: time.Date(2026, 9, 23, 3, 0, 0, 0, time.UTC), + CollectionSource: runtimeobs.CollectionSourcePeriodic, + ResolvedAt: time.Date(2026, 9, 23, 3, 0, 0, 0, time.UTC), } if err := exporter.Export(t.Context(), record); err != nil { t.Fatal(err) diff --git a/services/agents-api/internal/runtimeobs/sampler.go b/services/agents-api/internal/runtimeobs/sampler.go new file mode 100644 index 000000000..fa41677c9 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/sampler.go @@ -0,0 +1,215 @@ +package runtimeobs + +import ( + "context" + "errors" + "sync" + "time" +) + +const ( + defaultSamplerPageSize = 32 + defaultSamplerConcurrency = 8 + defaultSamplerSourceTimeout = 2 * time.Second + samplerOwnershipPollInterval = 100 * time.Millisecond + historyOwnershipCheckTimeout = 250 * time.Millisecond +) + +type SessionIdentity struct { + TenantID, SessionID string +} + +type SessionPage struct { + Sessions []SessionIdentity + NextCursor string +} + +type SessionLister interface { + ListRuntimeObservationSessions(context.Context, string, int) (SessionPage, error) +} + +type HistoryObserver interface { + ObserveSessionForHistory(context.Context, string, string, OwnershipChecker, time.Duration) (Observation, error) +} + +type OwnershipChecker interface { + CheckOwnership(context.Context) error +} + +type SamplerOptions struct { + Interval time.Duration + PageSize int + Concurrency int + SourceTimeout time.Duration + Report func(SweepResult) +} + +// SweepResult is deliberately low-cardinality. It reports collection coverage +// without exposing tenant, Session, allocation, provider-native, or error text. +type SweepResult struct { + StartedAt, CompletedAt time.Time + Listed, Observed, Failed int + Complete bool +} + +type Sampler struct { + lister SessionLister + observer HistoryObserver + owner OwnershipChecker + options SamplerOptions + now func() time.Time + afterSweep func(SweepResult) +} + +func NewSampler(lister SessionLister, observer HistoryObserver, owner OwnershipChecker, options SamplerOptions) (*Sampler, error) { + if lister == nil || observer == nil || owner == nil { + return nil, errors.New("Runtime history sampler dependencies are required") + } + if options.Interval <= 0 { + return nil, errors.New("Runtime history sampler interval must be positive") + } + if options.PageSize == 0 { + options.PageSize = defaultSamplerPageSize + } + if options.PageSize < 1 || options.PageSize > 100 { + return nil, errors.New("Runtime history sampler page size must be 1..100") + } + if options.Concurrency == 0 { + options.Concurrency = defaultSamplerConcurrency + } + if options.Concurrency < 1 || options.Concurrency > 32 { + return nil, errors.New("Runtime history sampler concurrency must be 1..32") + } + if options.SourceTimeout == 0 { + options.SourceTimeout = defaultSamplerSourceTimeout + } + if options.SourceTimeout < time.Millisecond || options.SourceTimeout > 30*time.Second { + return nil, errors.New("Runtime history sampler source timeout is out of range") + } + return &Sampler{lister: lister, observer: observer, owner: owner, options: options, now: time.Now, afterSweep: options.Report}, nil +} + +// Run performs one immediate full keyset sweep and then repeats without overlap. +// A failed sweep is isolated from execution and retried at the next interval. +func (s *Sampler) Run(ctx context.Context) error { + for { + result := s.sweep(ctx) + s.report(result) + if err := ctx.Err(); err != nil { + return err + } + timer := time.NewTimer(s.options.Interval) + select { + case <-timer.C: + case <-ctx.Done(): + timer.Stop() + return ctx.Err() + } + } +} + +func (s *Sampler) report(result SweepResult) { + if s.afterSweep == nil { + return + } + defer func() { _ = recover() }() + s.afterSweep(result) +} + +func (s *Sampler) sweep(ctx context.Context) (result SweepResult) { + result.StartedAt = s.now() + defer func() { result.CompletedAt = s.now() }() + sweepCtx, cancel := context.WithCancel(ctx) + defer cancel() + if err := s.checkOwnership(sweepCtx); err != nil { + return result + } + go s.watchOwnership(sweepCtx, cancel) + cursor := "" + for { + if err := s.checkOwnership(sweepCtx); err != nil { + return result + } + page, err := s.lister.ListRuntimeObservationSessions(sweepCtx, cursor, s.options.PageSize) + if err != nil || len(page.Sessions) > s.options.PageSize || + (page.NextCursor != "" && (len(page.Sessions) == 0 || page.NextCursor == cursor || page.NextCursor != page.Sessions[len(page.Sessions)-1].SessionID)) { + return result + } + result.Listed += len(page.Sessions) + observed, failed := s.samplePage(sweepCtx, page.Sessions) + result.Observed += observed + result.Failed += failed + if sweepCtx.Err() != nil { + return result + } + if page.NextCursor == "" { + result.Complete = true + return result + } + cursor = page.NextCursor + } +} + +func (s *Sampler) watchOwnership(ctx context.Context, cancel context.CancelFunc) { + ticker := time.NewTicker(samplerOwnershipPollInterval) + defer ticker.Stop() + for { + select { + case <-ticker.C: + if err := s.checkOwnership(ctx); err != nil { + cancel() + return + } + case <-ctx.Done(): + return + } + } +} + +func (s *Sampler) checkOwnership(ctx context.Context) error { + checkCtx, cancel := context.WithTimeout(ctx, historyOwnershipCheckTimeout) + defer cancel() + return s.owner.CheckOwnership(checkCtx) +} + +func (s *Sampler) samplePage(ctx context.Context, sessions []SessionIdentity) (int, int) { + semaphore := make(chan struct{}, s.options.Concurrency) + var wait sync.WaitGroup + var mu sync.Mutex + observed, failed := 0, 0 + for _, session := range sessions { + if session.TenantID == "" || session.SessionID == "" { + failed++ + continue + } + wait.Add(1) + go func(session SessionIdentity) { + defer wait.Done() + select { + case semaphore <- struct{}{}: + defer func() { <-semaphore }() + case <-ctx.Done(): + mu.Lock() + failed++ + mu.Unlock() + return + } + if err := s.checkOwnership(ctx); err != nil { + mu.Lock() + failed++ + mu.Unlock() + return + } + _, err := s.observer.ObserveSessionForHistory(ctx, session.TenantID, session.SessionID, s.owner, s.options.SourceTimeout) + mu.Lock() + if err == nil { + observed++ + } else { + failed++ + } + mu.Unlock() + }(session) + } + wait.Wait() + return observed, failed +} diff --git a/services/agents-api/internal/runtimeobs/sampler_test.go b/services/agents-api/internal/runtimeobs/sampler_test.go new file mode 100644 index 000000000..3185cedb7 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/sampler_test.go @@ -0,0 +1,304 @@ +package runtimeobs + +import ( + "context" + "errors" + "sync" + "sync/atomic" + "testing" + "time" +) + +type samplerLister struct { + mu sync.Mutex + pages map[string]SessionPage + cursors []string +} + +type delayedSamplerResolver struct { + *samplerLister + delay time.Duration + target Target +} + +func (r delayedSamplerResolver) Resolve(ctx context.Context, tenant, session string) (Target, error) { + select { + case <-time.After(r.delay): + case <-ctx.Done(): + return Target{}, ctx.Err() + } + target := r.target + target.TenantID = tenant + target.SessionID = session + return target, nil +} + +func (l *samplerLister) ListRuntimeObservationSessions(_ context.Context, after string, _ int) (SessionPage, error) { + l.mu.Lock() + defer l.mu.Unlock() + l.cursors = append(l.cursors, after) + return l.pages[after], nil +} + +type samplerObserver struct { + mu sync.Mutex + sessions []SessionIdentity + active, max int + fail map[string]bool + wait <-chan struct{} +} + +func (o *samplerObserver) ObserveSessionForHistory(ctx context.Context, tenant, session string, _ OwnershipChecker, sourceTimeout time.Duration) (Observation, error) { + o.mu.Lock() + o.active++ + if o.active > o.max { + o.max = o.active + } + o.mu.Unlock() + if o.wait != nil { + sourceCtx, cancel := context.WithTimeout(ctx, sourceTimeout) + defer cancel() + select { + case <-o.wait: + case <-sourceCtx.Done(): + o.mu.Lock() + o.active-- + o.mu.Unlock() + return Observation{}, sourceCtx.Err() + } + } + o.mu.Lock() + defer o.mu.Unlock() + o.active-- + o.sessions = append(o.sessions, SessionIdentity{TenantID: tenant, SessionID: session}) + if o.fail[session] { + return Observation{}, errors.New("test failure") + } + return Observation{}, nil +} + +type samplerOwner struct{ err error } + +func (o samplerOwner) CheckOwnership(context.Context) error { return o.err } + +type sequenceOwner struct { + calls atomic.Int32 + failAfter int32 + lost atomic.Bool +} + +func (o *sequenceOwner) CheckOwnership(context.Context) error { + call := o.calls.Add(1) + if o.lost.Load() || (o.failAfter > 0 && call >= o.failAfter) { + return errors.New("ownership lost") + } + return nil +} + +func TestSamplerSweepsEveryPageAndIsolatesSessionFailures(t *testing.T) { + lister := &samplerLister{pages: map[string]SessionPage{ + "": {Sessions: []SessionIdentity{{TenantID: "tenant-a", SessionID: "session-a"}, {TenantID: "tenant-b", SessionID: "session-b"}}, NextCursor: "session-b"}, + "session-b": {Sessions: []SessionIdentity{{TenantID: "tenant-c", SessionID: "session-c"}}}, + }} + observer := &samplerObserver{fail: map[string]bool{"session-b": true}} + sampler, err := NewSampler(lister, observer, samplerOwner{}, SamplerOptions{Interval: time.Minute, PageSize: 2, Concurrency: 2}) + if err != nil { + t.Fatal(err) + } + now := time.Date(2026, 9, 23, 4, 0, 0, 0, time.UTC) + sampler.now = func() time.Time { now = now.Add(time.Second); return now } + result := sampler.sweep(t.Context()) + if !result.Complete || result.Listed != 3 || result.Observed != 2 || result.Failed != 1 { + t.Fatalf("unexpected sweep result: %+v", result) + } + if got := lister.cursors; len(got) != 2 || got[0] != "" || got[1] != "session-b" { + t.Fatalf("unexpected keyset scan: %#v", got) + } + if len(observer.sessions) != 3 { + t.Fatalf("not every Session was attempted: %#v", observer.sessions) + } +} + +func TestSamplerBoundsConcurrencyAndSourceDeadline(t *testing.T) { + blocked := make(chan struct{}) + observer := &samplerObserver{wait: blocked} + lister := &samplerLister{pages: map[string]SessionPage{ + "": {Sessions: []SessionIdentity{ + {TenantID: "t", SessionID: "1"}, {TenantID: "t", SessionID: "2"}, {TenantID: "t", SessionID: "3"}, + }}, + }} + sampler, err := NewSampler(lister, observer, samplerOwner{}, SamplerOptions{Interval: time.Minute, Concurrency: 2, SourceTimeout: 20 * time.Millisecond}) + if err != nil { + t.Fatal(err) + } + result := sampler.sweep(t.Context()) + if !result.Complete || result.Observed != 0 || result.Failed != 3 || observer.max != 2 { + t.Fatalf("unexpected bounded result: result=%+v max=%d", result, observer.max) + } +} + +func TestSamplerStopsBeforeListingWithoutOwnership(t *testing.T) { + lister := &samplerLister{pages: map[string]SessionPage{}} + sampler, err := NewSampler(lister, &samplerObserver{}, samplerOwner{err: errors.New("lost")}, SamplerOptions{Interval: time.Minute}) + if err != nil { + t.Fatal(err) + } + result := sampler.sweep(t.Context()) + if result.Complete || len(lister.cursors) != 0 { + t.Fatalf("sampler ran without deployment ownership: %+v %#v", result, lister.cursors) + } +} + +func TestSamplerRejectsInvalidContinuationWithoutLooping(t *testing.T) { + lister := &samplerLister{pages: map[string]SessionPage{"": { + Sessions: []SessionIdentity{{TenantID: "tenant", SessionID: "session"}}, + NextCursor: "different-session", + }}} + observer := &samplerObserver{} + sampler, err := NewSampler(lister, observer, samplerOwner{}, SamplerOptions{Interval: time.Minute}) + if err != nil { + t.Fatal(err) + } + result := sampler.sweep(t.Context()) + if result.Complete || result.Listed != 0 || len(observer.sessions) != 0 || len(lister.cursors) != 1 { + t.Fatalf("invalid continuation was accepted: result=%+v sessions=%#v cursors=%#v", result, observer.sessions, lister.cursors) + } +} + +func TestSamplerReportPanicIsIsolated(t *testing.T) { + sampler, err := NewSampler( + &samplerLister{pages: map[string]SessionPage{}}, + &samplerObserver{}, + samplerOwner{}, + SamplerOptions{Interval: time.Minute, Report: func(SweepResult) { panic("test") }}, + ) + if err != nil { + t.Fatal(err) + } + sampler.report(SweepResult{Complete: true}) +} + +func TestSamplerRunDoesNotOverlapSweepsAndStops(t *testing.T) { + started := make(chan struct{}, 1) + release := make(chan struct{}) + observer := &samplerObserver{wait: release} + lister := &samplerLister{pages: map[string]SessionPage{"": {Sessions: []SessionIdentity{{TenantID: "t", SessionID: "s"}}}}} + sampler, err := NewSampler(lister, observer, samplerOwner{}, SamplerOptions{Interval: time.Millisecond, SourceTimeout: time.Second}) + if err != nil { + t.Fatal(err) + } + sampler.afterSweep = func(SweepResult) { started <- struct{}{} } + ctx, cancel := context.WithCancel(t.Context()) + done := make(chan error, 1) + go func() { done <- sampler.Run(ctx) }() + deadline := time.After(time.Second) + for { + observer.mu.Lock() + active := observer.active + max := observer.max + observer.mu.Unlock() + if active == 1 { + if max != 1 { + t.Fatalf("overlapping sweep observed: max=%d", max) + } + break + } + select { + case <-deadline: + t.Fatal("sampler did not start") + default: + time.Sleep(time.Millisecond) + } + } + cancel() + select { + case err := <-done: + if !errors.Is(err, context.Canceled) { + t.Fatalf("unexpected sampler exit: %v", err) + } + case <-time.After(time.Second): + t.Fatal("sampler did not stop") + } +} + +func TestSamplerCancelsProviderReadWhenOwnershipIsLost(t *testing.T) { + release := make(chan struct{}) + observer := &samplerObserver{wait: release} + owner := &sequenceOwner{} + lister := &samplerLister{pages: map[string]SessionPage{"": {Sessions: []SessionIdentity{{TenantID: "t", SessionID: "s"}}}}} + sampler, err := NewSampler(lister, observer, owner, SamplerOptions{Interval: time.Minute, SourceTimeout: time.Second}) + if err != nil { + t.Fatal(err) + } + done := make(chan SweepResult, 1) + go func() { done <- sampler.sweep(t.Context()) }() + deadline := time.After(time.Second) + for { + observer.mu.Lock() + active := observer.active + observer.mu.Unlock() + if active == 1 { + break + } + select { + case <-deadline: + t.Fatal("provider read did not start") + default: + time.Sleep(time.Millisecond) + } + } + owner.lost.Store(true) + select { + case result := <-done: + if result.Complete || result.Observed != 0 || result.Failed != 1 { + t.Fatalf("lease-lost sweep was not fenced: %+v", result) + } + case <-time.After(time.Second): + t.Fatal("lease loss did not cancel provider read") + } +} + +func TestSamplerPreservesProviderTimeoutAndFinalFenceAfterSlowResolution(t *testing.T) { + lister := &samplerLister{pages: map[string]SessionPage{ + "": {Sessions: []SessionIdentity{{TenantID: "tenant", SessionID: "session"}}}, + }} + resolver := delayedSamplerResolver{ + samplerLister: lister, + delay: historyOwnershipCheckTimeout + 25*time.Millisecond, + target: Target{ + EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}, + }, + } + records := make(chan ExportRecord, 1) + service, err := NewService( + resolver, + map[string]Source{"provider": blockingSource{}}, + WithExporter(channelExporter{records: records}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + owner := &sequenceOwner{} + sampler, err := NewSampler(resolver, service, owner, SamplerOptions{ + Interval: time.Minute, SourceTimeout: 10 * time.Millisecond, + }) + if err != nil { + t.Fatal(err) + } + result := sampler.sweep(t.Context()) + if !result.Complete || result.Observed != 1 || result.Failed != 0 { + t.Fatalf("slow-resolution sweep lost the timeout observation: %+v", result) + } + select { + case record := <-records: + if record.Status != StatusUnavailable || record.Reason != "sample_timeout" || record.CollectionSource != CollectionSourcePeriodic { + t.Fatalf("slow-resolution timeout export mismatch: %+v", record) + } + case <-time.After(time.Second): + t.Fatal("timed out waiting for slow-resolution timeout export") + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} diff --git a/services/agents-api/internal/runtimeobs/service.go b/services/agents-api/internal/runtimeobs/service.go index 34cf222a8..b414fbc5f 100644 --- a/services/agents-api/internal/runtimeobs/service.go +++ b/services/agents-api/internal/runtimeobs/service.go @@ -20,6 +20,13 @@ type Observation struct { SourceDuration time.Duration } +type CollectionSource string + +const ( + CollectionSourceOnRead CollectionSource = "on_read" + CollectionSourcePeriodic CollectionSource = "periodic" +) + type Service struct { resolver TargetResolver sources map[string]Source @@ -63,21 +70,44 @@ func (s *Service) Close(ctx context.Context) error { return s.exports.close(ctx) } -func (s *Service) finish(observation Observation) Observation { +func (s *Service) finish(ctx context.Context, observation Observation, source CollectionSource, owner OwnershipChecker) (Observation, error) { + if err := checkHistoryOwnership(ctx, owner); err != nil { + return Observation{}, err + } if s.exports != nil { - s.exports.enqueue(exportRecord(observation)) + s.exports.enqueue(exportRecord(observation, source)) } - return observation + return observation, nil } func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string) (Observation, error) { + return s.observeSession(ctx, tenantID, sessionID, CollectionSourceOnRead, nil, 0) +} + +// ObserveSessionForHistory performs the same provider-neutral current read, but +// marks its export as deployment-periodic so coverage queries can distinguish it +// from user-triggered API reads. It has no lifecycle side effects. +func (s *Service) ObserveSessionForHistory(ctx context.Context, tenantID, sessionID string, owner OwnershipChecker, sourceTimeout time.Duration) (Observation, error) { + if owner == nil { + return Observation{}, errors.New("Runtime history observation ownership is required") + } + if sourceTimeout <= 0 || sourceTimeout > 30*time.Second { + return Observation{}, errors.New("Runtime history source timeout is out of range") + } + return s.observeSession(ctx, tenantID, sessionID, CollectionSourcePeriodic, owner, sourceTimeout) +} + +func (s *Service) observeSession(ctx context.Context, tenantID, sessionID string, collectionSource CollectionSource, owner OwnershipChecker, sourceTimeout time.Duration) (Observation, error) { + if err := checkHistoryOwnership(ctx, owner); err != nil { + return Observation{}, err + } target, err := s.resolver.Resolve(ctx, tenantID, sessionID) resolvedAt := s.now() if errors.Is(err, ErrUnavailable) { if target.TenantID != tenantID || target.SessionID != sessionID || target.Mode != ModeManaged || target.EnvironmentID == "" { return Observation{}, errors.New("Runtime observation resolver returned invalid pending allocation identity") } - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, collectionSource, owner) } if err != nil { return Observation{}, err @@ -90,7 +120,7 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string return Observation{}, errors.New("Runtime observation resolver returned mismatched Environment identity") } if target.Mode == ModeNone || target.Mode == ModeSelfHosted { - return s.finish(Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: resolvedAt}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnsupported, Reason: "runtime_mode_not_observable", ResolvedAt: resolvedAt}, collectionSource, owner) } if target.Mode != ModeManaged || target.Instance.AllocationID == "" || target.Instance.ProviderKey == "" { return Observation{}, errors.New("invalid managed Runtime observation target") @@ -101,16 +131,16 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string } switch target.Instance.AllocationState { case "creating": - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "allocation_pending", ResolvedAt: resolvedAt}, collectionSource, owner) case "cleanup_pending", "released": - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ResolvedAt: resolvedAt}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ResolvedAt: resolvedAt}, collectionSource, owner) case "running": default: return Observation{}, errors.New("invalid managed Runtime allocation state") } source, ok := s.sources[target.Instance.ProviderKey] if !ok { - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "source_not_configured", ResolvedAt: resolvedAt}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "source_not_configured", ResolvedAt: resolvedAt}, collectionSource, owner) } providerType := "" if typed, ok := source.(interface{ ObservationProviderType() string }); ok { @@ -120,16 +150,22 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string } } sourceStarted := time.Now() - sample, err := source.Observe(ctx, target) + sourceCtx := ctx + stopSource := func() {} + if sourceTimeout > 0 { + sourceCtx, stopSource = context.WithTimeout(ctx, sourceTimeout) + } + sample, err := source.Observe(sourceCtx, target) + stopSource() sourceDuration := time.Since(sourceStarted) if errors.Is(err, context.DeadlineExceeded) { - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "sample_timeout", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "sample_timeout", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}, collectionSource, owner) } if errors.Is(err, ErrNotRunning) { - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "runtime_not_running", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}, collectionSource, owner) } if errors.Is(err, ErrUnavailable) { - return s.finish(Observation{Target: target, Status: StatusUnavailable, Reason: "sample_unavailable", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusUnavailable, Reason: "sample_unavailable", ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}, collectionSource, owner) } if err != nil { return Observation{}, fmt.Errorf("observe Runtime: %w", err) @@ -137,5 +173,14 @@ func (s *Service) ObserveSession(ctx context.Context, tenantID, sessionID string if err := sample.validate(s.now()); err != nil { return Observation{}, err } - return s.finish(Observation{Target: target, Status: StatusObserved, Sample: &sample, ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}), nil + return s.finish(ctx, Observation{Target: target, Status: StatusObserved, Sample: &sample, ProviderType: providerType, ResolvedAt: s.now(), SourceDuration: sourceDuration}, collectionSource, owner) +} + +func checkHistoryOwnership(ctx context.Context, owner OwnershipChecker) error { + if owner == nil { + return nil + } + checkCtx, cancel := context.WithTimeout(ctx, historyOwnershipCheckTimeout) + defer cancel() + return owner.CheckOwnership(checkCtx) } diff --git a/services/agents-api/internal/runtimeobs/service_test.go b/services/agents-api/internal/runtimeobs/service_test.go index ab9ab3cbe..01f76ef9b 100644 --- a/services/agents-api/internal/runtimeobs/service_test.go +++ b/services/agents-api/internal/runtimeobs/service_test.go @@ -199,7 +199,7 @@ func TestServiceExportsOnlySanitizedValidatedRecords(t *testing.T) { if record.TenantID != "tenant" || record.SessionID != "session" || record.EnvironmentID != "environment" || record.AllocationID != "allocation" { t.Fatalf("exported identity mismatch: %+v", record) } - if record.ProviderType != "docker" || record.Mode != ModeManaged || record.Status != StatusObserved || record.Reason != "" { + if record.ProviderType != "docker" || record.Mode != ModeManaged || record.Status != StatusObserved || record.Reason != "" || record.CollectionSource != CollectionSourceOnRead { t.Fatalf("exported classification mismatch: %+v", record) } if record.Sample == nil || record.Sample.CPUUsageSecondsTotal == nil || *record.Sample.CPUUsageSecondsTotal != 12.5 || record.Sample.MemoryUsageBytes == nil || *record.Sample.MemoryUsageBytes != 1024 { @@ -213,6 +213,70 @@ func TestServiceExportsOnlySanitizedValidatedRecords(t *testing.T) { } } +func TestServiceMarksPeriodicHistoryCollection(t *testing.T) { + now := time.Date(2026, 9, 23, 4, 0, 0, 0, time.UTC) + startedAt := now.Add(-time.Minute) + target := Target{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}, + } + records := make(chan ExportRecord, 1) + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now, StartedAt: &startedAt}}}, + WithExporter(channelExporter{records: records}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + if _, err := service.ObserveSessionForHistory(t.Context(), "tenant", "session", samplerOwner{}, time.Second); err != nil { + t.Fatal(err) + } + select { + case record := <-records: + if record.CollectionSource != CollectionSourcePeriodic { + t.Fatalf("history collection source = %q", record.CollectionSource) + } + case <-time.After(time.Second): + t.Fatal("timed out waiting for periodic Runtime export") + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} + +func TestServiceDoesNotExportPeriodicSampleAfterOwnershipLoss(t *testing.T) { + now := time.Date(2026, 9, 23, 4, 0, 0, 0, time.UTC) + startedAt := now.Add(-time.Minute) + target := Target{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}, + } + records := make(chan ExportRecord, 1) + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now, StartedAt: &startedAt}}}, + WithExporter(channelExporter{records: records}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + owner := &sequenceOwner{failAfter: 2} + if _, err := service.ObserveSessionForHistory(t.Context(), "tenant", "session", owner, time.Second); err == nil { + t.Fatal("periodic sample crossed a lost execution lease") + } + select { + case record := <-records: + t.Fatalf("lease-lost periodic sample reached exporter: %+v", record) + default: + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} + func TestServiceRejectsUnsafeProviderTypeBeforeSamplingOrExport(t *testing.T) { now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} @@ -423,6 +487,41 @@ func TestServiceMapsAnActualSourceDeadlineWithoutLeakingIt(t *testing.T) { } } +func TestServiceExportsPeriodicSourceTimeoutAfterFinalOwnershipFence(t *testing.T) { + target := Target{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged, + Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}, + } + records := make(chan ExportRecord, 1) + service, err := NewService( + fixedResolver{target: target}, + map[string]Source{"provider": blockingSource{}}, + WithExporter(channelExporter{records: records}, ExportOptions{}), + ) + if err != nil { + t.Fatal(err) + } + owner := &sequenceOwner{} + observation, err := service.ObserveSessionForHistory(t.Context(), "tenant", "session", owner, 10*time.Millisecond) + if err != nil || observation.Status != StatusUnavailable || observation.Reason != "sample_timeout" { + t.Fatalf("periodic source timeout was not safely classified: %+v %v", observation, err) + } + if calls := owner.calls.Load(); calls != 2 { + t.Fatalf("ownership checks = %d, want entry and pre-export fences", calls) + } + select { + case record := <-records: + if record.Status != StatusUnavailable || record.Reason != "sample_timeout" || record.CollectionSource != CollectionSourcePeriodic { + t.Fatalf("periodic timeout export mismatch: %+v", record) + } + case <-time.After(time.Second): + t.Fatal("timed out waiting for periodic timeout export") + } + if err := service.Close(t.Context()); err != nil { + t.Fatal(err) + } +} + func TestServiceClassifiesResolverAndTerminalAllocationUnavailability(t *testing.T) { now := time.Date(2026, 9, 22, 1, 0, 0, 0, time.UTC) service, err := NewService(fixedResolver{target: Target{SessionID: "session", EnvironmentID: "environment", Mode: ModeManaged}, err: ErrUnavailable}, nil) diff --git a/services/agents-api/internal/runtimeobs/storeresolver/resolver.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go index b09bff1bd..79a988a49 100644 --- a/services/agents-api/internal/runtimeobs/storeresolver/resolver.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go @@ -13,6 +13,7 @@ import ( type sessionStore interface { GetSession(context.Context, string, string) (store.Session, error) GetRuntimeAllocation(context.Context, string, string) (store.RuntimeAllocation, error) + ListRuntimeObservationSessions(context.Context, string, int) (store.RuntimeObservationSessionPage, error) } type Resolver struct{ store sessionStore } @@ -81,3 +82,15 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (run return runtimeobs.Target{}, errors.New("invalid stored Runtime environment type") } } + +func (r *Resolver) ListRuntimeObservationSessions(ctx context.Context, after string, limit int) (runtimeobs.SessionPage, error) { + page, err := r.store.ListRuntimeObservationSessions(ctx, after, limit) + if err != nil { + return runtimeobs.SessionPage{}, err + } + result := runtimeobs.SessionPage{Sessions: make([]runtimeobs.SessionIdentity, 0, len(page.Sessions)), NextCursor: page.NextCursor} + for _, session := range page.Sessions { + result.Sessions = append(result.Sessions, runtimeobs.SessionIdentity{TenantID: session.TenantID, SessionID: session.SessionID}) + } + return result, nil +} diff --git a/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go index 7a6ba56a6..9b86e304a 100644 --- a/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go @@ -13,6 +13,7 @@ type resolverStore struct { session store.Session allocation store.RuntimeAllocation allocationErr error + page store.RuntimeObservationSessionPage } func (s resolverStore) GetSession(context.Context, string, string) (store.Session, error) { @@ -23,6 +24,24 @@ func (s resolverStore) GetRuntimeAllocation(context.Context, string, string) (st return s.allocation, s.allocationErr } +func (s resolverStore) ListRuntimeObservationSessions(context.Context, string, int) (store.RuntimeObservationSessionPage, error) { + return s.page, nil +} + +func TestResolverListsOnlyProviderNeutralSessionIdentity(t *testing.T) { + r, err := NewResolver(resolverStore{page: store.RuntimeObservationSessionPage{ + Sessions: []store.RuntimeObservationSession{{TenantID: "tenant", SessionID: "session"}}, + NextCursor: "session", + }}) + if err != nil { + t.Fatal(err) + } + page, err := r.ListRuntimeObservationSessions(t.Context(), "", 32) + if err != nil || len(page.Sessions) != 1 || page.Sessions[0].TenantID != "tenant" || page.Sessions[0].SessionID != "session" || page.NextCursor != "session" { + t.Fatalf("unexpected observation scan identity: %+v %v", page, err) + } +} + func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { r, err := NewResolver(resolverStore{ session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, diff --git a/services/agents-api/internal/store/runtime_allocations.go b/services/agents-api/internal/store/runtime_allocations.go index dd643d28e..5b84bc51b 100644 --- a/services/agents-api/internal/store/runtime_allocations.go +++ b/services/agents-api/internal/store/runtime_allocations.go @@ -29,6 +29,18 @@ type RuntimeAllocation struct { CreatedAt, KeptAt time.Time } +// RuntimeObservationSession is the minimum durable Core identity needed by the +// deployment-wide read-only sampler. Provider identity is resolved again by the +// observation service before any external read. +type RuntimeObservationSession struct { + TenantID, SessionID string +} + +type RuntimeObservationSessionPage struct { + Sessions []RuntimeObservationSession + NextCursor string +} + // ReserveRuntimeAllocation commits the allocation and dedicated device together // before external Create. Only a fresh receipt authorizes that one Create call. func (s *Store) ReserveRuntimeAllocation(ctx context.Context, tenant, environment, providerKey, credentialHash string) (RuntimeAllocation, error) { @@ -139,6 +151,39 @@ func (s *Store) ListRuntimeAllocations(ctx context.Context, after string) ([]Run return result, nil } +// ListRuntimeObservationSessions performs a deployment-wide, read-only keyset +// scan of live managed Session identities. It excludes deleted Sessions and +// released allocations; it does not acquire, renew, or mutate Runtime state. +func (s *Store) ListRuntimeObservationSessions(ctx context.Context, after string, limit int) (RuntimeObservationSessionPage, error) { + if limit < 1 || limit > 100 { + return RuntimeObservationSessionPage{}, ErrInvalidInput + } + id := pgtype.UUID{Valid: true} + if after != "" { + var err error + id, err = parseID(after) + if err != nil { + return RuntimeObservationSessionPage{}, err + } + } + rows, err := s.queries.ListRuntimeObservationSessions(ctx, sqlc.ListRuntimeObservationSessionsParams{ID: id, Limit: int32(limit + 1)}) + if err != nil { + return RuntimeObservationSessionPage{}, err + } + page := RuntimeObservationSessionPage{Sessions: make([]RuntimeObservationSession, 0, min(limit, len(rows)))} + if len(rows) > limit { + page.NextCursor = uuid.UUID(rows[limit-1].ID.Bytes).String() + rows = rows[:limit] + } + for _, row := range rows { + page.Sessions = append(page.Sessions, RuntimeObservationSession{ + TenantID: uuid.UUID(row.TenantID.Bytes).String(), + SessionID: uuid.UUID(row.ID.Bytes).String(), + }) + } + return page, nil +} + func runtimeAllocationFromRow(row sqlc.RuntimeAllocation, session, tenant pgtype.UUID, deleted pgtype.Timestamptz, expired bool) RuntimeAllocation { var retainedUntil *time.Time if row.ComputeRetainedUntil.Valid { diff --git a/services/agents-api/internal/store/runtime_observation_scan_test.go b/services/agents-api/internal/store/runtime_observation_scan_test.go new file mode 100644 index 000000000..a392b45ac --- /dev/null +++ b/services/agents-api/internal/store/runtime_observation_scan_test.go @@ -0,0 +1,70 @@ +package store_test + +import ( + "slices" + "testing" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" +) + +func TestRuntimeObservationScanIsDeploymentWideBoundedAndExcludesDeleted(t *testing.T) { + s, _ := store.NewManagedTestStore(t) + var expected []string + for range 5 { + _, session, _ := managedSession(t, s) + expected = append(expected, session.ID) + } + slices.Sort(expected) + if err := s.DeleteSession(t.Context(), sessionTenant(t, s, expected[2]), expected[2]); err != nil { + t.Fatal(err) + } + expected = append(expected[:2], expected[3:]...) + + var got []string + cursor := "" + for { + page, err := s.ListRuntimeObservationSessions(t.Context(), cursor, 2) + if err != nil { + t.Fatal(err) + } + if len(page.Sessions) > 2 { + t.Fatalf("unbounded observation page: %+v", page) + } + for _, session := range page.Sessions { + if session.TenantID == "" || session.SessionID == "" { + t.Fatalf("incomplete observation identity: %+v", session) + } + got = append(got, session.SessionID) + } + if page.NextCursor == "" { + break + } + if len(page.Sessions) == 0 || page.NextCursor != page.Sessions[len(page.Sessions)-1].SessionID { + t.Fatalf("invalid observation cursor: %+v", page) + } + cursor = page.NextCursor + } + if !slices.Equal(got, expected) { + t.Fatalf("observation scan = %v, want %v", got, expected) + } + if _, err := s.ListRuntimeObservationSessions(t.Context(), "", 0); err == nil { + t.Fatal("zero observation page size was accepted") + } +} + +func sessionTenant(t *testing.T, s *store.Store, sessionID string) string { + t.Helper() + // The deployment-wide scan intentionally discovers tenant identity without + // enumerating configured API keys. Use that same read to locate this fixture. + page, err := s.ListRuntimeObservationSessions(t.Context(), "", 100) + if err != nil { + t.Fatal(err) + } + for _, session := range page.Sessions { + if session.SessionID == sessionID { + return session.TenantID + } + } + t.Fatalf("Session %s was not listed", sessionID) + return "" +} From 7169eabb17b6df44a1133a09fe3e6539febd7584 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 04:55:29 +0800 Subject: [PATCH 20/51] Add bounded Runtime history query seam --- .../runtime-observability-design.md | 42 +++- contracts/agents-api/runtime-observability.md | 9 +- .../internal/runtimehistory/service.go | 126 ++++++++++ .../internal/runtimehistory/service_test.go | 196 +++++++++++++++ .../runtimehistory/storeresolver/resolver.go | 48 ++++ .../storeresolver/resolver_test.go | 74 ++++++ .../internal/runtimehistory/types.go | 228 ++++++++++++++++++ .../runtimeobs/otlpexporter/exporter.go | 4 + .../runtimeobs/otlpexporter/exporter_test.go | 4 +- 9 files changed, 724 insertions(+), 7 deletions(-) create mode 100644 services/agents-api/internal/runtimehistory/service.go create mode 100644 services/agents-api/internal/runtimehistory/service_test.go create mode 100644 services/agents-api/internal/runtimehistory/storeresolver/resolver.go create mode 100644 services/agents-api/internal/runtimehistory/storeresolver/resolver_test.go create mode 100644 services/agents-api/internal/runtimehistory/types.go diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index cc6e455f8..e88b0811f 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -327,8 +327,10 @@ The OTLP request uses standard protobuf metrics and these instruments: | `agents.runtime.sample.duration` | delta histogram, seconds | bounded provider read duration | Core-owned tenant, Session, Environment, allocation, mode, provider type, -status, safe reason, and collection source (`on_read` or `periodic`) are metric -attributes. CPU, capacity, and memory points +status, safe reason, collection source (`on_read` or `periodic`), and Core +resolved/observed timestamps are metric attributes. The explicit nanosecond +timestamps preserve the record join key when a backend's generic OTLP tables +store metric event time at lower precision. CPU, capacity, and memory points are exported only when the sample also carries the compute `started_at` fence; that fence is included as an attribute on every such point. Provider keys, provider receipts, native container/pod/instance @@ -353,6 +355,33 @@ A future history API must expose actual sample coverage. Core Web must not advertise a durable range until the operator backend, query adapter, and a qualified periodic collection cadence are all configured. +### 10.3 Backend-neutral history query boundary + +`services/agents-api/internal/runtimehistory` defines the server-side query +contract independently from ClickHouse, OTLP, and the public HTTP shape. Its +service resolves the authenticated tenant and Session to durable Core identity +before calling a Reader. Reader queries always carry tenant, Session, and +Environment scope plus a bounded start, exclusive end, server-selected step, +and total point budget. Provider-native identity is never a query input. + +Reader results remain divided by allocation and compute `started_at` fence. +Every bucket reports explicit observation coverage and nullable CPU/memory +values. CPU utilization may be derived only from ordered cumulative counters +inside one fence; memory uses the final observed value in the bucket. Empty +buckets remain gaps. The service rejects cross-scope rows, duplicate series, +overlapping or out-of-range buckets, unsafe provider labels, invalid numeric +values, and results exceeding the total point budget. + +Capabilities contain only safe backend-neutral limits: collection mode, +qualified sample interval, retention, minimum step, maximum range, point budget, +and supported metrics. A configured Reader without qualified periodic sampling +is not sufficient to advertise a Durable Dashboard source. Backend identity, +URLs, credentials, and tenant data are never capability fields. + +This internal boundary is implemented, but no production Reader or public +history route is configured yet. The next qualification adds the ClickHouse +reference Reader, then a Session-scoped public extension and strict client. + ## 11. Dashboard information architecture ### 11.1 Overview @@ -510,12 +539,15 @@ Implemented for the browser-local current-snapshot live window. - Implemented: optional execution-owner singleton sampling across all nondeleted managed Sessions. Keyset scans, provider concurrency, source deadlines, and non-overlapping sweeps are bounded; collection source is exported explicitly. +- Implemented: backend-neutral `runtimehistory` types and service validation. + Tenant/Session/Environment scope precedes every Reader query; incarnation, + coverage, nullability, ordering, range and total-point invariants are enforced. - Qualified: optional OTLP Collector fan-out with a separate high-cardinality history store and server-side tenant-scoped query adapter. ClickHouse is the first reference backend; no backend is a Core execution dependency. -- Not implemented: `runtimehistory` query adapter, public history extension, - durable Web ranges, retention configuration, and exporter queue/drop/error - coverage telemetry. +- Not implemented: a production `runtimehistory` Reader, public history + extension, durable Web ranges, retention configuration, and exporter + queue/drop/error coverage telemetry. - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. - Replace or extend the ephemeral live window with explicitly advertised durable diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 22c200d3f..519f8c40b 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -87,7 +87,9 @@ ownership, or lifecycle authority. The file may contain endpoint authorization headers and is never returned to Web. Provider receipts, native identifiers, raw errors, paths, and credentials are excluded from metric attributes. CPU, capacity, and memory points require the provider-qualified compute `started_at` -fence; unfenced observations export coverage and read duration only. +fence; unfenced observations export coverage and read duration only. Core +resolved/observed nanosecond attributes provide a backend join key even when +generic OTLP storage lowers the event timestamp precision. The server-only file may also enable a bounded periodic cadence. Only the Core service holding the execution database lease runs that deployment-wide sampler; @@ -100,6 +102,11 @@ The Collector, high-cardinality history backend, server-side history query adapter, retention policy, and durable Web ranges remain separate optional capabilities in the [full design](runtime-observability-design.md). +The internal `runtimehistory` boundary is backend-neutral and validates Core +scope, incarnation fences, bucket coverage, nullability, time bounds and total +point limits. It does not make history available by itself: no production +Reader or public history route is configured in this phase. + ## First-phase boundary The current-snapshot implementation adds no migration, metrics backend, token diff --git a/services/agents-api/internal/runtimehistory/service.go b/services/agents-api/internal/runtimehistory/service.go new file mode 100644 index 000000000..1ec714702 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/service.go @@ -0,0 +1,126 @@ +package runtimehistory + +import ( + "context" + "errors" + "time" +) + +var ErrUnsupported = errors.New("Runtime history is unsupported for this Session") + +type ScopeResolver interface { + ResolveRuntimeHistoryScope(context.Context, string, string) (Scope, error) +} + +type Reader interface { + Capabilities() Capabilities + Query(context.Context, Query) (Result, error) +} + +type Service struct { + resolver ScopeResolver + reader Reader + capabilities Capabilities + now func() time.Time +} + +func NewService(resolver ScopeResolver, reader Reader) (*Service, error) { + if resolver == nil || reader == nil { + return nil, errors.New("Runtime history resolver and reader are required") + } + capabilities := reader.Capabilities() + if err := capabilities.validate(); err != nil { + return nil, err + } + return &Service{resolver: resolver, reader: reader, capabilities: cloneCapabilities(capabilities), now: time.Now}, nil +} + +func (s *Service) Capabilities() Capabilities { + return cloneCapabilities(s.capabilities) +} + +// QuerySession resolves authenticated Core ownership before issuing any backend +// request. The history Reader receives only Core identity and bounded range data; +// provider-native identity is never an authority-bearing input. +func (s *Service) QuerySession(ctx context.Context, tenantID, sessionID string, requested Range) (Response, error) { + now := s.now().UTC() + if requested.Start.IsZero() || requested.End.IsZero() || !requested.End.After(requested.Start) || requested.End.After(now.Add(time.Second)) || requested.End.Sub(requested.Start) > s.capabilities.MaximumRange || requested.MaxPoints < 2 || requested.MaxPoints > s.capabilities.MaximumPoints { + return Response{}, errors.New("invalid Runtime history range") + } + scope, err := s.resolver.ResolveRuntimeHistoryScope(ctx, tenantID, sessionID) + if err != nil { + return Response{}, err + } + if scope.TenantID != tenantID || scope.SessionID != sessionID { + return Response{}, errors.New("Runtime history resolver returned mismatched identity") + } + if err := scope.validate(); err != nil { + return Response{}, err + } + step := resolution(requested.End.Sub(requested.Start), requested.MaxPoints, s.capabilities.MinimumStep) + query := Query{Scope: scope, Start: requested.Start.UTC(), End: requested.End.UTC(), Step: step, MaxPoints: requested.MaxPoints} + result, err := s.reader.Query(ctx, query) + if err != nil { + return Response{}, err + } + if err := validateResult(query, result, now); err != nil { + return Response{}, err + } + return Response{ + Capabilities: cloneCapabilities(s.capabilities), + Scope: scope, + Requested: Range{Start: query.Start, End: query.End, MaxPoints: query.MaxPoints}, + Resolution: query.Step, + GeneratedAt: result.GeneratedAt.UTC(), + RetainedFrom: cloneTime(result.RetainedFrom), + Series: cloneSeries(result.Series), + }, nil +} + +func resolution(duration time.Duration, maxPoints int, minimum time.Duration) time.Duration { + step := (duration + time.Duration(maxPoints) - 1) / time.Duration(maxPoints) + if step < minimum { + return minimum + } + return ((step + time.Second - 1) / time.Second) * time.Second +} + +func cloneTime(value *time.Time) *time.Time { + if value == nil { + return nil + } + cloned := *value + return &cloned +} + +func cloneSeries(values []Series) []Series { + result := make([]Series, len(values)) + for index, value := range values { + result[index] = value + result[index].Points = append([]Point(nil), value.Points...) + for pointIndex := range result[index].Points { + point := &result[index].Points[pointIndex] + point.CPUUtilizationRatio = cloneFloat(point.CPUUtilizationRatio) + point.CPUCapacityCores = cloneFloat(point.CPUCapacityCores) + point.MemoryUsageBytes = cloneUint64(point.MemoryUsageBytes) + point.MemoryLimitBytes = cloneUint64(point.MemoryLimitBytes) + } + } + return result +} + +func cloneFloat(value *float64) *float64 { + if value == nil { + return nil + } + cloned := *value + return &cloned +} + +func cloneUint64(value *uint64) *uint64 { + if value == nil { + return nil + } + cloned := *value + return &cloned +} diff --git a/services/agents-api/internal/runtimehistory/service_test.go b/services/agents-api/internal/runtimehistory/service_test.go new file mode 100644 index 000000000..14202ff75 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/service_test.go @@ -0,0 +1,196 @@ +package runtimehistory + +import ( + "context" + "errors" + "testing" + "time" +) + +const ( + tenantID = "11111111-1111-4111-8111-111111111111" + sessionID = "22222222-2222-4222-8222-222222222222" + environmentID = "33333333-3333-4333-8333-333333333333" + allocationID = "44444444-4444-4444-8444-444444444444" +) + +type fixedScopeResolver struct { + scope Scope + err error +} + +func (r fixedScopeResolver) ResolveRuntimeHistoryScope(context.Context, string, string) (Scope, error) { + return r.scope, r.err +} + +type fakeReader struct { + capabilities Capabilities + result Result + err error + queries []Query +} + +func (r *fakeReader) Capabilities() Capabilities { return r.capabilities } + +func (r *fakeReader) Query(_ context.Context, query Query) (Result, error) { + r.queries = append(r.queries, query) + return r.result, r.err +} + +func capabilities() Capabilities { + return Capabilities{ + CollectionMode: CollectionPeriodic, + SampleInterval: 30 * time.Second, + Retention: 7 * 24 * time.Hour, + MinimumStep: 30 * time.Second, + MaximumRange: 24 * time.Hour, + MaximumPoints: 1_000, + Metrics: []Metric{MetricCPU, MetricMemory}, + } +} + +func TestServiceAuthorizesAndBoundsBackendQuery(t *testing.T) { + now := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + start := now.Add(-time.Hour) + startedAt := start.Add(-time.Minute) + ratio, capacity := .25, 2.0 + memory, limit := uint64(1024), uint64(2048) + scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} + reader := &fakeReader{capabilities: capabilities()} + reader.result = Result{GeneratedAt: now, RetainedFrom: &start, Series: []Series{{ + Scope: scope, AllocationID: allocationID, StartedAt: startedAt, ProviderType: "docker", + Points: []Point{{ + Start: start, End: start.Add(30 * time.Second), FirstObservedAt: start.Add(time.Second), LastObservedAt: start.Add(20 * time.Second), + ObservationCount: 2, ObservedCount: 2, CPUContributorCount: 2, MemoryContributorCount: 2, + CPUUtilizationRatio: &ratio, CPUCapacityCores: &capacity, MemoryUsageBytes: &memory, MemoryLimitBytes: &limit, + }}, + }}} + service, err := NewService(fixedScopeResolver{scope: scope}, reader) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + response, err := service.QuerySession(t.Context(), tenantID, sessionID, Range{Start: start, End: now, MaxPoints: 60}) + if err != nil { + t.Fatal(err) + } + if len(reader.queries) != 1 || reader.queries[0].Scope != scope || reader.queries[0].Step != time.Minute || response.Resolution != time.Minute || !response.Durable() { + t.Fatalf("unexpected bounded history query: query=%+v response=%+v", reader.queries, response) + } + ratio = .9 + if response.Series[0].Points[0].CPUUtilizationRatio == nil || *response.Series[0].Points[0].CPUUtilizationRatio != .25 { + t.Fatal("response aliases backend-owned metric memory") + } +} + +func TestServiceNeverQueriesBeforeOwnershipResolution(t *testing.T) { + reader := &fakeReader{capabilities: capabilities()} + denied := errors.New("not found") + service, err := NewService(fixedScopeResolver{err: denied}, reader) + if err != nil { + t.Fatal(err) + } + now := time.Now().UTC() + _, err = service.QuerySession(t.Context(), tenantID, sessionID, Range{Start: now.Add(-time.Hour), End: now, MaxPoints: 60}) + if !errors.Is(err, denied) || len(reader.queries) != 0 { + t.Fatalf("unauthorized history reached reader: err=%v queries=%+v", err, reader.queries) + } +} + +func TestServiceRejectsInvalidRangeAndResolverIdentity(t *testing.T) { + reader := &fakeReader{capabilities: capabilities()} + scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} + service, err := NewService(fixedScopeResolver{scope: scope}, reader) + if err != nil { + t.Fatal(err) + } + now := time.Now().UTC() + for _, requested := range []Range{ + {Start: now, End: now, MaxPoints: 60}, + {Start: now.Add(-25 * time.Hour), End: now, MaxPoints: 60}, + {Start: now.Add(-time.Hour), End: now, MaxPoints: 1}, + {Start: now.Add(-time.Hour), End: now, MaxPoints: 1_001}, + } { + if _, err := service.QuerySession(t.Context(), tenantID, sessionID, requested); err == nil { + t.Fatalf("invalid range accepted: %+v", requested) + } + } + if len(reader.queries) != 0 { + t.Fatalf("invalid range reached reader: %+v", reader.queries) + } + service.resolver = fixedScopeResolver{scope: Scope{TenantID: "55555555-5555-4555-8555-555555555555", SessionID: sessionID, EnvironmentID: environmentID}} + if _, err := service.QuerySession(t.Context(), tenantID, sessionID, Range{Start: now.Add(-time.Hour), End: now, MaxPoints: 60}); err == nil || len(reader.queries) != 0 { + t.Fatal("mismatched resolver identity reached reader") + } +} + +func TestServiceRejectsMalformedBackendResults(t *testing.T) { + now := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + start := now.Add(-time.Hour) + scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} + base := Series{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker"} + for name, mutate := range map[string]func(*Result){ + "cross tenant": func(result *Result) { result.Series[0].TenantID = "55555555-5555-4555-8555-555555555555" }, + "duplicate series": func(result *Result) { result.Series = append(result.Series, result.Series[0]) }, + "out of range": func(result *Result) { result.Series[0].Points[0].End = now.Add(time.Minute) }, + "zero coverage value": func(result *Result) { value := 1.0; result.Series[0].Points[0].CPUUtilizationRatio = &value }, + "exclusive end": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount, point.ObservedCount = 1, 1 + point.FirstObservedAt, point.LastObservedAt = point.Start, point.End + point.CPUContributorCount = 1 + value := 1.0 + point.CPUCapacityCores = &value + }, + "zero capacity": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount, point.ObservedCount, point.CPUContributorCount = 1, 1, 1 + point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start + value := 0.0 + point.CPUCapacityCores = &value + }, + "zero memory limit": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount, point.ObservedCount, point.MemoryContributorCount = 1, 1, 1 + point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start + value := uint64(0) + point.MemoryLimitBytes = &value + }, + "coverage without value": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount, point.ObservedCount, point.CPUContributorCount = 1, 1, 1 + point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start + }, + } { + t.Run(name, func(t *testing.T) { + result := Result{GeneratedAt: now, Series: []Series{base}} + result.Series[0].Points = []Point{{Start: start, End: start.Add(time.Minute)}} + mutate(&result) + reader := &fakeReader{capabilities: capabilities(), result: result} + service, err := NewService(fixedScopeResolver{scope: scope}, reader) + if err != nil { + t.Fatal(err) + } + service.now = func() time.Time { return now } + if _, err := service.QuerySession(t.Context(), tenantID, sessionID, Range{Start: start, End: now, MaxPoints: 60}); err == nil { + t.Fatal("malformed backend result accepted") + } + }) + } +} + +func TestServiceRejectsUnsafeCapabilities(t *testing.T) { + base := capabilities() + for _, mutate := range []func(*Capabilities){ + func(value *Capabilities) { value.CollectionMode = CollectionOnRead }, + func(value *Capabilities) { value.MaximumRange = value.Retention + time.Second }, + func(value *Capabilities) { value.Metrics = []Metric{MetricCPU, MetricCPU} }, + } { + value := base + value.Metrics = append([]Metric(nil), base.Metrics...) + mutate(&value) + if _, err := NewService(fixedScopeResolver{}, &fakeReader{capabilities: value}); err == nil { + t.Fatalf("unsafe capabilities accepted: %+v", value) + } + } +} diff --git a/services/agents-api/internal/runtimehistory/storeresolver/resolver.go b/services/agents-api/internal/runtimehistory/storeresolver/resolver.go new file mode 100644 index 000000000..e18ee3912 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/storeresolver/resolver.go @@ -0,0 +1,48 @@ +package storeresolver + +import ( + "context" + "encoding/json" + "errors" + "fmt" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" +) + +type environmentStore interface { + GetSessionEnvironment(context.Context, string, string) (store.Environment, error) +} + +type Resolver struct{ store environmentStore } + +func NewResolver(value environmentStore) (*Resolver, error) { + if value == nil { + return nil, errors.New("Runtime history store is required") + } + return &Resolver{store: value}, nil +} + +// ResolveRuntimeHistoryScope authorizes the Session through the Core store and +// returns only durable Core identity. It intentionally does not resolve a +// current allocation: retained history may contain earlier allocations or +// multiple compute incarnations that the Reader must keep separately fenced. +func (r *Resolver) ResolveRuntimeHistoryScope(ctx context.Context, tenantID, sessionID string) (runtimehistory.Scope, error) { + environment, err := r.store.GetSessionEnvironment(ctx, tenantID, sessionID) + if err != nil { + return runtimehistory.Scope{}, fmt.Errorf("resolve Runtime history Environment: %w", err) + } + if environment.TenantID != tenantID || environment.SessionID != sessionID { + return runtimehistory.Scope{}, errors.New("Runtime history Environment does not match resolved ownership") + } + var configuration struct { + Type string `json:"type"` + } + if json.Unmarshal(environment.Configuration, &configuration) != nil || configuration.Type == "" { + return runtimehistory.Scope{}, errors.New("invalid stored Runtime history environment configuration") + } + if configuration.Type != "openai_hosted" { + return runtimehistory.Scope{}, runtimehistory.ErrUnsupported + } + return runtimehistory.Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environment.ID}, nil +} diff --git a/services/agents-api/internal/runtimehistory/storeresolver/resolver_test.go b/services/agents-api/internal/runtimehistory/storeresolver/resolver_test.go new file mode 100644 index 000000000..0b4533348 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/storeresolver/resolver_test.go @@ -0,0 +1,74 @@ +package storeresolver + +import ( + "context" + "errors" + "testing" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" +) + +type resolverStore struct { + environment store.Environment + err error + calls int +} + +func (s *resolverStore) GetSessionEnvironment(context.Context, string, string) (store.Environment, error) { + s.calls++ + return s.environment, s.err +} + +func TestResolverAuthorizesManagedSessionWithoutSelectingCurrentAllocation(t *testing.T) { + backend := &resolverStore{environment: store.Environment{ + ID: environmentID, TenantID: tenantID, SessionID: sessionID, + Configuration: []byte(`{"type":"openai_hosted"}`), + }} + resolver, err := NewResolver(backend) + if err != nil { + t.Fatal(err) + } + scope, err := resolver.ResolveRuntimeHistoryScope(t.Context(), tenantID, sessionID) + if err != nil { + t.Fatal(err) + } + if scope != (runtimehistory.Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID}) || backend.calls != 1 { + t.Fatalf("unexpected Runtime history scope: %+v calls=%d", scope, backend.calls) + } +} + +func TestResolverPreservesTenantScopedNotFound(t *testing.T) { + backend := &resolverStore{err: store.ErrNotFound} + resolver, err := NewResolver(backend) + if err != nil { + t.Fatal(err) + } + _, err = resolver.ResolveRuntimeHistoryScope(t.Context(), tenantID, sessionID) + if !errors.Is(err, store.ErrNotFound) { + t.Fatalf("tenant-scoped not found was not preserved: %v", err) + } +} + +func TestResolverRejectsUnsupportedOrMismatchedEnvironment(t *testing.T) { + for _, environment := range []store.Environment{ + {ID: environmentID, TenantID: tenantID, SessionID: sessionID, Configuration: []byte(`{"type":"self_hosted"}`)}, + {ID: environmentID, TenantID: "55555555-5555-4555-8555-555555555555", SessionID: sessionID, Configuration: []byte(`{"type":"openai_hosted"}`)}, + {ID: environmentID, TenantID: tenantID, SessionID: "66666666-6666-4666-8666-666666666666", Configuration: []byte(`{"type":"openai_hosted"}`)}, + {ID: environmentID, TenantID: tenantID, SessionID: sessionID, Configuration: []byte(`{"type":`)}, + } { + resolver, err := NewResolver(&resolverStore{environment: environment}) + if err != nil { + t.Fatal(err) + } + if _, err := resolver.ResolveRuntimeHistoryScope(t.Context(), tenantID, sessionID); err == nil { + t.Fatalf("unsafe Runtime history Environment accepted: %+v", environment) + } + } +} + +const ( + tenantID = "11111111-1111-4111-8111-111111111111" + sessionID = "22222222-2222-4222-8222-222222222222" + environmentID = "33333333-3333-4333-8333-333333333333" +) diff --git a/services/agents-api/internal/runtimehistory/types.go b/services/agents-api/internal/runtimehistory/types.go new file mode 100644 index 000000000..2bdfb8868 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/types.go @@ -0,0 +1,228 @@ +package runtimehistory + +import ( + "errors" + "math" + "regexp" + "slices" + "time" + + "github.com/google/uuid" +) + +type CollectionMode string + +const ( + CollectionOnRead CollectionMode = "on_read" + CollectionPeriodic CollectionMode = "periodic" +) + +type Metric string + +const ( + MetricCPU Metric = "cpu" + MetricMemory Metric = "memory" +) + +var providerTypePattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) + +// Capabilities contains only backend-neutral facts safe to expose through a +// future public capability response. Backend names, endpoints, credentials and +// tenant identity are deliberately absent. +type Capabilities struct { + CollectionMode CollectionMode + SampleInterval time.Duration + Retention time.Duration + MinimumStep time.Duration + MaximumRange time.Duration + MaximumPoints int + Metrics []Metric +} + +func (c Capabilities) Durable() bool { + return c.CollectionMode == CollectionPeriodic && c.SampleInterval > 0 +} + +func (c Capabilities) validate() error { + switch c.CollectionMode { + case CollectionOnRead: + if c.SampleInterval != 0 { + return errors.New("on-read Runtime history cannot declare a sampling interval") + } + case CollectionPeriodic: + if c.SampleInterval <= 0 { + return errors.New("periodic Runtime history requires a sampling interval") + } + default: + return errors.New("invalid Runtime history collection mode") + } + if c.Retention <= 0 || c.MinimumStep <= 0 || c.MaximumRange <= 0 || c.MaximumRange > c.Retention { + return errors.New("invalid Runtime history time bounds") + } + if c.MaximumPoints < 2 || c.MaximumPoints > 10_000 { + return errors.New("invalid Runtime history point limit") + } + if len(c.Metrics) == 0 || len(c.Metrics) > 2 { + return errors.New("invalid Runtime history metrics") + } + seen := map[Metric]bool{} + for _, metric := range c.Metrics { + if metric != MetricCPU && metric != MetricMemory || seen[metric] { + return errors.New("invalid Runtime history metrics") + } + seen[metric] = true + } + return nil +} + +type Scope struct { + TenantID, SessionID, EnvironmentID string +} + +func (s Scope) validate() error { + for _, value := range []string{s.TenantID, s.SessionID, s.EnvironmentID} { + if uuid.Validate(value) != nil { + return errors.New("invalid Runtime history scope") + } + } + return nil +} + +type Range struct { + Start, End time.Time + MaxPoints int +} + +type Query struct { + Scope + Start, End time.Time + Step time.Duration + MaxPoints int +} + +// Result is returned by a backend Reader before the service validates identity, +// ordering, bounds and values. Empty retained ranges and missing metric values +// remain explicit; a Reader must never synthesize zeroes for absent samples. +type Result struct { + GeneratedAt time.Time + RetainedFrom *time.Time + Series []Series +} + +type Series struct { + Scope + AllocationID string + StartedAt time.Time + ProviderType string + Points []Point +} + +// Point represents one server-selected bucket. CPUUtilizationRatio is derived +// only from ordered cumulative counters within this Series' allocation and +// StartedAt fence. Memory values are the final observed values in the bucket. +type Point struct { + Start, End time.Time + FirstObservedAt, LastObservedAt time.Time + ObservationCount int + ObservedCount, UnavailableCount int + CPUContributorCount int + MemoryContributorCount int + CPUUtilizationRatio, CPUCapacityCores *float64 + MemoryUsageBytes, MemoryLimitBytes *uint64 +} + +type Response struct { + Capabilities + Scope + Requested Range + Resolution time.Duration + GeneratedAt time.Time + RetainedFrom *time.Time + Series []Series +} + +func validateResult(query Query, result Result, now time.Time) error { + if result.GeneratedAt.IsZero() || result.GeneratedAt.After(now.Add(time.Second)) { + return errors.New("invalid Runtime history generation time") + } + if result.RetainedFrom != nil && (result.RetainedFrom.IsZero() || result.RetainedFrom.After(query.End)) { + return errors.New("invalid Runtime history retention boundary") + } + seriesKeys := map[string]bool{} + totalPoints := 0 + if len(result.Series) > query.MaxPoints { + return errors.New("Runtime history result exceeds series limit") + } + for index := range result.Series { + series := &result.Series[index] + if series.Scope != query.Scope || uuid.Validate(series.AllocationID) != nil || series.StartedAt.IsZero() || !series.StartedAt.Before(query.End) || !providerTypePattern.MatchString(series.ProviderType) { + return errors.New("invalid Runtime history series identity") + } + key := series.AllocationID + "\x00" + series.StartedAt.UTC().Format(time.RFC3339Nano) + if seriesKeys[key] { + return errors.New("duplicate Runtime history series") + } + seriesKeys[key] = true + totalPoints += len(series.Points) + if totalPoints > query.MaxPoints { + return errors.New("Runtime history result exceeds point limit") + } + for pointIndex := range series.Points { + point := series.Points[pointIndex] + if err := validatePoint(query, series.StartedAt, point); err != nil { + return err + } + if pointIndex > 0 && series.Points[pointIndex-1].End.After(point.Start) { + return errors.New("Runtime history points overlap or are out of order") + } + } + } + return nil +} + +func validatePoint(query Query, startedAt time.Time, point Point) error { + if point.Start.Before(query.Start) || !point.End.After(point.Start) || point.End.After(query.End) || point.End.Sub(point.Start) > query.Step || point.End.Before(startedAt) { + return errors.New("invalid Runtime history point bounds") + } + if point.ObservationCount < 0 || point.ObservedCount < 0 || point.UnavailableCount < 0 || point.ObservedCount+point.UnavailableCount != point.ObservationCount || + point.CPUContributorCount < 0 || point.CPUContributorCount > point.ObservedCount || point.MemoryContributorCount < 0 || point.MemoryContributorCount > point.ObservedCount { + return errors.New("invalid Runtime history point coverage") + } + if point.ObservationCount == 0 { + if !point.FirstObservedAt.IsZero() || !point.LastObservedAt.IsZero() || point.CPUUtilizationRatio != nil || point.CPUCapacityCores != nil || point.MemoryUsageBytes != nil || point.MemoryLimitBytes != nil { + return errors.New("empty Runtime history bucket contains observations") + } + return nil + } + if point.FirstObservedAt.Before(point.Start) || point.FirstObservedAt.Before(startedAt) || point.LastObservedAt.Before(point.FirstObservedAt) || !point.LastObservedAt.Before(point.End) { + return errors.New("invalid Runtime history observation bounds") + } + if value := point.CPUUtilizationRatio; value != nil && (math.IsNaN(*value) || math.IsInf(*value, 0) || *value < 0) { + return errors.New("invalid Runtime history CPU utilization") + } + if value := point.CPUCapacityCores; value != nil && (math.IsNaN(*value) || math.IsInf(*value, 0) || *value <= 0) { + return errors.New("invalid Runtime history CPU capacity") + } + if point.CPUContributorCount == 0 && (point.CPUUtilizationRatio != nil || point.CPUCapacityCores != nil) { + return errors.New("Runtime history CPU value lacks coverage") + } + if point.CPUContributorCount > 0 && point.CPUUtilizationRatio == nil && point.CPUCapacityCores == nil { + return errors.New("Runtime history CPU coverage lacks a value") + } + const maxSafeInteger = uint64(1<<53 - 1) + if point.MemoryUsageBytes != nil && *point.MemoryUsageBytes > maxSafeInteger || point.MemoryLimitBytes != nil && (*point.MemoryLimitBytes == 0 || *point.MemoryLimitBytes > maxSafeInteger) { + return errors.New("invalid Runtime history memory value") + } + if point.MemoryContributorCount == 0 && (point.MemoryUsageBytes != nil || point.MemoryLimitBytes != nil) { + return errors.New("Runtime history memory value lacks coverage") + } + if point.MemoryContributorCount > 0 && point.MemoryUsageBytes == nil && point.MemoryLimitBytes == nil { + return errors.New("Runtime history memory coverage lacks a value") + } + return nil +} + +func cloneCapabilities(value Capabilities) Capabilities { + value.Metrics = slices.Clone(value.Metrics) + return value +} diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go index 5620471d2..1e173c10e 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go @@ -212,6 +212,7 @@ func recordAttributes(record runtimeobs.ExportRecord) attribute.Set { attribute.String("agents.runtime.mode", string(record.Mode)), attribute.String("agents.runtime.status", string(record.Status)), attribute.String("agents.runtime.collection.source", string(record.CollectionSource)), + attribute.Int64("agents.runtime.resolved_at_unix_nano", record.ResolvedAt.UnixNano()), } if record.EnvironmentID != "" { values = append(values, attribute.String("agents.environment.id", record.EnvironmentID)) @@ -228,6 +229,9 @@ func recordAttributes(record runtimeobs.ExportRecord) attribute.Set { if record.Sample != nil && record.Sample.StartedAt != nil { values = append(values, attribute.Int64("agents.runtime.compute.started_at_unix_nano", record.Sample.StartedAt.UnixNano())) } + if record.Sample != nil { + values = append(values, attribute.Int64("agents.runtime.observed_at_unix_nano", record.Sample.ObservedAt.UnixNano())) + } return attribute.NewSet(values...) } diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go index a346dabad..1f59f010f 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go @@ -84,7 +84,9 @@ func TestExporterBuildsFencedProviderNeutralMetrics(t *testing.T) { "agents.tenant.id": "tenant", "agents.session.id": "session", "agents.environment.id": "environment", "agents.runtime.allocation.id": "allocation", "agents.runtime.mode": "openai_hosted", "agents.runtime.provider.type": "docker", "agents.runtime.status": "observed", - "agents.runtime.collection.source": "on_read", + "agents.runtime.collection.source": "on_read", + "agents.runtime.resolved_at_unix_nano": "1790132400000000000", + "agents.runtime.observed_at_unix_nano": "1790132400000000000", } { if attrs[key] != want { t.Fatalf("attribute %q = %q, want %q", key, attrs[key], want) From 2d4808361b7c072366cd66335925f273b418c93f Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 06:56:23 +0800 Subject: [PATCH 21/51] Expose bounded Runtime history API --- Makefile | 1 + contracts/agents-api/README.md | 3 +- contracts/agents-api/openapi.yaml | 429 ++++++++++++++++++ contracts/agents-api/runtime-history-api.md | 153 +++++++ .../agents-api/runtime-observability-api.md | 7 +- .../runtime-observability-design.md | 19 +- contracts/agents-api/runtime-observability.md | 11 +- contracts/agents-api/v1/runtime_history.go | 88 ++++ .../agents-api/v1/runtime_history_test.go | 85 ++++ docs/web/architecture.md | 2 +- docs/web/protocol-coverage.md | 20 +- packages/agents-client/src/client.ts | 39 ++ .../src/runtime-history-projection.ts | 298 ++++++++++++ .../agents-client/src/runtime-history.test.ts | 248 ++++++++++ packages/agents-client/src/types.ts | 94 ++++ scripts/patch-agents-openapi.py | 42 ++ services/agents-api/internal/api/handler.go | 3 + .../internal/api/runtime_history.go | 247 ++++++++++ .../internal/api/runtime_history_test.go | 250 ++++++++++ .../internal/runtimehistory/service.go | 43 +- .../internal/runtimehistory/service_test.go | 111 ++++- .../internal/runtimehistory/types.go | 167 ++++++- 22 files changed, 2295 insertions(+), 65 deletions(-) create mode 100644 contracts/agents-api/runtime-history-api.md create mode 100644 contracts/agents-api/v1/runtime_history.go create mode 100644 contracts/agents-api/v1/runtime_history_test.go create mode 100644 packages/agents-client/src/runtime-history-projection.ts create mode 100644 packages/agents-client/src/runtime-history.test.ts create mode 100644 scripts/patch-agents-openapi.py create mode 100644 services/agents-api/internal/api/runtime_history.go create mode 100644 services/agents-api/internal/api/runtime_history_test.go diff --git a/Makefile b/Makefile index f9142df47..22204fdf8 100644 --- a/Makefile +++ b/Makefile @@ -25,6 +25,7 @@ openapi: -g cmd/server/main.go --dir ./services/agents-api,./contracts/agents-api/v1 \ --output "$$output" \ --outputTypes yaml --parseInternal; \ + python3 scripts/patch-agents-openapi.py "$$output/swagger.yaml"; \ mv "$$output/swagger.yaml" contracts/agents-api/openapi.yaml check-sqlc: diff --git a/contracts/agents-api/README.md b/contracts/agents-api/README.md index 49250c862..b85c6ea6e 100644 --- a/contracts/agents-api/README.md +++ b/contracts/agents-api/README.md @@ -110,7 +110,8 @@ compatibility. | Extension | Operations | Current coverage | | --- | --- | --- | -| Runtime observations | `GET /v1/agents/runtime-observations`; `GET /v1/agents/sessions/{session_id}/runtime-observation` | Current, read-only, tenant-scoped Session contexts with stable Session-keyset pagination, bounded concurrent sampling, Docker metrics, explicit unsupported/unavailable states, strict `packages/agents-client` projection, and no lifecycle mutation. Kubernetes, E2B, self-hosted telemetry, history, CPU-rate derivation, and automatic idle policy remain unimplemented. See [Runtime observation API](runtime-observability-api.md). | +| Runtime observations | `GET /v1/agents/runtime-observations`; `GET /v1/agents/sessions/{session_id}/runtime-observation` | Current, read-only, tenant-scoped Session contexts with stable Session-keyset pagination, bounded concurrent sampling, Docker and microsandbox metrics, explicit unsupported/unavailable states, strict `packages/agents-client` projection, and no lifecycle mutation. Kubernetes, E2B, self-hosted telemetry, and automatic idle policy remain unimplemented. See [Runtime observation API](runtime-observability-api.md). | +| Runtime history | `GET /v1/agents/runtime-history/capabilities`; `GET /v1/agents/sessions/{session_id}/runtime-history` | Optional backend-neutral capability and bounded tenant/Session-scoped history contract with allocation/incarnation fencing, explicit coverage and strict client projection. Disabled by default until a production Reader and qualified periodic collection are configured; Durable Web rendering remains pending. See [Runtime history API](runtime-history-api.md). | For each resource, verify the referenced request/response unions and observable behavior, not just the route. Non-text initial input, configuration diff --git a/contracts/agents-api/openapi.yaml b/contracts/agents-api/openapi.yaml index d81a55fa5..242f72a42 100644 --- a/contracts/agents-api/openapi.yaml +++ b/contracts/agents-api/openapi.yaml @@ -1051,6 +1051,333 @@ definitions: - usage_seconds_total - utilization_ratio type: object + v1.RuntimeHistory: + properties: + coverage: + $ref: '#/definitions/v1.RuntimeHistoryCoverage' + generated_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + object: + enum: + - agent.runtime_history + type: string + requested_range: + $ref: '#/definitions/v1.RuntimeHistoryRange' + resolution_seconds: + maximum: 9007199254740991 + minimum: 1 + type: integer + series: + items: + $ref: '#/definitions/v1.RuntimeHistorySeries' + maxItems: 1000 + type: array + session_id: + format: uuid + type: string + source: + enum: + - durable + type: string + required: + - coverage + - generated_at + - object + - requested_range + - resolution_seconds + - series + - session_id + - source + type: object + v1.RuntimeHistoryCPU: + properties: + capacity_cores: + minimum: 5e-324 + type: number + x-nullable: true + contributor_count: + maximum: 9007199254740991 + minimum: 1 + type: integer + utilization_ratio: + minimum: 0 + type: number + x-nullable: true + required: + - capacity_cores + - contributor_count + - utilization_ratio + type: object + v1.RuntimeHistoryCapabilities: + properties: + available: + type: boolean + collection_mode: + enum: + - on_read + - periodic + type: string + x-nullable: true + maximum_points: + maximum: 10000 + minimum: 2 + type: integer + x-nullable: true + maximum_range_seconds: + maximum: 9007199254740991 + minimum: 1 + type: integer + x-nullable: true + metrics: + items: + enum: + - cpu + - memory + type: string + maxItems: 2 + type: array + minimum_step_seconds: + maximum: 9007199254740991 + minimum: 1 + type: integer + x-nullable: true + object: + enum: + - agent.runtime_history_capabilities + type: string + reason: + enum: + - not_configured + - periodic_collection_required + type: string + x-nullable: true + retention_seconds: + maximum: 9007199254740991 + minimum: 1 + type: integer + x-nullable: true + sample_interval_seconds: + maximum: 9007199254740991 + minimum: 1 + type: integer + x-nullable: true + required: + - available + - collection_mode + - maximum_points + - maximum_range_seconds + - metrics + - minimum_step_seconds + - object + - reason + - retention_seconds + - sample_interval_seconds + type: object + v1.RuntimeHistoryCoverage: + properties: + buckets: + items: + $ref: '#/definitions/v1.RuntimeHistoryCoveragePoint' + maxItems: 10000 + type: array + expected_sample_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + first_sample_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + last_sample_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + retained_start: + maximum: 9007199254740991 + minimum: 0 + type: integer + sample_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - buckets + - expected_sample_count + - first_sample_at + - last_sample_at + - retained_start + - sample_count + type: object + v1.RuntimeHistoryCoveragePoint: + properties: + end: + maximum: 9007199254740991 + minimum: 1 + type: integer + first_observed_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + last_observed_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + observation_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + observed_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + start: + maximum: 9007199254740991 + minimum: 0 + type: integer + unavailable_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - end + - first_observed_at + - last_observed_at + - observation_count + - observed_count + - start + - unavailable_count + type: object + v1.RuntimeHistoryMemory: + properties: + contributor_count: + maximum: 9007199254740991 + minimum: 1 + type: integer + limit_bytes: + maximum: 9007199254740991 + minimum: 1 + type: integer + x-nullable: true + usage_bytes: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + required: + - contributor_count + - limit_bytes + - usage_bytes + type: object + v1.RuntimeHistoryPoint: + properties: + cpu: + allOf: + - $ref: '#/definitions/v1.RuntimeHistoryCPU' + x-nullable: true + end: + maximum: 9007199254740991 + minimum: 1 + type: integer + first_observed_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + last_observed_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + x-nullable: true + memory: + allOf: + - $ref: '#/definitions/v1.RuntimeHistoryMemory' + x-nullable: true + observation_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + observed_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + start: + maximum: 9007199254740991 + minimum: 0 + type: integer + unavailable_count: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - cpu + - end + - first_observed_at + - last_observed_at + - memory + - observation_count + - observed_count + - start + - unavailable_count + type: object + v1.RuntimeHistoryRange: + properties: + end: + maximum: 9007199254740991 + minimum: 1 + type: integer + start: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - end + - start + type: object + v1.RuntimeHistorySeries: + properties: + allocation_id: + format: uuid + type: string + environment_id: + format: uuid + type: string + points: + items: + $ref: '#/definitions/v1.RuntimeHistoryPoint' + maxItems: 10000 + type: array + provider_type: + pattern: '^[a-z][a-z0-9_]{0,31}$' + type: string + started_at: + $ref: '#/definitions/v1.RuntimeHistoryTime' + required: + - allocation_id + - environment_id + - points + - provider_type + - started_at + type: object + v1.RuntimeHistoryTime: + properties: + nanoseconds: + maximum: 999999999 + minimum: 0 + type: integer + seconds: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - nanoseconds + - seconds + type: object v1.RuntimeInstance: properties: allocation_id: @@ -2877,6 +3204,37 @@ paths: summary: Update an Environment Template tags: - Environment Templates + /agents/runtime-history/capabilities: + get: + description: Core extension advertising only safe backend-neutral Durable history + availability and bounds. Available is true only when a query Reader and qualified + periodic collection are both configured. + parameters: + - description: agents=v1 + in: header + name: OpenAI-Beta + required: true + type: string + produces: + - application/json + responses: + "200": + description: OK + schema: + $ref: '#/definitions/v1.RuntimeHistoryCapabilities' + "400": + description: Bad Request + schema: + $ref: '#/definitions/v1.ErrorResponse' + "401": + description: Unauthorized + schema: + $ref: '#/definitions/v1.ErrorResponse' + security: + - BearerAuth: [] + summary: Retrieve Runtime history capabilities + tags: + - Runtime history /agents/runtime-observations: get: description: Core extension listing one current Runtime context per tenant-owned @@ -3725,6 +4083,77 @@ paths: summary: List persisted execution Items tags: - Items + /agents/sessions/{session_id}/runtime-history: + get: + description: Core extension returning tenant-scoped stored Runtime observations + for one Session. End is exclusive; the server selects a bounded resolution. + Responses contain at most 1,000 series, 10,000 points per coverage/series + array, and 100,000 total coverage plus series points. It never reads or changes + live compute. + parameters: + - description: agents=v1 + in: header + name: OpenAI-Beta + required: true + type: string + - description: Session ID + in: path + name: session_id + required: true + type: string + - description: Inclusive Unix-second start + in: query + maximum: 9007199254740991 + minimum: 0 + name: start + required: true + type: integer + - description: Exclusive Unix-second end + in: query + maximum: 9007199254740991 + minimum: 1 + name: end + required: true + type: integer + - description: Maximum points per series; defaults to the lower of 120 and the + advertised service maximum + in: query + maximum: 10000 + minimum: 2 + name: max_points + type: integer + produces: + - application/json + responses: + "200": + description: OK + schema: + $ref: '#/definitions/v1.RuntimeHistory' + "400": + description: Bad Request + schema: + $ref: '#/definitions/v1.ErrorResponse' + "401": + description: Unauthorized + schema: + $ref: '#/definitions/v1.ErrorResponse' + "404": + description: Not Found + schema: + $ref: '#/definitions/v1.ErrorResponse' + "409": + description: Conflict + schema: + $ref: '#/definitions/v1.ErrorResponse' + "503": + description: Service Unavailable + schema: + $ref: '#/definitions/v1.ErrorResponse' + security: + - BearerAuth: [] + summary: Retrieve Session Runtime history + tags: + - Runtime history /agents/sessions/{session_id}/runtime-observation: get: description: Core extension returning one tenant-scoped, read-only current Runtime diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md new file mode 100644 index 000000000..175abd6e4 --- /dev/null +++ b/contracts/agents-api/runtime-history-api.md @@ -0,0 +1,153 @@ +# Runtime history API + +Status: public contract and strict TypeScript client implemented; no production +Reader is configured by default. The capability route therefore advertises +`available=false` until an operator supplies both a validated Reader and qualified +periodic collection. + +This is an optional Agents Core extension. It is read-only and backend-neutral. +The browser never receives a ClickHouse endpoint, OTLP credential, provider-native +identity, or tenant selector. + +## Capability discovery + +```http +GET /v1/agents/runtime-history/capabilities +OpenAI-Beta: agents=v1 +Authorization: Bearer ... +``` + +The route accepts no query parameters. An unconfigured Core returns: + +```json +{ + "object": "agent.runtime_history_capabilities", + "available": false, + "reason": "not_configured", + "collection_mode": null, + "sample_interval_seconds": null, + "retention_seconds": null, + "minimum_step_seconds": null, + "maximum_range_seconds": null, + "maximum_points": null, + "metrics": [] +} +``` + +`available=true` requires `collection_mode=periodic`, a qualified positive sample +interval, a validated Reader, retention and query bounds, and at least one of +`cpu` or `memory`. A Reader backed only by request-triggered samples returns +`reason=periodic_collection_required`; Web must not call that data Durable. +Malformed capability configuration fails closed as unconfigured. Backend names, +URLs, credentials, table names, and tenant data are never capability fields. + +## Session history + +```http +GET /v1/agents/sessions/{session_id}/runtime-history?start=1789951200&end=1789954800&max_points=120 +OpenAI-Beta: agents=v1 +Authorization: Bearer ... +``` + +`start` is inclusive and `end` is exclusive, both in whole Unix seconds. +`max_points` defaults to the lower of 120 and the advertised maximum. Core +selects an effective whole-second resolution. Unknown parameters, duplicate +parameters, negative timestamps, invalid ranges, and invalid point limits are +rejected before a Reader query. + +The caller supplies only a Session ID. Core obtains the tenant from authentication, +resolves the Session and Environment from its store, and only then calls the +Reader. Missing and foreign Sessions remain indistinguishable. Allocation IDs, +provider IDs, and backend labels are result identity, never authority-bearing +query inputs. + +```json +{ + "object": "agent.runtime_history", + "source": "durable", + "session_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", + "requested_range": { "start": 1789951200, "end": 1789954800 }, + "resolution_seconds": 60, + "generated_at": 1789954801, + "coverage": { + "retained_start": 1789951200, + "first_sample_at": 1789951210, + "last_sample_at": 1789954750, + "sample_count": 118, + "expected_sample_count": 120, + "buckets": [] + }, + "series": [] +} +``` + +`coverage` describes all resolved samples, including unavailable observations that +cannot safely be attached to one compute incarnation. `retained_start` is the +latest of the requested start, configured retention boundary, and backend-reported +retention boundary. Expected coverage is calculated only over that retained +period and only from qualified periodic cadence. + +Resource `series` are split by `allocation_id` plus the lossless compute +`started_at` fence: + +```json +{ + "started_at": { + "seconds": 1789951210, + "nanoseconds": 123456789 + } +} +``` + +The two integers are JSON-safe, nonnegative Unix seconds and a 0–999,999,999 +nanosecond remainder. They are identity, not merely a display timestamp. CPU +utilization is derived only from ordered cumulative counters inside that fence; +memory values are the last observed values in a bucket. Every point contains +observation and contributor counts. Missing values are null and gaps remain gaps. +Numeric zero is retained as an observed value. The endpoint does not return token +history: token throughput remains sourced from canonical Session Usage and is +Live-only until a separate durable usage contract exists. + +## Errors and bounds + +| HTTP | Code | Meaning | +| --- | --- | --- | +| 400 | `unsupported_parameter` or `invalid_request` | Invalid query shape or range. | +| 401 | `authentication_error` | Missing or invalid authentication. | +| 404 | `not_found` | Missing or foreign Session. | +| 409 | `runtime_history_unsupported` | The Session has no supported managed Runtime history scope. | +| 503 | `runtime_history_unavailable` | Durable history is unconfigured, timed out, unavailable, or returned malformed data. | + +Reader failures and malformed results are sanitized. Raw backend errors are not +logged or returned. Results are bounded by per-series points, series count, and +total response points. Coverage and series points must be ordered, non-overlapping, +inside the half-open request range, and contain valid safe JSON values. + +## Client contract + +`packages/agents-client` exposes: + +```ts +interface AgentCore { + getRuntimeHistoryCapabilities(options?: ReadOptions): Promise; + retrieveRuntimeHistory(sessionId: string, query: { + start: number; + end: number; + maxPoints?: number; + signal?: AbortSignal; + }): Promise; +} +``` + +The client validates exact fields, capability consistency, requested-range echo, +Session identity, half-open bucket ordering, coverage totals, incarnation identity, +contributor counts, nullability, finite numbers, and response size. Unknown fields +or malformed data reject the entire response with a 502 client projection error. + +## Explicit boundaries + +- The routes never sample a live provider, provision compute, or mutate lifecycle. +- The contract does not choose ClickHouse, Prometheus, Mimir, or another backend. +- History availability does not imply current Runtime readiness. +- Current observations and Durable history have separate freshness and retention + semantics and must remain separately labelled in Web. diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 7c967cb18..668d5c28d 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -3,7 +3,9 @@ Status: Phase 2 and initial Core Web consumption implemented. The current-snapshot routes, strict `packages/agents-client` projection, and generated `openapi.yaml` contract are implemented and consumed by the Dashboard through complete Session/observation -identity joins. Historical queries and lifecycle controls remain outside this phase. +identity joins. Durable history uses the separate optional +[Runtime history API](runtime-history-api.md); lifecycle controls remain outside +this phase. This is an Agents Core extension, not an upstream OpenAI Agents resource. The implementation must record that status in the coverage ledger and generated @@ -264,7 +266,8 @@ deletion makes the candidate incomplete and prevents publication. - Token usage: use existing Session/Turn Usage. - Billing and cost: product/backend concern. -- Historical series: optional later capability with a separate contract. +- Historical series in these routes: the optional capability uses the separate + [Runtime history API](runtime-history-api.md). - Container logs and command output. - Provider credentials or native configuration. - Start, stop, pause, resume, restart, renew, or delete operations. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index e88b0811f..b503eb96a 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -351,7 +351,7 @@ lease before every periodic export handoff. `on_read` remains distinct from `periodic`, so ad hoc API traffic cannot be counted as qualified cadence coverage. -A future history API must expose actual sample coverage. Core Web must not +A history API must expose actual sample coverage. Core Web must not advertise a durable range until the operator backend, query adapter, and a qualified periodic collection cadence are all configured. @@ -364,7 +364,8 @@ before calling a Reader. Reader queries always carry tenant, Session, and Environment scope plus a bounded start, exclusive end, server-selected step, and total point budget. Provider-native identity is never a query input. -Reader results remain divided by allocation and compute `started_at` fence. +Reader results remain divided by allocation and the lossless compute `started_at` +seconds-plus-nanoseconds fence. Every bucket reports explicit observation coverage and nullable CPU/memory values. CPU utilization may be derived only from ordered cumulative counters inside one fence; memory uses the final observed value in the bucket. Empty @@ -378,9 +379,10 @@ and supported metrics. A configured Reader without qualified periodic sampling is not sufficient to advertise a Durable Dashboard source. Backend identity, URLs, credentials, and tenant data are never capability fields. -This internal boundary is implemented, but no production Reader or public -history route is configured yet. The next qualification adds the ClickHouse -reference Reader, then a Session-scoped public extension and strict client. +This internal boundary, the Session-scoped public extension, capability discovery, +and strict client are implemented. No production Reader is configured yet. The +next qualification adds the ClickHouse reference Reader and end-to-end retention, +isolation, restart, and incarnation evidence before Web advertises Durable ranges. ## 11. Dashboard information architecture @@ -542,11 +544,14 @@ Implemented for the browser-local current-snapshot live window. - Implemented: backend-neutral `runtimehistory` types and service validation. Tenant/Session/Environment scope precedes every Reader query; incarnation, coverage, nullability, ordering, range and total-point invariants are enforced. +- Implemented: safe public capability discovery, bounded Session-scoped history + query routes, generated OpenAPI schemas, and strict `packages/agents-client` + projection. Unconfigured or on-read-only deployments cannot advertise Durable. - Qualified: optional OTLP Collector fan-out with a separate high-cardinality history store and server-side tenant-scoped query adapter. ClickHouse is the first reference backend; no backend is a Core execution dependency. -- Not implemented: a production `runtimehistory` Reader, public history - extension, durable Web ranges, retention configuration, and exporter +- Not implemented: a production `runtimehistory` Reader, durable Web ranges, + retention deployment configuration, and exporter queue/drop/error coverage telemetry. - Add telemetry exporter and qualified operator backend. - Define a separate history query adapter and retention/security policy. diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 519f8c40b..e9a45c4ea 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -104,8 +104,9 @@ capabilities in the [full design](runtime-observability-design.md). The internal `runtimehistory` boundary is backend-neutral and validates Core scope, incarnation fences, bucket coverage, nullability, time bounds and total -point limits. It does not make history available by itself: no production -Reader or public history route is configured in this phase. +point limits. The public capability and Session history routes plus strict client +projection are implemented, but they do not make history available by themselves: +no production Reader is configured in this phase. ## First-phase boundary @@ -116,5 +117,7 @@ attribution or the existing sandbox lifecycle interface. The current API and browser-local Web live window are documented in the [full design](runtime-observability-design.md) and the -[public extension](runtime-observability-api.md). History queries, additional -providers, self-hosted telemetry, and idle-policy authority remain later phases. +[current-snapshot extension](runtime-observability-api.md), and the optional +[history extension](runtime-history-api.md). The production history Reader, +additional providers, self-hosted telemetry, and idle-policy authority remain +later phases. diff --git a/contracts/agents-api/v1/runtime_history.go b/contracts/agents-api/v1/runtime_history.go new file mode 100644 index 000000000..741c1ba04 --- /dev/null +++ b/contracts/agents-api/v1/runtime_history.go @@ -0,0 +1,88 @@ +package v1 + +type RuntimeHistoryCapabilities struct { + Object string `json:"object" enums:"agent.runtime_history_capabilities" binding:"required"` + Available bool `json:"available" binding:"required"` + Reason *string `json:"reason" extensions:"x-nullable" binding:"required" enums:"not_configured,periodic_collection_required"` + CollectionMode *string `json:"collection_mode" extensions:"x-nullable" binding:"required" enums:"on_read,periodic"` + SampleIntervalSeconds *int64 `json:"sample_interval_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` + RetentionSeconds *int64 `json:"retention_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` + MinimumStepSeconds *int64 `json:"minimum_step_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` + MaximumRangeSeconds *int64 `json:"maximum_range_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` + MaximumPoints *int `json:"maximum_points" extensions:"x-nullable" binding:"required" minimum:"2" maximum:"10000"` + Metrics []string `json:"metrics" binding:"required" validate:"max=2" enums:"cpu,memory"` +} + +type RuntimeHistory struct { + Object string `json:"object" enums:"agent.runtime_history" binding:"required"` + Source string `json:"source" enums:"durable" binding:"required"` + SessionID string `json:"session_id" binding:"required" format:"uuid"` + RequestedRange RuntimeHistoryRange `json:"requested_range" binding:"required"` + ResolutionSeconds int64 `json:"resolution_seconds" binding:"required" minimum:"1" maximum:"9007199254740991"` + GeneratedAt int64 `json:"generated_at" binding:"required" minimum:"0" maximum:"9007199254740991"` + Coverage RuntimeHistoryCoverage `json:"coverage" binding:"required"` + Series []RuntimeHistorySeries `json:"series" binding:"required" validate:"max=1000"` +} + +type RuntimeHistoryRange struct { + Start int64 `json:"start" binding:"required" minimum:"0" maximum:"9007199254740991"` + End int64 `json:"end" binding:"required" minimum:"1" maximum:"9007199254740991"` +} + +type RuntimeHistoryCoverage struct { + RetainedStart int64 `json:"retained_start" binding:"required" minimum:"0" maximum:"9007199254740991"` + FirstSampleAt *int64 `json:"first_sample_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + LastSampleAt *int64 `json:"last_sample_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + SampleCount int64 `json:"sample_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + ExpectedSampleCount int64 `json:"expected_sample_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + Buckets []RuntimeHistoryCoveragePoint `json:"buckets" binding:"required" validate:"max=10000"` +} + +type RuntimeHistoryCoveragePoint struct { + Start int64 `json:"start" binding:"required" minimum:"0" maximum:"9007199254740991"` + End int64 `json:"end" binding:"required" minimum:"1" maximum:"9007199254740991"` + FirstObservedAt *int64 `json:"first_observed_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + LastObservedAt *int64 `json:"last_observed_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + ObservationCount int `json:"observation_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + ObservedCount int `json:"observed_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + UnavailableCount int `json:"unavailable_count" binding:"required" minimum:"0" maximum:"9007199254740991"` +} + +type RuntimeHistorySeries struct { + EnvironmentID string `json:"environment_id" binding:"required" format:"uuid"` + AllocationID string `json:"allocation_id" binding:"required" format:"uuid"` + StartedAt RuntimeHistoryTime `json:"started_at" binding:"required"` + ProviderType string `json:"provider_type" binding:"required" pattern:"^[a-z][a-z0-9_]{0,31}$"` + Points []RuntimeHistoryPoint `json:"points" binding:"required" validate:"max=10000"` +} + +// RuntimeHistoryTime is a lossless JSON-safe timestamp used as an incarnation +// identity fence. Other public history timestamps are display/query seconds. +type RuntimeHistoryTime struct { + Seconds int64 `json:"seconds" binding:"required" minimum:"0" maximum:"9007199254740991"` + Nanoseconds int `json:"nanoseconds" binding:"required" minimum:"0" maximum:"999999999"` +} + +type RuntimeHistoryPoint struct { + Start int64 `json:"start" binding:"required" minimum:"0" maximum:"9007199254740991"` + End int64 `json:"end" binding:"required" minimum:"1" maximum:"9007199254740991"` + FirstObservedAt *int64 `json:"first_observed_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + LastObservedAt *int64 `json:"last_observed_at" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + ObservationCount int `json:"observation_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + ObservedCount int `json:"observed_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + UnavailableCount int `json:"unavailable_count" binding:"required" minimum:"0" maximum:"9007199254740991"` + CPU *RuntimeHistoryCPU `json:"cpu" extensions:"x-nullable" binding:"required"` + Memory *RuntimeHistoryMemory `json:"memory" extensions:"x-nullable" binding:"required"` +} + +type RuntimeHistoryCPU struct { + ContributorCount int `json:"contributor_count" binding:"required" minimum:"1" maximum:"9007199254740991"` + UtilizationRatio *float64 `json:"utilization_ratio" extensions:"x-nullable" binding:"required" minimum:"0"` + CapacityCores *float64 `json:"capacity_cores" extensions:"x-nullable" binding:"required" minimum:"5e-324"` +} + +type RuntimeHistoryMemory struct { + ContributorCount int `json:"contributor_count" binding:"required" minimum:"1" maximum:"9007199254740991"` + UsageBytes *uint64 `json:"usage_bytes" extensions:"x-nullable" binding:"required" minimum:"0" maximum:"9007199254740991"` + LimitBytes *uint64 `json:"limit_bytes" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` +} diff --git a/contracts/agents-api/v1/runtime_history_test.go b/contracts/agents-api/v1/runtime_history_test.go new file mode 100644 index 000000000..bd986d68d --- /dev/null +++ b/contracts/agents-api/v1/runtime_history_test.go @@ -0,0 +1,85 @@ +package v1 + +import ( + "os" + "strings" + "testing" + + "gopkg.in/yaml.v3" +) + +func TestRuntimeHistoryOpenAPICollectionLimits(t *testing.T) { + t.Parallel() + + raw, err := os.ReadFile("../openapi.yaml") + if err != nil { + t.Fatalf("read generated OpenAPI contract: %v", err) + } + + var document map[string]any + if err := yaml.Unmarshal(raw, &document); err != nil { + t.Fatalf("parse generated OpenAPI contract: %v", err) + } + + definitions := openAPIMap(t, document, "definitions") + assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistoryCapabilities", "metrics", 2) + assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistory", "series", 1000) + assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistoryCoverage", "buckets", 10000) + assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistorySeries", "points", 10000) + capabilities := openAPIMap(t, definitions, "v1.RuntimeHistoryCapabilities") + metrics := openAPIMap(t, openAPIMap(t, capabilities, "properties"), "metrics") + metricItems := openAPIMap(t, metrics, "items") + assertOpenAPIStrings(t, metricItems, "enum", []string{"cpu", "memory"}) + series := openAPIMap(t, definitions, "v1.RuntimeHistorySeries") + providerType := openAPIMap(t, openAPIMap(t, series, "properties"), "provider_type") + if actual, ok := providerType["pattern"].(string); !ok || actual != "^[a-z][a-z0-9_]{0,31}$" { + t.Fatalf("Runtime history provider_type pattern = %#v, want canonical provider pattern", providerType["pattern"]) + } + + if !strings.Contains(string(raw), "100,000 total coverage plus series points") { + t.Fatal("runtime history operation must document the 100,000 total point limit") + } +} + +func assertOpenAPIStrings(t *testing.T, parent map[string]any, key string, expected []string) { + t.Helper() + + values, ok := parent[key].([]any) + if !ok || len(values) != len(expected) { + t.Fatalf("OpenAPI key %q = %#v, want %v", key, parent[key], expected) + } + for index, expectedValue := range expected { + if actual, ok := values[index].(string); !ok || actual != expectedValue { + t.Fatalf("OpenAPI key %q item %d = %#v, want %q", key, index, values[index], expectedValue) + } + } +} + +func assertOpenAPIArrayLimit(t *testing.T, definitions map[string]any, definition, property string, expected int) { + t.Helper() + + schema := openAPIMap(t, definitions, definition) + properties := openAPIMap(t, schema, "properties") + field := openAPIMap(t, properties, property) + actual, ok := field["maxItems"].(int) + if !ok { + t.Fatalf("%s.%s maxItems has type %T, want int", definition, property, field["maxItems"]) + } + if actual != expected { + t.Fatalf("%s.%s maxItems = %d, want %d", definition, property, actual, expected) + } +} + +func openAPIMap(t *testing.T, parent map[string]any, key string) map[string]any { + t.Helper() + + value, ok := parent[key] + if !ok { + t.Fatalf("OpenAPI key %q is missing", key) + } + result, ok := value.(map[string]any) + if !ok { + t.Fatalf("OpenAPI key %q has type %T, want object", key, value) + } + return result +} diff --git a/docs/web/architecture.md b/docs/web/architecture.md index e7852f84a..4836a13a7 100644 --- a/docs/web/architecture.md +++ b/docs/web/architecture.md @@ -80,7 +80,7 @@ that is a separate product layer. | Layer | Owns | Does not own | | --- | --- | --- | | Agents Core Web | navigation, Agent forms, Vault/Credential metadata controls, Session timeline, live-state projection, reconnect recovery, function-result UI | durable truth, Credential decryption, engine selection, sandbox or provider credentials | -| `@agents-core-web/agents-client` | `/v1/agents/**`, `/v1/vaults/**`, and Source `/v1/files**` wire types; route-specific beta/auth headers; pagination; strict Environment/File/Vault metadata projection; write-only Credential create/replace; fetch-based SSE parsing | uncertain-write retries, secret readback, native protocol translation | +| `@agents-core-web/agents-client` | `/v1/agents/**`, `/v1/vaults/**`, and Source `/v1/files**` wire types; route-specific beta/auth headers; pagination; strict Environment/File/Vault and Runtime current/history projection; write-only Credential create/replace; fetch-based SSE parsing | uncertain-write retries, secret readback, native protocol translation, Runtime persistence | | Agent Core | principal authentication, validation, idempotency, durable Agent/Session/Turn/Item and Vault/Credential truth, server-side Credential encryption/use, live events, scheduling | product organization UI or Web user sessions | | `parsar-daemon` | device connection, host capability advertisement, native process lifecycle and translation | public Agents HTTP semantics or product policy | | Native adapter | Codex app-server or Claude Agent SDK integration | public API and Web deployment policy | diff --git a/docs/web/protocol-coverage.md b/docs/web/protocol-coverage.md index 8379a5c81..d11ae5477 100644 --- a/docs/web/protocol-coverage.md +++ b/docs/web/protocol-coverage.md @@ -83,6 +83,8 @@ the Core key binding. Agents Core Web's local proxy owns the bearer server-side. | Environment Templates retrieve/update/delete | Yes | No | Reusable client methods only; no Web management surface, so a Template is never edited or deleted from the browser | | Environment keys | No public browser API | Hidden | Operator-issued executor credentials stay on executor compute and never enter browser state, request previews, navigation, or Create actions | | Vaults and static-bearer Credentials | Yes | Yes, when discovered | Web traverses the Vault and per-Vault Credential page chains to their Core end markers before publishing one loaded metadata result, provides safe lifecycle controls and write-only token create/replace, and deterministically attaches each selected Credential's owning Vault to Session creation. The reads are non-atomic and not a current Core total. Tokens are never returned; catalog success is not runtime proof | +| Runtime observations | Core extension | Yes, read-only | Dashboard traverses the complete tenant-scoped current-observation page chain, requires an exact Session identity join, and publishes only complete snapshots. Docker and microsandbox values remain provider evidence; unsupported, unavailable, stale, and unknown are distinct. Browser-local trends are explicitly Live and ephemeral | +| Runtime history | Optional Core extension | Client implemented; Dashboard pending | Capability discovery and bounded Session history are strictly projected through `packages/agents-client`. Durable UI must remain disabled until capability discovery reports qualified periodic collection and a production Reader; token throughput is not synthesized into history | | Protocol Subagents / enabled multi-agent | Later | No | Distinct from storing multiple Agent configurations | | Usage/observability | Response types | Yes, scoped | Session aggregate and per-Turn token Usage are labelled separately; unavailable measurements remain unknown, not zero | @@ -397,13 +399,16 @@ upstream. ## Dashboard and System boundary -- Dashboard is a Web-derived view over the last successfully traversed Agent and - Session page-chain results for the configured Core access scope. Web follows every +- Dashboard is a Web-derived view over the last successfully traversed Agent, + Session, and Runtime-observation page-chain results for the configured Core + access scope. Web follows every continuation with `limit=100&order=desc`, rejects duplicate identities and invalid or cyclic cursors, and allows at most 100 pages per collection. If Core reports that page 101 is required, the refresh fails closed and does not publish - the partial result. Dashboard performs no additional Turn, Item, Environment, - or execution-readiness requests and makes no writes. + the partial result. The Runtime collection must complete and have exactly the + same Session ID set before it replaces the prior snapshot. Dashboard performs + no Turn, Item, Environment, lifecycle, or execution-readiness requests and makes + no writes. - When both top-level collection reads fail through a gateway/network condition, Dashboard labels the local Agent Core backend as not ready and links the whole notice to connection recovery. An explicitly configured local Docker guide may @@ -428,6 +433,13 @@ upstream. at most eight rows ordered by valid Core-reported `last_active_at`. A `self_hosted` label identifies only the Session profile, not an executor connection. Row actions navigate to the exact loaded Session. +- Runtime charts use bounded browser-local samples from complete joined snapshots. + CPU rate is derived only from ordered cumulative counters for the same allocation + and compute incarnation; memory remains point-in-time; gaps are not interpolated. + The 15-minute and one-hour ranges are labelled Live and reset across browser + lifecycle. Durable history capability and Session-query client methods exist, + but Web must not advertise Durable until the operator Reader and qualified + periodic sampling pass end-to-end acceptance. Token throughput remains Live-only. - Agent and Session collection states remain independent. A failed refresh may leave an explicitly labelled prior loaded result visible, including a previously confirmed empty result. An initial failed empty collection is unavailable diff --git a/packages/agents-client/src/client.ts b/packages/agents-client/src/client.ts index 24cc80cac..4229c74e2 100644 --- a/packages/agents-client/src/client.ts +++ b/packages/agents-client/src/client.ts @@ -1,6 +1,7 @@ import { exactFields, onlyFields, isRecord, hasOwn, canonicalUuid, isNonnegativeInteger, sameResourceId } from "./response-projection"; import { projectTokenUsage } from "./usage-projection"; import { projectAgentTurn, projectSessionItem, projectItemContent, projectHistoryPage, validateHistoryPageOptions } from "./history-projection"; +import { projectRuntimeHistory, projectRuntimeHistoryCapabilities } from "./runtime-history-projection"; import { createSSEDecoder } from "./sse"; import { projectVaultCredentialAuth, validCredentialURL } from "./vault-credential-auth"; import type { @@ -48,6 +49,9 @@ import type { ReplaceVaultCredentialTokenInput, RuntimeObservation, RuntimeObservationList, + RuntimeHistory, + RuntimeHistoryCapabilities, + RuntimeHistoryQuery, Vault, VaultCredential, VaultCredentialDeleted, @@ -923,6 +927,14 @@ function invalidRuntimeObservation(message = "Agent Core returned an invalid Run throw new AgentCoreError(message, 502, "invalid_runtime_observation"); } +function invalidRuntimeHistoryCapabilities(message = "Agent Core returned invalid Runtime history capabilities."): never { + throw new AgentCoreError(message, 502, "invalid_runtime_history_capabilities"); +} + +function invalidRuntimeHistory(message = "Agent Core returned invalid Runtime history."): never { + throw new AgentCoreError(message, 502, "invalid_runtime_history"); +} + function nullableRuntimeNumber(value: unknown): number | null { if (value === null) return null; if (typeof value !== "number" || !Number.isFinite(value) || value < 0) { @@ -1975,6 +1987,33 @@ export class OpenAIAgentsClient implements AgentCore { return projectRuntimeObservation(value, sessionId); } + async getRuntimeHistoryCapabilities(options?: ReadOptions): Promise { + const value = await this.request( + "/agents/runtime-history/capabilities", + { signal: options?.signal }, + ); + return projectRuntimeHistoryCapabilities(value, invalidRuntimeHistoryCapabilities); + } + + async retrieveRuntimeHistory(sessionId: string, query: RuntimeHistoryQuery): Promise { + const canonicalSessionId = canonicalUuid(sessionId); + if ( + query == null || canonicalSessionId === null || !isNonnegativeInteger(query.start) || !isNonnegativeInteger(query.end) || + query.end <= query.start || (query.maxPoints !== undefined && ( + !Number.isSafeInteger(query.maxPoints) || query.maxPoints < 2 || query.maxPoints > 10_000 + )) + ) throw new TypeError("Runtime history query is invalid."); + const params = new URLSearchParams(); + params.set("start", String(query.start)); + params.set("end", String(query.end)); + if (query.maxPoints !== undefined) params.set("max_points", String(query.maxPoints)); + const value = await this.request( + withQuery(`/agents/sessions/${encodeURIComponent(canonicalSessionId)}/runtime-history`, params), + { signal: query.signal }, + ); + return projectRuntimeHistory(value, canonicalSessionId, query, invalidRuntimeHistory); + } + async createSession(input: CreateSessionInput, idempotencyKey = createIdempotencyKey()): Promise { if ((input as { stream?: boolean }).stream === true) { throw new TypeError("createSession only supports the JSON response; connect streamEvents after creation."); diff --git a/packages/agents-client/src/runtime-history-projection.ts b/packages/agents-client/src/runtime-history-projection.ts new file mode 100644 index 000000000..38c79df43 --- /dev/null +++ b/packages/agents-client/src/runtime-history-projection.ts @@ -0,0 +1,298 @@ +import { canonicalUuid, exactFields, isNonnegativeInteger, isRecord, sameResourceId } from "./response-projection"; +import type { + RuntimeHistory, + RuntimeHistoryCPU, + RuntimeHistoryCapabilities, + RuntimeHistoryCoveragePoint, + RuntimeHistoryMemory, + RuntimeHistoryPoint, + RuntimeHistoryQuery, + RuntimeHistorySeries, +} from "./types"; + +type InvalidRuntimeHistory = (message?: string) => never; + +const capabilityFields = new Set([ + "object", "available", "reason", "collection_mode", "sample_interval_seconds", + "retention_seconds", "minimum_step_seconds", "maximum_range_seconds", "maximum_points", "metrics", +]); +const historyFields = new Set([ + "object", "source", "session_id", "requested_range", "resolution_seconds", "generated_at", "coverage", "series", +]); +const rangeFields = new Set(["start", "end"]); +const coverageFields = new Set([ + "retained_start", "first_sample_at", "last_sample_at", "sample_count", "expected_sample_count", "buckets", +]); +const coveragePointFields = new Set([ + "start", "end", "first_observed_at", "last_observed_at", "observation_count", "observed_count", "unavailable_count", +]); +const seriesFields = new Set(["environment_id", "allocation_id", "started_at", "provider_type", "points"]); +const timeFields = new Set(["seconds", "nanoseconds"]); +const pointFields = new Set([...coveragePointFields, "cpu", "memory"]); +const cpuFields = new Set(["contributor_count", "utilization_ratio", "capacity_cores"]); +const memoryFields = new Set(["contributor_count", "usage_bytes", "limit_bytes"]); +const providerTypePattern = /^[a-z][a-z0-9_]{0,31}$/; +const metrics = new Set(["cpu", "memory"]); +const reasons = new Set(["not_configured", "periodic_collection_required"]); +const maximumSeries = 1_000; +const maximumTotalPoints = 100_000; + +function positiveInteger(value: unknown): value is number { + return isNonnegativeInteger(value) && value > 0; +} + +function nullableNonnegativeInteger(value: unknown, invalid: InvalidRuntimeHistory): number | null { + if (value === null) return null; + if (!isNonnegativeInteger(value)) return invalid(); + return value; +} + +function nullablePositiveInteger(value: unknown, invalid: InvalidRuntimeHistory): number | null { + if (value === null) return null; + if (!positiveInteger(value)) return invalid(); + return value; +} + +function nullableNonnegativeNumber(value: unknown, invalid: InvalidRuntimeHistory): number | null { + if (value === null) return null; + if (typeof value !== "number" || !Number.isFinite(value) || value < 0) return invalid(); + return value; +} + +function nullablePositiveNumber(value: unknown, invalid: InvalidRuntimeHistory): number | null { + const projected = nullableNonnegativeNumber(value, invalid); + if (projected === 0) return invalid(); + return projected; +} + +export function projectRuntimeHistoryCapabilities( + value: unknown, + invalid: InvalidRuntimeHistory, +): RuntimeHistoryCapabilities { + if ( + !isRecord(value) || !exactFields(value, capabilityFields) || + value.object !== "agent.runtime_history_capabilities" || typeof value.available !== "boolean" || + !(value.reason === null || (typeof value.reason === "string" && reasons.has(value.reason))) || + !(value.collection_mode === null || value.collection_mode === "on_read" || value.collection_mode === "periodic") || + !Array.isArray(value.metrics) || value.metrics.some((metric) => typeof metric !== "string" || !metrics.has(metric)) || + new Set(value.metrics).size !== value.metrics.length || value.metrics.length > metrics.size + ) return invalid(); + + const sampleInterval = nullablePositiveInteger(value.sample_interval_seconds, invalid); + const retention = nullablePositiveInteger(value.retention_seconds, invalid); + const minimumStep = nullablePositiveInteger(value.minimum_step_seconds, invalid); + const maximumRange = nullablePositiveInteger(value.maximum_range_seconds, invalid); + const maximumPoints = nullablePositiveInteger(value.maximum_points, invalid); + const disabled = value.reason === "not_configured"; + const onRead = value.reason === "periodic_collection_required"; + if ( + (disabled && ( + value.available || value.collection_mode !== null || sampleInterval !== null || retention !== null || + minimumStep !== null || maximumRange !== null || maximumPoints !== null || value.metrics.length !== 0 + )) || + (onRead && ( + value.available || value.collection_mode !== "on_read" || sampleInterval !== null || retention === null || + minimumStep === null || maximumRange === null || maximumPoints === null || value.metrics.length === 0 + )) || + (value.available && ( + value.reason !== null || value.collection_mode !== "periodic" || sampleInterval === null || retention === null || + minimumStep === null || maximumRange === null || maximumPoints === null || value.metrics.length === 0 + )) || + (!value.available && value.reason === null) || + (retention !== null && maximumRange !== null && maximumRange > retention) || + (maximumPoints !== null && (maximumPoints < 2 || maximumPoints > 10_000)) + ) return invalid(); + + return { + object: "agent.runtime_history_capabilities", + available: value.available, + reason: value.reason as RuntimeHistoryCapabilities["reason"], + collection_mode: value.collection_mode as RuntimeHistoryCapabilities["collection_mode"], + sample_interval_seconds: sampleInterval, + retention_seconds: retention, + minimum_step_seconds: minimumStep, + maximum_range_seconds: maximumRange, + maximum_points: maximumPoints, + metrics: [...value.metrics] as RuntimeHistoryCapabilities["metrics"], + }; +} + +function projectCoveragePoint( + value: unknown, + fields: Set, + requestedStart: number, + requestedEnd: number, + resolution: number, + invalid: InvalidRuntimeHistory, +): RuntimeHistoryCoveragePoint { + if ( + !isRecord(value) || !exactFields(value, fields) || + !isNonnegativeInteger(value.start) || !positiveInteger(value.end) || + value.start < requestedStart || value.end <= value.start || value.end > requestedEnd || value.end - value.start > resolution || + !isNonnegativeInteger(value.observation_count) || !isNonnegativeInteger(value.observed_count) || + !isNonnegativeInteger(value.unavailable_count) || value.observed_count + value.unavailable_count !== value.observation_count + ) return invalid(); + const first = nullableNonnegativeInteger(value.first_observed_at, invalid); + const last = nullableNonnegativeInteger(value.last_observed_at, invalid); + if ( + (value.observation_count === 0 && (first !== null || last !== null)) || + (value.observation_count > 0 && ( + first === null || last === null || first < value.start || last < first || last >= value.end + )) + ) return invalid(); + return { + start: value.start, + end: value.end, + first_observed_at: first, + last_observed_at: last, + observation_count: value.observation_count, + observed_count: value.observed_count, + unavailable_count: value.unavailable_count, + }; +} + +function projectCPU(value: unknown, observedCount: number, invalid: InvalidRuntimeHistory): RuntimeHistoryCPU | null { + if (value === null) return null; + if (!isRecord(value) || !exactFields(value, cpuFields) || !positiveInteger(value.contributor_count) || value.contributor_count > observedCount) return invalid(); + const utilization = nullableNonnegativeNumber(value.utilization_ratio, invalid); + const capacity = nullablePositiveNumber(value.capacity_cores, invalid); + if (utilization === null && capacity === null) return invalid(); + return { contributor_count: value.contributor_count, utilization_ratio: utilization, capacity_cores: capacity }; +} + +function projectMemory(value: unknown, observedCount: number, invalid: InvalidRuntimeHistory): RuntimeHistoryMemory | null { + if (value === null) return null; + if (!isRecord(value) || !exactFields(value, memoryFields) || !positiveInteger(value.contributor_count) || value.contributor_count > observedCount) return invalid(); + const usage = nullableNonnegativeInteger(value.usage_bytes, invalid); + const limit = nullablePositiveInteger(value.limit_bytes, invalid); + if (usage === null && limit === null) return invalid(); + return { contributor_count: value.contributor_count, usage_bytes: usage, limit_bytes: limit }; +} + +function projectPoint( + value: unknown, + requestedStart: number, + requestedEnd: number, + resolution: number, + startedAt: number, + invalid: InvalidRuntimeHistory, +): RuntimeHistoryPoint { + const coverage = projectCoveragePoint(value, pointFields, requestedStart, requestedEnd, resolution, invalid); + if (!isRecord(value) || coverage.end <= startedAt || (coverage.first_observed_at !== null && coverage.first_observed_at < startedAt)) return invalid(); + const cpu = projectCPU(value.cpu, coverage.observed_count, invalid); + const memory = projectMemory(value.memory, coverage.observed_count, invalid); + if (coverage.observation_count === 0 && (cpu !== null || memory !== null)) return invalid(); + return { ...coverage, cpu, memory }; +} + +function ordered(points: readonly RuntimeHistoryCoveragePoint[]): boolean { + return points.every((point, index) => index === 0 || points[index - 1]!.end <= point.start); +} + +function projectSeries( + value: unknown, + requestedStart: number, + requestedEnd: number, + resolution: number, + generatedAt: number, + maximumPoints: number, + invalid: InvalidRuntimeHistory, +): RuntimeHistorySeries { + if (!isRecord(value) || !exactFields(value, seriesFields) || !Array.isArray(value.points)) return invalid(); + const environmentId = canonicalUuid(value.environment_id); + const allocationId = canonicalUuid(value.allocation_id); + if (!isRecord(value.started_at) || !exactFields(value.started_at, timeFields)) return invalid(); + const startedAtSeconds = value.started_at.seconds; + const startedAtNanoseconds = value.started_at.nanoseconds; + if ( + environmentId === null || allocationId === null || !isNonnegativeInteger(startedAtSeconds) || + !isNonnegativeInteger(startedAtNanoseconds) || startedAtNanoseconds > 999_999_999 || + startedAtSeconds > generatedAt || startedAtSeconds >= requestedEnd || + typeof value.provider_type !== "string" || !providerTypePattern.test(value.provider_type) || value.points.length > maximumPoints + ) return invalid(); + const points = value.points.map((point) => projectPoint( + point, requestedStart, requestedEnd, resolution, startedAtSeconds, invalid, + )); + if (!ordered(points)) return invalid(); + return { + environment_id: environmentId, + allocation_id: allocationId, + started_at: { seconds: startedAtSeconds, nanoseconds: startedAtNanoseconds }, + provider_type: value.provider_type, + points, + }; +} + +export function projectRuntimeHistory( + value: unknown, + expectedSessionId: string, + requested: RuntimeHistoryQuery, + invalid: InvalidRuntimeHistory, +): RuntimeHistory { + if ( + !isRecord(value) || !exactFields(value, historyFields) || value.object !== "agent.runtime_history" || value.source !== "durable" || + typeof value.session_id !== "string" || !sameResourceId(value.session_id, expectedSessionId) || + !isRecord(value.requested_range) || !exactFields(value.requested_range, rangeFields) || + value.requested_range.start !== requested.start || value.requested_range.end !== requested.end || + !positiveInteger(value.resolution_seconds) || !isNonnegativeInteger(value.generated_at) || + !isRecord(value.coverage) || !exactFields(value.coverage, coverageFields) || !Array.isArray(value.coverage.buckets) || + !Array.isArray(value.series) + ) return invalid(); + const sessionId = canonicalUuid(value.session_id); + if (sessionId === null) return invalid(); + const generatedAt = value.generated_at as number; + const maximumPoints = requested.maxPoints ?? 120; + if (value.coverage.buckets.length > maximumPoints || value.series.length > maximumSeries) return invalid(); + const retainedStart = value.coverage.retained_start; + const firstSample = nullableNonnegativeInteger(value.coverage.first_sample_at, invalid); + const lastSample = nullableNonnegativeInteger(value.coverage.last_sample_at, invalid); + if ( + generatedAt < requested.start || + !isNonnegativeInteger(retainedStart) || retainedStart < requested.start || retainedStart > requested.end || retainedStart > generatedAt || + !isNonnegativeInteger(value.coverage.sample_count) || !isNonnegativeInteger(value.coverage.expected_sample_count) || + ((firstSample === null) !== (lastSample === null)) || + (firstSample !== null && (firstSample < retainedStart || lastSample! < firstSample || lastSample! >= requested.end || lastSample! > generatedAt)) + ) return invalid(); + const buckets = value.coverage.buckets.map((point) => projectCoveragePoint( + point, coveragePointFields, requested.start, requested.end, value.resolution_seconds as number, invalid, + )); + if (!ordered(buckets)) return invalid(); + const sampleCount = buckets.reduce((sum, point) => sum + point.observation_count, 0); + const observedTimes = buckets.flatMap((point) => point.first_observed_at === null + ? [] + : [point.first_observed_at, point.last_observed_at as number]); + if ( + sampleCount !== value.coverage.sample_count || + (observedTimes.length === 0 && firstSample !== null) || + (observedTimes.length > 0 && ( + firstSample !== Math.min(...observedTimes) || lastSample !== Math.max(...observedTimes) + )) + ) return invalid(); + const series = value.series.map((entry) => projectSeries( + entry, requested.start, requested.end, value.resolution_seconds as number, generatedAt, maximumPoints, invalid, + )); + if ( + new Set(series.map((entry) => entry.environment_id)).size > 1 || + new Set(series.map((entry) => `${entry.allocation_id}\u0000${entry.started_at.seconds}\u0000${entry.started_at.nanoseconds}`)).size !== series.length || + buckets.length + series.reduce((sum, entry) => sum + entry.points.length, 0) > maximumTotalPoints || + series.some((entry) => entry.points.some((point) => point.first_observed_at !== null && point.first_observed_at < retainedStart)) || + series.some((entry) => entry.points.some((point) => point.last_observed_at !== null && point.last_observed_at > generatedAt)) + ) return invalid(); + return { + object: "agent.runtime_history", + source: "durable", + session_id: sessionId, + requested_range: { start: requested.start, end: requested.end }, + resolution_seconds: value.resolution_seconds, + generated_at: generatedAt, + coverage: { + retained_start: retainedStart, + first_sample_at: firstSample, + last_sample_at: lastSample, + sample_count: value.coverage.sample_count, + expected_sample_count: value.coverage.expected_sample_count, + buckets, + }, + series, + }; +} diff --git a/packages/agents-client/src/runtime-history.test.ts b/packages/agents-client/src/runtime-history.test.ts new file mode 100644 index 000000000..948933d50 --- /dev/null +++ b/packages/agents-client/src/runtime-history.test.ts @@ -0,0 +1,248 @@ +import { describe, expect, it } from "vitest"; + +import { OpenAIAgentsClient } from "./client"; + +const sessionId = "11111111-1111-4111-8111-111111111111"; +const environmentId = "22222222-2222-4222-8222-222222222222"; +const allocationId = "33333333-3333-4333-8333-333333333333"; + +interface FetchCall { + input: RequestInfo | URL; + init?: RequestInit; +} + +function jsonResponse(body: unknown): Response { + return new Response(JSON.stringify(body), { headers: { "Content-Type": "application/json" } }); +} + +function clientFor(body: unknown, calls: FetchCall[] = []): OpenAIAgentsClient { + return new OpenAIAgentsClient({ + baseUrl: "https://core.example/v1", + fetch: (async (input: RequestInfo | URL, init?: RequestInit) => { + calls.push({ input, init }); + return jsonResponse(body); + }) as typeof fetch, + }); +} + +function capabilities(overrides: Record = {}): Record { + return { + object: "agent.runtime_history_capabilities", + available: true, + reason: null, + collection_mode: "periodic", + sample_interval_seconds: 30, + retention_seconds: 604800, + minimum_step_seconds: 30, + maximum_range_seconds: 86400, + maximum_points: 1000, + metrics: ["cpu", "memory"], + ...overrides, + }; +} + +function history(overrides: Record = {}): Record { + const first = { + start: 1000, + end: 1060, + first_observed_at: 1010, + last_observed_at: 1010, + observation_count: 1, + observed_count: 1, + unavailable_count: 0, + }; + const second = { + start: 1060, + end: 1120, + first_observed_at: 1070, + last_observed_at: 1070, + observation_count: 1, + observed_count: 0, + unavailable_count: 1, + }; + return { + object: "agent.runtime_history", + source: "durable", + session_id: sessionId, + requested_range: { start: 1000, end: 1120 }, + resolution_seconds: 60, + generated_at: 1120, + coverage: { + retained_start: 1000, + first_sample_at: 1010, + last_sample_at: 1070, + sample_count: 2, + expected_sample_count: 4, + buckets: [first, second], + }, + series: [{ + environment_id: environmentId, + allocation_id: allocationId, + started_at: { seconds: 900, nanoseconds: 0 }, + provider_type: "docker", + points: [{ + ...first, + cpu: { contributor_count: 1, utilization_ratio: 0, capacity_cores: 2 }, + memory: { contributor_count: 1, usage_bytes: 0, limit_bytes: 2048 }, + }, { ...second, cpu: null, memory: null }], + }], + ...overrides, + }; +} + +describe("Runtime history client", () => { + it("retrieves strict capabilities without exposing backend details", async () => { + const calls: FetchCall[] = []; + await expect(clientFor(capabilities(), calls).getRuntimeHistoryCapabilities()).resolves.toEqual(capabilities()); + expect(String(calls[0]?.input)).toBe("https://core.example/v1/agents/runtime-history/capabilities"); + }); + + it("retrieves bounded Session history and preserves observed zeroes", async () => { + const calls: FetchCall[] = []; + const controller = new AbortController(); + const response = history(); + const incarnations = response.series as Array>; + const secondIncarnation = structuredClone(incarnations[0]!); + secondIncarnation.started_at = { seconds: 900, nanoseconds: 1 }; + incarnations.push(secondIncarnation); + const value = await clientFor(response, calls).retrieveRuntimeHistory(sessionId.toUpperCase(), { + start: 1000, + end: 1120, + maxPoints: 60, + signal: controller.signal, + }); + expect(String(calls[0]?.input)).toBe( + `https://core.example/v1/agents/sessions/${sessionId}/runtime-history?start=1000&end=1120&max_points=60`, + ); + expect(calls[0]?.init?.signal).toBe(controller.signal); + expect(value.series[0]?.points[0]?.cpu?.utilization_ratio).toBe(0); + expect(value.series[0]?.points[0]?.memory?.usage_bytes).toBe(0); + expect(value.series.map((series) => series.started_at)).toEqual([ + { seconds: 900, nanoseconds: 0 }, + { seconds: 900, nanoseconds: 1 }, + ]); + expect(value.coverage.sample_count).toBe(2); + }); + + it("rejects invalid requests before fetch", async () => { + const calls: FetchCall[] = []; + const client = clientFor(history(), calls); + for (const [id, query] of [ + ["not-a-session", { start: 1000, end: 1120 }], + [sessionId, { start: 1120, end: 1000 }], + [sessionId, { start: -1, end: 1000 }], + [sessionId, { start: 1000, end: 1120, maxPoints: 1 }], + [sessionId, { start: 1000, end: 1120, maxPoints: 10001 }], + ] as const) { + await expect(client.retrieveRuntimeHistory(id, query)).rejects.toThrow(TypeError); + } + expect(calls).toHaveLength(0); + }); + + for (const [name, body] of [ + ["unknown field", capabilities({ backend: "clickhouse" })], + ["inconsistent availability", capabilities({ available: false })], + ["duplicate metrics", capabilities({ metrics: ["cpu", "cpu"] })], + ["invalid retention", capabilities({ retention_seconds: 60, maximum_range_seconds: 120 })], + ["unsafe maximum", capabilities({ maximum_points: 10001 })], + ["unqualified periodic", capabilities({ sample_interval_seconds: null })], + ] as const) { + it(`rejects ${name} in capabilities`, async () => { + await expect(clientFor(body).getRuntimeHistoryCapabilities()).rejects.toMatchObject({ + status: 502, + code: "invalid_runtime_history_capabilities", + }); + }); + } + + for (const [name, mutate] of [ + ["unknown field", (value: Record) => { value.tokens = 10; }], + ["foreign Session", (value: Record) => { value.session_id = "55555555-5555-4555-8555-555555555555"; }], + ["mismatched range", (value: Record) => { value.requested_range = { start: 999, end: 1120 }; }], + ["empty result generated before the requested range", (value: Record) => { + value.generated_at = 999; + value.coverage = { + retained_start: 1000, first_sample_at: null, last_sample_at: null, + sample_count: 0, expected_sample_count: 0, buckets: [], + }; + value.series = []; + }], + ["empty result retained after generation", (value: Record) => { + value.generated_at = 1050; + value.coverage = { + retained_start: 1060, first_sample_at: null, last_sample_at: null, + sample_count: 0, expected_sample_count: 0, buckets: [], + }; + value.series = []; + }], + ["empty incarnation newer than generation", (value: Record) => { + value.generated_at = 1050; + value.coverage = { + retained_start: 1000, first_sample_at: null, last_sample_at: null, + sample_count: 0, expected_sample_count: 0, buckets: [], + }; + const series = (value.series as Array>)[0]!; + series.started_at = { seconds: 1051, nanoseconds: 0 }; + series.points = []; + }], + ["lossy incarnation fence", (value: Record) => { + (value.series as Array>)[0]!.started_at = 900; + }], + ["invalid incarnation nanoseconds", (value: Record) => { + (value.series as Array>)[0]!.started_at = { seconds: 900, nanoseconds: 1_000_000_000 }; + }], + ["bucket ends at incarnation", (value: Record) => { + const series = (value.series as Array<{ started_at: unknown; points: Array> }>)[0]!; + series.started_at = { seconds: 1060, nanoseconds: 0 }; + series.points[0] = { + start: 1000, end: 1060, first_observed_at: null, last_observed_at: null, + observation_count: 0, observed_count: 0, unavailable_count: 0, cpu: null, memory: null, + }; + }], + ["noncanonical allocation", (value: Record) => { + (value.series as Array>)[0]!.allocation_id = `urn:uuid:${allocationId}`; + }], + ["incorrect coverage total", (value: Record) => { (value.coverage as Record).sample_count = 3; }], + ["series predates retention", (value: Record) => { + const coverage = value.coverage as { retained_start: number; first_sample_at: number; last_sample_at: number; sample_count: number; buckets: Array> }; + coverage.retained_start = 1020; + coverage.first_sample_at = 1070; + coverage.last_sample_at = 1070; + coverage.sample_count = 1; + coverage.buckets = [coverage.buckets[1]!]; + }], + ["overlapping buckets", (value: Record) => { + const buckets = (value.coverage as { buckets: Array> }).buckets; + buckets[1]!.start = 1050; + }], + ["zero CPU capacity", (value: Record) => { + const point = (value.series as Array<{ points: Array> }>)[0]!.points[0]!; + point.cpu = { contributor_count: 1, utilization_ratio: null, capacity_cores: 0 }; + }], + ["unsafe memory", (value: Record) => { + const point = (value.series as Array<{ points: Array> }>)[0]!.points[0]!; + point.memory = { contributor_count: 1, usage_bytes: Number.MAX_SAFE_INTEGER + 1, limit_bytes: 2048 }; + }], + ["duplicate incarnation", (value: Record) => { + const series = value.series as Array>; + series.push(structuredClone(series[0]!)); + }], + ["mixed Environment scope", (value: Record) => { + const series = value.series as Array>; + const foreign = structuredClone(series[0]!); + foreign.environment_id = "44444444-4444-4444-8444-444444444444"; + foreign.allocation_id = "55555555-5555-4555-8555-555555555555"; + series.push(foreign); + }], + ] as const) { + it(`rejects ${name} in history`, async () => { + const value = history(); + mutate(value); + await expect(clientFor(value).retrieveRuntimeHistory(sessionId, { + start: 1000, + end: 1120, + maxPoints: 60, + })).rejects.toMatchObject({ status: 502, code: "invalid_runtime_history" }); + }); + } +}); diff --git a/packages/agents-client/src/types.ts b/packages/agents-client/src/types.ts index 8d38d2a57..d93ab5938 100644 --- a/packages/agents-client/src/types.ts +++ b/packages/agents-client/src/types.ts @@ -828,6 +828,98 @@ export interface RuntimeObservationList extends ListPage { last_id: string | null; } +export type RuntimeHistoryCollectionMode = "on_read" | "periodic"; +export type RuntimeHistoryCapabilityReason = "not_configured" | "periodic_collection_required"; +export type RuntimeHistoryMetric = "cpu" | "memory"; + +export interface RuntimeHistoryCapabilities { + object: "agent.runtime_history_capabilities"; + available: boolean; + reason: RuntimeHistoryCapabilityReason | null; + collection_mode: RuntimeHistoryCollectionMode | null; + sample_interval_seconds: number | null; + retention_seconds: number | null; + minimum_step_seconds: number | null; + maximum_range_seconds: number | null; + maximum_points: number | null; + metrics: RuntimeHistoryMetric[]; +} + +export interface RuntimeHistoryQuery extends ReadOptions { + /** Inclusive Unix-second boundary. */ + start: number; + /** Exclusive Unix-second boundary. */ + end: number; + /** Requested maximum buckets per series. Core selects the effective resolution. */ + maxPoints?: number; +} + +export interface RuntimeHistoryRange { + start: number; + end: number; +} + +export interface RuntimeHistoryCoveragePoint { + start: number; + end: number; + first_observed_at: number | null; + last_observed_at: number | null; + observation_count: number; + observed_count: number; + unavailable_count: number; +} + +export interface RuntimeHistoryCPU { + contributor_count: number; + utilization_ratio: number | null; + capacity_cores: number | null; +} + +export interface RuntimeHistoryMemory { + contributor_count: number; + usage_bytes: number | null; + limit_bytes: number | null; +} + +export interface RuntimeHistoryPoint extends RuntimeHistoryCoveragePoint { + cpu: RuntimeHistoryCPU | null; + memory: RuntimeHistoryMemory | null; +} + +export interface RuntimeHistorySeries { + environment_id: string; + allocation_id: string; + /** Lossless compute-incarnation identity fence. */ + started_at: RuntimeHistoryTime; + provider_type: string; + points: RuntimeHistoryPoint[]; +} + +export interface RuntimeHistoryTime { + seconds: number; + nanoseconds: number; +} + +export interface RuntimeHistoryCoverage { + retained_start: number; + first_sample_at: number | null; + last_sample_at: number | null; + sample_count: number; + expected_sample_count: number; + buckets: RuntimeHistoryCoveragePoint[]; +} + +export interface RuntimeHistory { + object: "agent.runtime_history"; + source: "durable"; + session_id: string; + requested_range: RuntimeHistoryRange; + resolution_seconds: number; + generated_at: number; + coverage: RuntimeHistoryCoverage; + series: RuntimeHistorySeries[]; +} + export interface AgentCore { listAgents(options?: PageOptions): Promise>; createAgent(input: CreateAgentInput): Promise; @@ -846,6 +938,8 @@ export interface AgentCore { listSessions(options?: PageOptions & { agentId?: string }): Promise>; listRuntimeObservations(options?: PageOptions): Promise; retrieveRuntimeObservation(sessionId: string, options?: ReadOptions): Promise; + getRuntimeHistoryCapabilities(options?: ReadOptions): Promise; + retrieveRuntimeHistory(sessionId: string, query: RuntimeHistoryQuery): Promise; createSession(input: CreateSessionInput, idempotencyKey?: string): Promise; createSessionStream( input: Omit, diff --git a/scripts/patch-agents-openapi.py b/scripts/patch-agents-openapi.py new file mode 100644 index 000000000..7064da65c --- /dev/null +++ b/scripts/patch-agents-openapi.py @@ -0,0 +1,42 @@ +#!/usr/bin/env python3 +"""Apply schema constraints that swag cannot express on response fields.""" + +from pathlib import Path +import sys + + +PROVIDER_TYPE_SCHEMA = """ provider_type: + type: string +""" +BOUNDED_PROVIDER_TYPE_SCHEMA = """ provider_type: + pattern: '^[a-z][a-z0-9_]{0,31}$' + type: string +""" + + +def main() -> None: + if len(sys.argv) != 2: + raise SystemExit("usage: patch-agents-openapi.py OPENAPI_YAML") + + path = Path(sys.argv[1]) + document = path.read_text(encoding="utf-8") + series_start = " v1.RuntimeHistorySeries:\n" + series_end = "\n v1.RuntimeHistoryTime:\n" + if document.count(series_start) != 1 or document.count(series_end) != 1: + raise SystemExit("expected exactly one Runtime history series schema") + before, remainder = document.split(series_start) + series, after = remainder.split(series_end) + if series.count(PROVIDER_TYPE_SCHEMA) != 1: + raise SystemExit("expected exactly one Runtime history provider_type field") + path.write_text( + before + + series_start + + series.replace(PROVIDER_TYPE_SCHEMA, BOUNDED_PROVIDER_TYPE_SCHEMA) + + series_end + + after, + encoding="utf-8", + ) + + +if __name__ == "__main__": + main() diff --git a/services/agents-api/internal/api/handler.go b/services/agents-api/internal/api/handler.go index 3d923eb6a..cab226b43 100644 --- a/services/agents-api/internal/api/handler.go +++ b/services/agents-api/internal/api/handler.go @@ -50,6 +50,7 @@ type Handler struct { artifacts SessionArtifactStore subagents SubagentStore runtimeObservations RuntimeObservationService + runtimeHistory RuntimeHistoryService } func NewHandler(s ResourceStore, auth *Authenticator, engine string, options ...Option) (http.Handler, error) { @@ -103,6 +104,8 @@ func NewHandler(s ResourceStore, auth *Authenticator, engine string, options ... r.Get("/agents/sessions/{session_id}", h.getSession) r.Get("/agents/sessions/{session_id}/runtime-observation", h.getRuntimeObservation) r.Get("/agents/runtime-observations", h.listRuntimeObservations) + r.Get("/agents/runtime-history/capabilities", h.getRuntimeHistoryCapabilities) + r.Get("/agents/sessions/{session_id}/runtime-history", h.getRuntimeHistory) r.Post("/agents/sessions/{session_id}", h.updateSession) r.Delete("/agents/sessions/{session_id}", h.deleteSession) r.Post("/agents/sessions/{session_id}/events", h.createEvents) diff --git a/services/agents-api/internal/api/runtime_history.go b/services/agents-api/internal/api/runtime_history.go new file mode 100644 index 000000000..b99c7a596 --- /dev/null +++ b/services/agents-api/internal/api/runtime_history.go @@ -0,0 +1,247 @@ +package api + +import ( + "context" + "errors" + "net/http" + "strconv" + "time" + + v1 "github.com/MiniMax-AI-Dev/parsar/contracts/agents-api/v1" + "github.com/MiniMax-AI-Dev/parsar/internal/obs/log" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/go-chi/chi/v5" +) + +const ( + defaultRuntimeHistoryPoints = 120 + runtimeHistoryRequestBudget = 15 * time.Second +) + +type RuntimeHistoryService interface { + Capabilities() runtimehistory.Capabilities + QuerySession(context.Context, string, string, runtimehistory.Range) (runtimehistory.Response, error) +} + +func WithRuntimeHistory(service RuntimeHistoryService) Option { + return func(h *Handler) { h.runtimeHistory = service } +} + +// @Summary Retrieve Runtime history capabilities +// @Description Core extension advertising only safe backend-neutral Durable history availability and bounds. Available is true only when a query Reader and qualified periodic collection are both configured. +// @Tags Runtime history +// @Produce json +// @Security BearerAuth +// @Param OpenAI-Beta header string true "agents=v1" +// @Success 200 {object} v1.RuntimeHistoryCapabilities +// @Failure 400,401 {object} v1.ErrorResponse +// @Router /agents/runtime-history/capabilities [get] +func (h *Handler) getRuntimeHistoryCapabilities(w http.ResponseWriter, r *http.Request) { + if len(r.URL.Query()) != 0 { + writeError(w, http.StatusBadRequest, "unsupported_parameter", "Runtime history capabilities do not accept query parameters.") + return + } + writeJSON(w, http.StatusOK, runtimeHistoryCapabilitiesResponse(h.runtimeHistory)) +} + +// @Summary Retrieve Session Runtime history +// @Description Core extension returning tenant-scoped stored Runtime observations for one Session. End is exclusive; the server selects a bounded resolution. Responses contain at most 1,000 series, 10,000 points per coverage/series array, and 100,000 total coverage plus series points. It never reads or changes live compute. +// @Tags Runtime history +// @Produce json +// @Security BearerAuth +// @Param OpenAI-Beta header string true "agents=v1" +// @Param session_id path string true "Session ID" +// @Param start query integer true "Inclusive Unix-second start" minimum(0) maximum(9007199254740991) +// @Param end query integer true "Exclusive Unix-second end" minimum(1) maximum(9007199254740991) +// @Param max_points query integer false "Maximum points per series; defaults to the lower of 120 and the advertised service maximum" minimum(2) maximum(10000) +// @Success 200 {object} v1.RuntimeHistory +// @Failure 400,401,404,409,503 {object} v1.ErrorResponse +// @Router /agents/sessions/{session_id}/runtime-history [get] +func (h *Handler) getRuntimeHistory(w http.ResponseWriter, r *http.Request) { + if h.runtimeHistory == nil { + writeError(w, http.StatusServiceUnavailable, "runtime_history_unavailable", "Durable Runtime history is not configured on this service.") + return + } + capabilities := h.runtimeHistory.Capabilities() + if capabilities.Validate() != nil || !capabilities.Durable() { + writeError(w, http.StatusServiceUnavailable, "runtime_history_unavailable", "Durable Runtime history is not configured on this service.") + return + } + requested, ok := readRuntimeHistoryRange(w, r, capabilities, time.Now().UTC()) + if !ok { + return + } + expectedTenantID := tenantID(r) + expectedSessionID := chi.URLParam(r, "session_id") + ctx, cancel := context.WithTimeout(r.Context(), runtimeHistoryRequestBudget) + defer cancel() + value, err := h.runtimeHistory.QuerySession(ctx, expectedTenantID, expectedSessionID, requested) + if err != nil { + switch { + case errors.Is(err, runtimehistory.ErrInvalidRange): + writeError(w, http.StatusBadRequest, "invalid_request", "Runtime history range is invalid.") + case errors.Is(err, runtimehistory.ErrUnsupported): + writeError(w, http.StatusConflict, "runtime_history_unsupported", "Runtime history is not supported for this Session.") + case errors.Is(err, runtimehistory.ErrUnavailable), errors.Is(err, runtimehistory.ErrInvalidResult), errors.Is(err, context.DeadlineExceeded): + log.Ctx(r.Context()).Warn("Runtime history query unavailable") + writeError(w, http.StatusServiceUnavailable, "runtime_history_unavailable", "Durable Runtime history is temporarily unavailable.") + default: + writeStoreError(w, r, err) + } + return + } + response, err := runtimeHistoryResponse(value, expectedTenantID, expectedSessionID, requested) + if err != nil { + log.Ctx(r.Context()).Error("Runtime history response validation failed") + writeError(w, http.StatusServiceUnavailable, "runtime_history_unavailable", "Durable Runtime history is temporarily unavailable.") + return + } + writeJSON(w, http.StatusOK, response) +} + +func readRuntimeHistoryRange(w http.ResponseWriter, r *http.Request, capabilities runtimehistory.Capabilities, now time.Time) (runtimehistory.Range, bool) { + query := r.URL.Query() + for key, values := range query { + if key != "start" && key != "end" && key != "max_points" || len(values) != 1 { + writeError(w, http.StatusBadRequest, "unsupported_parameter", "Runtime history accepts one start, end and optional max_points value.") + return runtimehistory.Range{}, false + } + } + start, startErr := strconv.ParseInt(query.Get("start"), 10, 64) + end, endErr := strconv.ParseInt(query.Get("end"), 10, 64) + points := defaultRuntimeHistoryPoints + if capabilities.MaximumPoints < points { + points = capabilities.MaximumPoints + } + if raw := query.Get("max_points"); raw != "" { + var err error + points, err = strconv.Atoi(raw) + if err != nil || points < 2 || points > capabilities.MaximumPoints { + writeError(w, http.StatusBadRequest, "invalid_request", "Runtime history range is invalid.") + return runtimehistory.Range{}, false + } + } + if startErr != nil || endErr != nil || start < 0 || end <= start || end > now.Unix()+1 || end-start > int64(capabilities.MaximumRange/time.Second) { + writeError(w, http.StatusBadRequest, "invalid_request", "Runtime history range is invalid.") + return runtimehistory.Range{}, false + } + return runtimehistory.Range{Start: time.Unix(start, 0).UTC(), End: time.Unix(end, 0).UTC(), MaxPoints: points}, true +} + +func runtimeHistoryCapabilitiesResponse(service RuntimeHistoryService) v1.RuntimeHistoryCapabilities { + response := v1.RuntimeHistoryCapabilities{Object: "agent.runtime_history_capabilities", Metrics: []string{}} + if service == nil { + reason := "not_configured" + response.Reason = &reason + return response + } + capabilities := service.Capabilities() + if capabilities.Validate() != nil { + reason := "not_configured" + response.Reason = &reason + return response + } + mode := string(capabilities.CollectionMode) + retention := int64(capabilities.Retention / time.Second) + minimumStep := int64(capabilities.MinimumStep / time.Second) + maximumRange := int64(capabilities.MaximumRange / time.Second) + maximumPoints := capabilities.MaximumPoints + response.CollectionMode = &mode + response.RetentionSeconds = &retention + response.MinimumStepSeconds = &minimumStep + response.MaximumRangeSeconds = &maximumRange + response.MaximumPoints = &maximumPoints + response.Metrics = make([]string, len(capabilities.Metrics)) + for index, metric := range capabilities.Metrics { + response.Metrics[index] = string(metric) + } + if capabilities.SampleInterval > 0 { + seconds := int64(capabilities.SampleInterval / time.Second) + response.SampleIntervalSeconds = &seconds + } + response.Available = capabilities.Durable() + if !response.Available { + reason := "periodic_collection_required" + response.Reason = &reason + } + return response +} + +func runtimeHistoryResponse(value runtimehistory.Response, expectedTenantID, expectedSessionID string, expectedRange runtimehistory.Range) (v1.RuntimeHistory, error) { + if err := value.Validate(time.Now().UTC()); err != nil || !value.Durable() || + value.TenantID != expectedTenantID || value.SessionID != expectedSessionID || + !value.Requested.Start.Equal(expectedRange.Start) || !value.Requested.End.Equal(expectedRange.End) || value.Requested.MaxPoints != expectedRange.MaxPoints { + return v1.RuntimeHistory{}, runtimehistory.ErrInvalidResult + } + retainedStart := value.Requested.Start + configuredStart := value.GeneratedAt.Add(-value.Retention) + if configuredStart.After(retainedStart) { + retainedStart = configuredStart + } + if value.RetainedFrom != nil && value.RetainedFrom.After(retainedStart) { + retainedStart = *value.RetainedFrom + } + if retainedStart.After(value.Requested.End) { + retainedStart = value.Requested.End + } + coverage := v1.RuntimeHistoryCoverage{RetainedStart: retainedStart.Unix(), Buckets: make([]v1.RuntimeHistoryCoveragePoint, 0, len(value.Coverage))} + if retainedStart.Before(value.Requested.End) { + duration := value.Requested.End.Sub(retainedStart) + coverage.ExpectedSampleCount = int64(duration / value.SampleInterval) + if duration%value.SampleInterval != 0 { + coverage.ExpectedSampleCount++ + } + } + for _, point := range value.Coverage { + converted := v1.RuntimeHistoryCoveragePoint{ + Start: point.Start.Unix(), End: point.End.Unix(), ObservationCount: point.ObservationCount, + ObservedCount: point.ObservedCount, UnavailableCount: point.UnavailableCount, + } + if point.ObservationCount > 0 { + first, last := point.FirstObservedAt.Unix(), point.LastObservedAt.Unix() + converted.FirstObservedAt, converted.LastObservedAt = &first, &last + if coverage.FirstSampleAt == nil || first < *coverage.FirstSampleAt { + value := first + coverage.FirstSampleAt = &value + } + if coverage.LastSampleAt == nil || last > *coverage.LastSampleAt { + value := last + coverage.LastSampleAt = &value + } + } + coverage.SampleCount += int64(point.ObservationCount) + coverage.Buckets = append(coverage.Buckets, converted) + } + response := v1.RuntimeHistory{ + Object: "agent.runtime_history", Source: "durable", SessionID: value.SessionID, + RequestedRange: v1.RuntimeHistoryRange{Start: value.Requested.Start.Unix(), End: value.Requested.End.Unix()}, + ResolutionSeconds: int64(value.Resolution / time.Second), GeneratedAt: value.GeneratedAt.Unix(), Coverage: coverage, + Series: make([]v1.RuntimeHistorySeries, 0, len(value.Series)), + } + for _, series := range value.Series { + converted := v1.RuntimeHistorySeries{ + EnvironmentID: series.EnvironmentID, AllocationID: series.AllocationID, + StartedAt: v1.RuntimeHistoryTime{Seconds: series.StartedAt.Unix(), Nanoseconds: series.StartedAt.Nanosecond()}, ProviderType: series.ProviderType, + Points: make([]v1.RuntimeHistoryPoint, 0, len(series.Points)), + } + for _, point := range series.Points { + item := v1.RuntimeHistoryPoint{ + Start: point.Start.Unix(), End: point.End.Unix(), ObservationCount: point.ObservationCount, + ObservedCount: point.ObservedCount, UnavailableCount: point.UnavailableCount, + } + if point.ObservationCount > 0 { + first, last := point.FirstObservedAt.Unix(), point.LastObservedAt.Unix() + item.FirstObservedAt, item.LastObservedAt = &first, &last + } + if point.CPUContributorCount > 0 { + item.CPU = &v1.RuntimeHistoryCPU{ContributorCount: point.CPUContributorCount, UtilizationRatio: point.CPUUtilizationRatio, CapacityCores: point.CPUCapacityCores} + } + if point.MemoryContributorCount > 0 { + item.Memory = &v1.RuntimeHistoryMemory{ContributorCount: point.MemoryContributorCount, UsageBytes: point.MemoryUsageBytes, LimitBytes: point.MemoryLimitBytes} + } + converted.Points = append(converted.Points, item) + } + response.Series = append(response.Series, converted) + } + return response, nil +} diff --git a/services/agents-api/internal/api/runtime_history_test.go b/services/agents-api/internal/api/runtime_history_test.go new file mode 100644 index 000000000..3ff2532eb --- /dev/null +++ b/services/agents-api/internal/api/runtime_history_test.go @@ -0,0 +1,250 @@ +package api + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + "testing" + "time" + + v1 "github.com/MiniMax-AI-Dev/parsar/contracts/agents-api/v1" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" + "github.com/google/uuid" +) + +type runtimeHistoryFixture struct { + capabilities runtimehistory.Capabilities + response runtimehistory.Response + err error + tenant string + session string + requested runtimehistory.Range + calls int +} + +func (f *runtimeHistoryFixture) Capabilities() runtimehistory.Capabilities { return f.capabilities } + +func (f *runtimeHistoryFixture) QuerySession(_ context.Context, tenant, session string, requested runtimehistory.Range) (runtimehistory.Response, error) { + f.calls++ + f.tenant, f.session, f.requested = tenant, session, requested + return f.response, f.err +} + +func historyCapabilities(mode runtimehistory.CollectionMode) runtimehistory.Capabilities { + value := runtimehistory.Capabilities{ + CollectionMode: mode, Retention: 7 * 24 * time.Hour, MinimumStep: 30 * time.Second, + MaximumRange: 24 * time.Hour, MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory}, + } + if mode == runtimehistory.CollectionPeriodic { + value.SampleInterval = 30 * time.Second + } + return value +} + +func TestRuntimeHistoryCapabilitiesAreSafeAndDisabledByDefault(t *testing.T) { + handler, _, _ := testHandler(t) + response := runtimeObservationRequest(handler, "/v1/agents/runtime-history/capabilities") + if response.Code != http.StatusOK { + t.Fatalf("capabilities returned %d: %s", response.Code, response.Body) + } + var value v1.RuntimeHistoryCapabilities + if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history_capabilities" || value.Available || value.Reason == nil || *value.Reason != "not_configured" || value.CollectionMode != nil || value.SampleIntervalSeconds != nil || value.RetentionSeconds != nil || value.MaximumPoints != nil || value.Metrics == nil || len(value.Metrics) != 0 { + t.Fatalf("unsafe disabled capabilities: %s", response.Body) + } + invalid := runtimeObservationRequest(handler, "/v1/agents/runtime-history/capabilities?backend=clickhouse") + if invalid.Code != http.StatusBadRequest { + t.Fatalf("capability query was accepted: %d %s", invalid.Code, invalid.Body) + } +} + +func TestRuntimeHistoryRequiresQualifiedPeriodicCollection(t *testing.T) { + service := &runtimeHistoryFixture{capabilities: historyCapabilities(runtimehistory.CollectionOnRead)} + handler, _, _ := testHandler(t, WithRuntimeHistory(service)) + capabilityResponse := runtimeObservationRequest(handler, "/v1/agents/runtime-history/capabilities") + var capabilities v1.RuntimeHistoryCapabilities + if capabilityResponse.Code != http.StatusOK || json.Unmarshal(capabilityResponse.Body.Bytes(), &capabilities) != nil || capabilities.Available || capabilities.Reason == nil || *capabilities.Reason != "periodic_collection_required" || capabilities.CollectionMode == nil || *capabilities.CollectionMode != "on_read" || capabilities.SampleIntervalSeconds != nil { + t.Fatalf("on-read capability was advertised as durable: %d %s", capabilityResponse.Code, capabilityResponse.Body) + } + response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+uuid.NewString()+"/runtime-history?start=1&end=2") + if response.Code != http.StatusServiceUnavailable || service.calls != 0 { + t.Fatalf("on-read history reached query service: %d calls=%d body=%s", response.Code, service.calls, response.Body) + } +} + +func TestRuntimeHistoryFailsClosedForMalformedCapabilities(t *testing.T) { + service := &runtimeHistoryFixture{capabilities: historyCapabilities(runtimehistory.CollectionPeriodic)} + service.capabilities.Retention = 0 + handler, _, _ := testHandler(t, WithRuntimeHistory(service)) + capabilityResponse := runtimeObservationRequest(handler, "/v1/agents/runtime-history/capabilities") + var capabilities v1.RuntimeHistoryCapabilities + if capabilityResponse.Code != http.StatusOK || json.Unmarshal(capabilityResponse.Body.Bytes(), &capabilities) != nil || capabilities.Available || capabilities.Reason == nil || *capabilities.Reason != "not_configured" || capabilities.CollectionMode != nil || len(capabilities.Metrics) != 0 { + t.Fatalf("malformed capabilities did not fail closed: %d %s", capabilityResponse.Code, capabilityResponse.Body) + } + response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+uuid.NewString()+"/runtime-history?start=1&end=2") + if response.Code != http.StatusServiceUnavailable || service.calls != 0 { + t.Fatalf("malformed capabilities reached query service: %d calls=%d body=%s", response.Code, service.calls, response.Body) + } +} + +func TestRuntimeHistoryRouteBindsAuthenticatedSessionAndPreservesCoverage(t *testing.T) { + now := time.Now().UTC().Truncate(time.Second) + start := now.Add(-time.Hour) + sessionID, environmentID, allocationID := uuid.NewString(), uuid.NewString(), uuid.NewString() + scope := runtimehistory.Scope{SessionID: sessionID, EnvironmentID: environmentID} + zeroRatio := float64(0) + capacity := 2.0 + zeroMemory := uint64(0) + limit := uint64(2048) + coveragePoint := runtimehistory.CoveragePoint{ + Start: start, End: start.Add(time.Minute), FirstObservedAt: start.Add(10 * time.Second), LastObservedAt: start.Add(10 * time.Second), + ObservationCount: 1, ObservedCount: 1, + } + resourcePoint := runtimehistory.Point{ + Start: start, End: start.Add(time.Minute), FirstObservedAt: start.Add(10 * time.Second), LastObservedAt: start.Add(10 * time.Second), + ObservationCount: 1, ObservedCount: 1, CPUContributorCount: 1, MemoryContributorCount: 1, + CPUUtilizationRatio: &zeroRatio, CPUCapacityCores: &capacity, MemoryUsageBytes: &zeroMemory, MemoryLimitBytes: &limit, + } + service := &runtimeHistoryFixture{capabilities: historyCapabilities(runtimehistory.CollectionPeriodic)} + handler, _, tenant := testHandler(t, WithRuntimeHistory(service)) + scope.TenantID = tenant + service.response = runtimehistory.Response{ + Capabilities: service.capabilities, Scope: scope, + Requested: runtimehistory.Range{Start: start, End: now, MaxPoints: 60}, Resolution: time.Minute, GeneratedAt: now, RetainedFrom: &start, + Coverage: []runtimehistory.CoveragePoint{coveragePoint}, + Series: []runtimehistory.Series{{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker", Points: []runtimehistory.Point{resourcePoint}}}, + } + secondIncarnation := service.response.Series[0] + secondIncarnation.StartedAt = secondIncarnation.StartedAt.Add(time.Nanosecond) + service.response.Series = append(service.response.Series, secondIncarnation) + response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+sessionID+"/runtime-history?start="+timeString(start)+"&end="+timeString(now)+"&max_points=60") + if response.Code != http.StatusOK { + t.Fatalf("history returned %d: %s", response.Code, response.Body) + } + var value v1.RuntimeHistory + if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history" || value.Source != "durable" || value.SessionID != sessionID || value.ResolutionSeconds != 60 || value.Coverage.SampleCount != 1 || value.Coverage.ExpectedSampleCount != 120 || len(value.Coverage.Buckets) != 1 || len(value.Series) != 2 || len(value.Series[0].Points) != 1 || value.Series[0].StartedAt.Seconds != value.Series[1].StartedAt.Seconds || value.Series[0].StartedAt.Nanoseconds == value.Series[1].StartedAt.Nanoseconds { + t.Fatalf("invalid history response: %s", response.Body) + } + point := value.Series[0].Points[0] + if point.CPU == nil || point.CPU.UtilizationRatio == nil || *point.CPU.UtilizationRatio != 0 || point.Memory == nil || point.Memory.UsageBytes == nil || *point.Memory.UsageBytes != 0 { + t.Fatalf("observed zero was lost: %+v", point) + } + if service.calls != 1 || service.tenant != tenant || service.session != sessionID || !service.requested.Start.Equal(start) || !service.requested.End.Equal(now) || service.requested.MaxPoints != 60 { + t.Fatalf("history authority/range mismatch: %+v", service) + } +} + +func TestRuntimeHistoryRejectsUnsafeQueriesAndFailures(t *testing.T) { + service := &runtimeHistoryFixture{capabilities: historyCapabilities(runtimehistory.CollectionPeriodic)} + handler, _, _ := testHandler(t, WithRuntimeHistory(service)) + sessionID := uuid.NewString() + for _, query := range []string{ + "", "?start=1", "?start=2&end=1", "?start=x&end=2", "?start=1&end=2&provider=docker", "?start=1&start=1&end=2", "?start=1&end=2&max_points=x", "?start=1&end=2&max_points=1", "?start=1&end=2&max_points=1001", "?start=1&end=90002", "?start=1&end=9223372036854775807", + } { + response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+sessionID+"/runtime-history"+query) + if response.Code != http.StatusBadRequest { + t.Fatalf("unsafe query %q returned %d: %s", query, response.Code, response.Body) + } + } + if service.calls != 0 { + t.Fatalf("unsafe query reached history service %d times", service.calls) + } + + now := time.Now().UTC().Truncate(time.Second) + path := "/v1/agents/sessions/" + sessionID + "/runtime-history?start=" + timeString(now.Add(-time.Hour)) + "&end=" + timeString(now) + for _, tc := range []struct { + err error + code int + }{ + {runtimehistory.ErrInvalidRange, http.StatusBadRequest}, + {runtimehistory.ErrUnsupported, http.StatusConflict}, + {runtimehistory.ErrUnavailable, http.StatusServiceUnavailable}, + {runtimehistory.ErrInvalidResult, http.StatusServiceUnavailable}, + {store.ErrNotFound, http.StatusNotFound}, + } { + service.err = tc.err + response := runtimeObservationRequest(handler, path) + if response.Code != tc.code { + t.Fatalf("history error %v returned %d: %s", tc.err, response.Code, response.Body) + } + } + service.err = nil + service.response = runtimehistory.Response{} + response := runtimeObservationRequest(handler, path) + if response.Code != http.StatusServiceUnavailable { + t.Fatalf("malformed history response returned %d: %s", response.Code, response.Body) + } +} + +func TestRuntimeHistoryRejectsMismatchedServiceResponses(t *testing.T) { + now := time.Now().UTC().Truncate(time.Second) + start := now.Add(-time.Hour) + sessionID := uuid.NewString() + service := &runtimeHistoryFixture{capabilities: historyCapabilities(runtimehistory.CollectionPeriodic)} + handler, _, tenant := testHandler(t, WithRuntimeHistory(service)) + base := runtimehistory.Response{ + Capabilities: service.capabilities, + Scope: runtimehistory.Scope{ + TenantID: tenant, SessionID: sessionID, EnvironmentID: uuid.NewString(), + }, + Requested: runtimehistory.Range{Start: start, End: now, MaxPoints: 60}, + Resolution: time.Minute, GeneratedAt: now, + } + path := "/v1/agents/sessions/" + sessionID + "/runtime-history?start=" + timeString(start) + "&end=" + timeString(now) + "&max_points=60" + + for name, mutate := range map[string]func(*runtimehistory.Response){ + "tenant": func(value *runtimehistory.Response) { value.TenantID = uuid.NewString() }, + "Session": func(value *runtimehistory.Response) { value.SessionID = uuid.NewString() }, + "range": func(value *runtimehistory.Response) { + value.Requested.Start = value.Requested.Start.Add(30 * time.Second) + }, + } { + t.Run(name, func(t *testing.T) { + service.response = base + mutate(&service.response) + response := runtimeObservationRequest(handler, path) + if response.Code != http.StatusServiceUnavailable { + t.Fatalf("mismatched %s response returned %d: %s", name, response.Code, response.Body) + } + }) + } +} + +func TestRuntimeHistoryDefaultPointBudgetRespectsCapabilities(t *testing.T) { + capabilities := historyCapabilities(runtimehistory.CollectionPeriodic) + capabilities.MaximumPoints = 60 + service := &runtimeHistoryFixture{capabilities: capabilities, err: runtimehistory.ErrUnavailable} + handler, _, _ := testHandler(t, WithRuntimeHistory(service)) + now := time.Now().UTC().Truncate(time.Second) + response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+uuid.NewString()+"/runtime-history?start="+timeString(now.Add(-time.Hour))+"&end="+timeString(now)) + if response.Code != http.StatusServiceUnavailable || service.calls != 1 || service.requested.MaxPoints != 60 { + t.Fatalf("default point budget ignored capabilities: status=%d calls=%d requested=%+v", response.Code, service.calls, service.requested) + } +} + +func TestRuntimeHistoryExpectedCoverageUsesOverflowSafeCeiling(t *testing.T) { + now := time.Now().UTC().Truncate(time.Second) + start := now.Add(-time.Hour) + huge := (time.Duration(1<<63-1) / time.Second) * time.Second + capabilities := historyCapabilities(runtimehistory.CollectionPeriodic) + capabilities.SampleInterval = huge + capabilities.Retention = huge + capabilities.MaximumRange = time.Hour + value := runtimehistory.Response{ + Capabilities: capabilities, + Scope: runtimehistory.Scope{ + TenantID: uuid.NewString(), SessionID: uuid.NewString(), EnvironmentID: uuid.NewString(), + }, + Requested: runtimehistory.Range{Start: start, End: now, MaxPoints: 60}, + Resolution: time.Minute, + GeneratedAt: now, + } + response, err := runtimeHistoryResponse(value, value.TenantID, value.SessionID, value.Requested) + if err != nil || response.Coverage.ExpectedSampleCount != 1 { + t.Fatalf("unsafe expected sample ceiling: count=%d err=%v", response.Coverage.ExpectedSampleCount, err) + } +} + +func timeString(value time.Time) string { return fmt.Sprintf("%d", value.Unix()) } diff --git a/services/agents-api/internal/runtimehistory/service.go b/services/agents-api/internal/runtimehistory/service.go index 1ec714702..40859ce0f 100644 --- a/services/agents-api/internal/runtimehistory/service.go +++ b/services/agents-api/internal/runtimehistory/service.go @@ -6,7 +6,12 @@ import ( "time" ) -var ErrUnsupported = errors.New("Runtime history is unsupported for this Session") +var ( + ErrInvalidRange = errors.New("invalid Runtime history range") + ErrInvalidResult = errors.New("invalid Runtime history result") + ErrUnavailable = errors.New("Runtime history is unavailable") + ErrUnsupported = errors.New("Runtime history is unsupported for this Session") +) type ScopeResolver interface { ResolveRuntimeHistoryScope(context.Context, string, string) (Scope, error) @@ -29,7 +34,7 @@ func NewService(resolver ScopeResolver, reader Reader) (*Service, error) { return nil, errors.New("Runtime history resolver and reader are required") } capabilities := reader.Capabilities() - if err := capabilities.validate(); err != nil { + if err := capabilities.Validate(); err != nil { return nil, err } return &Service{resolver: resolver, reader: reader, capabilities: cloneCapabilities(capabilities), now: time.Now}, nil @@ -43,28 +48,31 @@ func (s *Service) Capabilities() Capabilities { // request. The history Reader receives only Core identity and bounded range data; // provider-native identity is never an authority-bearing input. func (s *Service) QuerySession(ctx context.Context, tenantID, sessionID string, requested Range) (Response, error) { - now := s.now().UTC() - if requested.Start.IsZero() || requested.End.IsZero() || !requested.End.After(requested.Start) || requested.End.After(now.Add(time.Second)) || requested.End.Sub(requested.Start) > s.capabilities.MaximumRange || requested.MaxPoints < 2 || requested.MaxPoints > s.capabilities.MaximumPoints { - return Response{}, errors.New("invalid Runtime history range") + requestNow := s.now().UTC() + if !validPublicBoundary(requested.Start) || !validPublicBoundary(requested.End) || !requested.End.After(requested.Start) || requested.End.After(requestNow.Add(time.Second)) || requested.End.Sub(requested.Start) > s.capabilities.MaximumRange || requested.MaxPoints < 2 || requested.MaxPoints > s.capabilities.MaximumPoints { + return Response{}, ErrInvalidRange } scope, err := s.resolver.ResolveRuntimeHistoryScope(ctx, tenantID, sessionID) if err != nil { return Response{}, err } if scope.TenantID != tenantID || scope.SessionID != sessionID { - return Response{}, errors.New("Runtime history resolver returned mismatched identity") + return Response{}, ErrInvalidResult } if err := scope.validate(); err != nil { - return Response{}, err + return Response{}, ErrInvalidResult } step := resolution(requested.End.Sub(requested.Start), requested.MaxPoints, s.capabilities.MinimumStep) - query := Query{Scope: scope, Start: requested.Start.UTC(), End: requested.End.UTC(), Step: step, MaxPoints: requested.MaxPoints} + query := Query{ + Scope: scope, Start: requested.Start.UTC(), End: requested.End.UTC(), Step: step, Retention: s.capabilities.Retention, MaxPoints: requested.MaxPoints, + MaximumSeries: s.capabilities.MaximumSeries, MaximumTotalPoints: s.capabilities.MaximumTotalPoints, + } result, err := s.reader.Query(ctx, query) if err != nil { - return Response{}, err + return Response{}, ErrUnavailable } - if err := validateResult(query, result, now); err != nil { - return Response{}, err + if err := validateResult(query, result, s.now().UTC()); err != nil { + return Response{}, ErrInvalidResult } return Response{ Capabilities: cloneCapabilities(s.capabilities), @@ -73,16 +81,25 @@ func (s *Service) QuerySession(ctx context.Context, tenantID, sessionID string, Resolution: query.Step, GeneratedAt: result.GeneratedAt.UTC(), RetainedFrom: cloneTime(result.RetainedFrom), + Coverage: append([]CoveragePoint(nil), result.Coverage...), Series: cloneSeries(result.Series), }, nil } func resolution(duration time.Duration, maxPoints int, minimum time.Duration) time.Duration { - step := (duration + time.Duration(maxPoints) - 1) / time.Duration(maxPoints) + divisor := time.Duration(maxPoints) + step := duration / divisor + if duration%divisor != 0 { + step++ + } if step < minimum { return minimum } - return ((step + time.Second - 1) / time.Second) * time.Second + seconds := step / time.Second + if step%time.Second != 0 { + seconds++ + } + return seconds * time.Second } func cloneTime(value *time.Time) *time.Time { diff --git a/services/agents-api/internal/runtimehistory/service_test.go b/services/agents-api/internal/runtimehistory/service_test.go index 14202ff75..4e100dbb7 100644 --- a/services/agents-api/internal/runtimehistory/service_test.go +++ b/services/agents-api/internal/runtimehistory/service_test.go @@ -39,13 +39,15 @@ func (r *fakeReader) Query(_ context.Context, query Query) (Result, error) { func capabilities() Capabilities { return Capabilities{ - CollectionMode: CollectionPeriodic, - SampleInterval: 30 * time.Second, - Retention: 7 * 24 * time.Hour, - MinimumStep: 30 * time.Second, - MaximumRange: 24 * time.Hour, - MaximumPoints: 1_000, - Metrics: []Metric{MetricCPU, MetricMemory}, + CollectionMode: CollectionPeriodic, + SampleInterval: 30 * time.Second, + Retention: 7 * 24 * time.Hour, + MinimumStep: 30 * time.Second, + MaximumRange: 24 * time.Hour, + MaximumPoints: 1_000, + MaximumSeries: 64, + MaximumTotalPoints: 10_000, + Metrics: []Metric{MetricCPU, MetricMemory}, } } @@ -90,13 +92,42 @@ func TestServiceNeverQueriesBeforeOwnershipResolution(t *testing.T) { if err != nil { t.Fatal(err) } - now := time.Now().UTC() + now := time.Now().UTC().Truncate(time.Second) _, err = service.QuerySession(t.Context(), tenantID, sessionID, Range{Start: now.Add(-time.Hour), End: now, MaxPoints: 60}) if !errors.Is(err, denied) || len(reader.queries) != 0 { t.Fatalf("unauthorized history reached reader: err=%v queries=%+v", err, reader.queries) } } +func TestServiceValidatesResultAgainstTimeAfterReaderReturns(t *testing.T) { + requestNow := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + resultNow := requestNow.Add(2 * time.Second) + scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} + reader := &fakeReader{capabilities: capabilities(), result: Result{GeneratedAt: resultNow}} + service, err := NewService(fixedScopeResolver{scope: scope}, reader) + if err != nil { + t.Fatal(err) + } + clockCalls := 0 + service.now = func() time.Time { + clockCalls++ + if clockCalls == 1 { + return requestNow + } + return resultNow + } + + _, err = service.QuerySession(t.Context(), tenantID, sessionID, Range{ + Start: requestNow.Add(-time.Hour), End: requestNow, MaxPoints: 60, + }) + if err != nil { + t.Fatalf("valid result from slow Reader rejected: %v", err) + } + if clockCalls != 2 { + t.Fatalf("clock calls = %d, want request and result validation times", clockCalls) + } +} + func TestServiceRejectsInvalidRangeAndResolverIdentity(t *testing.T) { reader := &fakeReader{capabilities: capabilities()} scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} @@ -104,7 +135,7 @@ func TestServiceRejectsInvalidRangeAndResolverIdentity(t *testing.T) { if err != nil { t.Fatal(err) } - now := time.Now().UTC() + now := time.Now().UTC().Truncate(time.Second) for _, requested := range []Range{ {Start: now, End: now, MaxPoints: 60}, {Start: now.Add(-25 * time.Hour), End: now, MaxPoints: 60}, @@ -130,9 +161,24 @@ func TestServiceRejectsMalformedBackendResults(t *testing.T) { scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} base := Series{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker"} for name, mutate := range map[string]func(*Result){ - "cross tenant": func(result *Result) { result.Series[0].TenantID = "55555555-5555-4555-8555-555555555555" }, + "cross tenant": func(result *Result) { result.Series[0].TenantID = "55555555-5555-4555-8555-555555555555" }, + "noncanonical allocation": func(result *Result) { result.Series[0].AllocationID = "urn:uuid:" + allocationID }, + "nil allocation": func(result *Result) { result.Series[0].AllocationID = "00000000-0000-0000-0000-000000000000" }, + "generation before range": func(result *Result) { result.GeneratedAt = start.Add(-time.Second) }, + "pre-epoch generation": func(result *Result) { result.GeneratedAt = time.Unix(-1, 0) }, + "pre-epoch incarnation": func(result *Result) { result.Series[0].StartedAt = time.Unix(-1, 0) }, + "incarnation newer than generation": func(result *Result) { + result.GeneratedAt = start.Add(10 * time.Minute) + result.Series[0].StartedAt = start.Add(10*time.Minute + time.Nanosecond) + result.Series[0].Points = nil + }, + "subsecond retention": func(result *Result) { value := start.Add(time.Nanosecond); result.RetainedFrom = &value }, "duplicate series": func(result *Result) { result.Series = append(result.Series, result.Series[0]) }, "out of range": func(result *Result) { result.Series[0].Points[0].End = now.Add(time.Minute) }, + "subsecond bucket": func(result *Result) { result.Series[0].Points[0].Start = start.Add(time.Nanosecond) }, + "bucket ends at incarnation": func(result *Result) { + result.Series[0].StartedAt = result.Series[0].Points[0].End + }, "zero coverage value": func(result *Result) { value := 1.0; result.Series[0].Points[0].CPUUtilizationRatio = &value }, "exclusive end": func(result *Result) { point := &result.Series[0].Points[0] @@ -161,6 +207,43 @@ func TestServiceRejectsMalformedBackendResults(t *testing.T) { point.ObservationCount, point.ObservedCount, point.CPUContributorCount = 1, 1, 1 point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start }, + "unsafe observation count": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount = int(maxSafeInteger) + 1 + point.UnavailableCount = point.ObservationCount + }, + "unsafe coverage count": func(result *Result) { + count := int(maxSafeInteger) + 1 + result.Coverage = []CoveragePoint{{Start: start, End: start.Add(time.Minute), ObservationCount: count, UnavailableCount: count}} + }, + "unsafe coverage sum": func(result *Result) { + count := int(maxSafeInteger/2 + 1) + result.Coverage = []CoveragePoint{ + {Start: start, End: start.Add(time.Minute), FirstObservedAt: start, LastObservedAt: start, ObservationCount: count, UnavailableCount: count}, + {Start: start.Add(time.Minute), End: start.Add(2 * time.Minute), FirstObservedAt: start.Add(time.Minute), LastObservedAt: start.Add(time.Minute), ObservationCount: count, UnavailableCount: count}, + } + }, + "coverage predates retention": func(result *Result) { + retained := start.Add(30 * time.Second) + result.RetainedFrom = &retained + result.Coverage = []CoveragePoint{{ + Start: start, End: start.Add(time.Minute), FirstObservedAt: start.Add(10 * time.Second), LastObservedAt: start.Add(10 * time.Second), + ObservationCount: 1, UnavailableCount: 1, + }} + }, + "series predates retention": func(result *Result) { + retained := start.Add(30 * time.Second) + result.RetainedFrom = &retained + point := &result.Series[0].Points[0] + point.FirstObservedAt, point.LastObservedAt = start.Add(10*time.Second), start.Add(10*time.Second) + point.ObservationCount, point.UnavailableCount = 1, 1 + }, + "observation newer than generation": func(result *Result) { + point := &result.Series[0].Points[0] + point.ObservationCount, point.ObservedCount = 1, 1 + point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start.Add(time.Second) + result.GeneratedAt = point.Start + }, } { t.Run(name, func(t *testing.T) { result := Result{GeneratedAt: now, Series: []Series{base}} @@ -194,3 +277,11 @@ func TestServiceRejectsUnsafeCapabilities(t *testing.T) { } } } + +func TestResolutionUsesOverflowSafeCeiling(t *testing.T) { + duration := (time.Duration(1<<63-1) / time.Second) * time.Second + value := resolution(duration, 2, time.Second) + if value <= 0 || value < duration/2 || value%time.Second != 0 { + t.Fatalf("unsafe resolution for large duration: %v", value) + } +} diff --git a/services/agents-api/internal/runtimehistory/types.go b/services/agents-api/internal/runtimehistory/types.go index 2bdfb8868..55b83caa5 100644 --- a/services/agents-api/internal/runtimehistory/types.go +++ b/services/agents-api/internal/runtimehistory/types.go @@ -26,24 +26,49 @@ const ( var providerTypePattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) +const maxSafeInteger = uint64(1<<53 - 1) + +func canonicalUUID(value string) bool { + parsed, err := uuid.Parse(value) + return err == nil && parsed != uuid.Nil && parsed.String() == value +} + +func validPublicTime(value time.Time) bool { + if value.IsZero() { + return false + } + seconds := value.Unix() + return seconds >= 0 && uint64(seconds) <= maxSafeInteger +} + +func validPublicBoundary(value time.Time) bool { + return validPublicTime(value) && value.Nanosecond() == 0 +} + +func validCount(value int) bool { + return value >= 0 && uint64(value) <= maxSafeInteger +} + // Capabilities contains only backend-neutral facts safe to expose through a // future public capability response. Backend names, endpoints, credentials and // tenant identity are deliberately absent. type Capabilities struct { - CollectionMode CollectionMode - SampleInterval time.Duration - Retention time.Duration - MinimumStep time.Duration - MaximumRange time.Duration - MaximumPoints int - Metrics []Metric + CollectionMode CollectionMode + SampleInterval time.Duration + Retention time.Duration + MinimumStep time.Duration + MaximumRange time.Duration + MaximumPoints int + MaximumSeries int + MaximumTotalPoints int + Metrics []Metric } func (c Capabilities) Durable() bool { return c.CollectionMode == CollectionPeriodic && c.SampleInterval > 0 } -func (c Capabilities) validate() error { +func (c Capabilities) Validate() error { switch c.CollectionMode { case CollectionOnRead: if c.SampleInterval != 0 { @@ -59,9 +84,17 @@ func (c Capabilities) validate() error { if c.Retention <= 0 || c.MinimumStep <= 0 || c.MaximumRange <= 0 || c.MaximumRange > c.Retention { return errors.New("invalid Runtime history time bounds") } + for _, value := range []time.Duration{c.SampleInterval, c.Retention, c.MinimumStep, c.MaximumRange} { + if value%time.Second != 0 { + return errors.New("Runtime history public time bounds must use whole seconds") + } + } if c.MaximumPoints < 2 || c.MaximumPoints > 10_000 { return errors.New("invalid Runtime history point limit") } + if c.MaximumSeries < 1 || c.MaximumSeries > 1_000 || c.MaximumTotalPoints < c.MaximumPoints || c.MaximumTotalPoints > 100_000 { + return errors.New("invalid Runtime history result limits") + } if len(c.Metrics) == 0 || len(c.Metrics) > 2 { return errors.New("invalid Runtime history metrics") } @@ -81,7 +114,7 @@ type Scope struct { func (s Scope) validate() error { for _, value := range []string{s.TenantID, s.SessionID, s.EnvironmentID} { - if uuid.Validate(value) != nil { + if !canonicalUUID(value) { return errors.New("invalid Runtime history scope") } } @@ -95,9 +128,11 @@ type Range struct { type Query struct { Scope - Start, End time.Time - Step time.Duration - MaxPoints int + Start, End time.Time + Step time.Duration + Retention time.Duration + MaxPoints int + MaximumSeries, MaximumTotalPoints int } // Result is returned by a backend Reader before the service validates identity, @@ -106,9 +141,17 @@ type Query struct { type Result struct { GeneratedAt time.Time RetainedFrom *time.Time + Coverage []CoveragePoint Series []Series } +type CoveragePoint struct { + Start, End time.Time + FirstObservedAt, LastObservedAt time.Time + ObservationCount int + ObservedCount, UnavailableCount int +} + type Series struct { Scope AllocationID string @@ -138,24 +181,78 @@ type Response struct { Resolution time.Duration GeneratedAt time.Time RetainedFrom *time.Time + Coverage []CoveragePoint Series []Series } +func (r Response) Validate(now time.Time) error { + if err := r.Capabilities.Validate(); err != nil || r.Scope.validate() != nil { + return ErrInvalidResult + } + if !validPublicBoundary(r.Requested.Start) || !validPublicBoundary(r.Requested.End) || !r.Requested.End.After(r.Requested.Start) || r.Requested.End.Sub(r.Requested.Start) > r.MaximumRange || r.Requested.MaxPoints < 2 || r.Requested.MaxPoints > r.MaximumPoints { + return ErrInvalidResult + } + wantResolution := resolution(r.Requested.End.Sub(r.Requested.Start), r.Requested.MaxPoints, r.MinimumStep) + if r.Resolution != wantResolution { + return ErrInvalidResult + } + query := Query{ + Scope: r.Scope, Start: r.Requested.Start, End: r.Requested.End, Step: r.Resolution, Retention: r.Retention, MaxPoints: r.Requested.MaxPoints, + MaximumSeries: r.MaximumSeries, MaximumTotalPoints: r.MaximumTotalPoints, + } + if err := validateResult(query, Result{GeneratedAt: r.GeneratedAt, RetainedFrom: r.RetainedFrom, Coverage: r.Coverage, Series: r.Series}, now); err != nil { + return ErrInvalidResult + } + return nil +} + func validateResult(query Query, result Result, now time.Time) error { - if result.GeneratedAt.IsZero() || result.GeneratedAt.After(now.Add(time.Second)) { + if query.Retention <= 0 { + return errors.New("invalid Runtime history query retention") + } + if !validPublicTime(result.GeneratedAt) || result.GeneratedAt.Before(query.Start) || result.GeneratedAt.After(now.Add(time.Second)) { return errors.New("invalid Runtime history generation time") } - if result.RetainedFrom != nil && (result.RetainedFrom.IsZero() || result.RetainedFrom.After(query.End)) { + if result.RetainedFrom != nil && (!validPublicBoundary(*result.RetainedFrom) || result.RetainedFrom.After(query.End) || result.RetainedFrom.After(result.GeneratedAt)) { return errors.New("invalid Runtime history retention boundary") } + retainedStart := query.Start + if configuredStart := result.GeneratedAt.Add(-query.Retention); configuredStart.After(retainedStart) { + retainedStart = configuredStart + } + if result.RetainedFrom != nil && result.RetainedFrom.After(retainedStart) { + retainedStart = *result.RetainedFrom + } + if len(result.Coverage) > query.MaxPoints { + return errors.New("Runtime history result exceeds coverage point limit") + } + var coverageObservationCount uint64 + for index, point := range result.Coverage { + if err := validateCoveragePoint(query, point); err != nil { + return err + } + if uint64(point.ObservationCount) > maxSafeInteger-coverageObservationCount { + return errors.New("Runtime history coverage sample count exceeds safe integer limit") + } + coverageObservationCount += uint64(point.ObservationCount) + if point.LastObservedAt.After(result.GeneratedAt) { + return errors.New("Runtime history coverage observation is newer than result") + } + if point.ObservationCount > 0 && point.FirstObservedAt.Before(retainedStart) { + return errors.New("Runtime history coverage predates retention") + } + if index > 0 && result.Coverage[index-1].End.After(point.Start) { + return errors.New("Runtime history coverage points overlap or are out of order") + } + } seriesKeys := map[string]bool{} - totalPoints := 0 - if len(result.Series) > query.MaxPoints { + totalPoints := len(result.Coverage) + if len(result.Series) > query.MaximumSeries { return errors.New("Runtime history result exceeds series limit") } for index := range result.Series { series := &result.Series[index] - if series.Scope != query.Scope || uuid.Validate(series.AllocationID) != nil || series.StartedAt.IsZero() || !series.StartedAt.Before(query.End) || !providerTypePattern.MatchString(series.ProviderType) { + if series.Scope != query.Scope || !canonicalUUID(series.AllocationID) || !validPublicTime(series.StartedAt) || series.StartedAt.After(result.GeneratedAt) || !series.StartedAt.Before(query.End) || !providerTypePattern.MatchString(series.ProviderType) { return errors.New("invalid Runtime history series identity") } key := series.AllocationID + "\x00" + series.StartedAt.UTC().Format(time.RFC3339Nano) @@ -164,7 +261,7 @@ func validateResult(query Query, result Result, now time.Time) error { } seriesKeys[key] = true totalPoints += len(series.Points) - if totalPoints > query.MaxPoints { + if len(series.Points) > query.MaxPoints || totalPoints > query.MaximumTotalPoints { return errors.New("Runtime history result exceeds point limit") } for pointIndex := range series.Points { @@ -172,6 +269,12 @@ func validateResult(query Query, result Result, now time.Time) error { if err := validatePoint(query, series.StartedAt, point); err != nil { return err } + if point.LastObservedAt.After(result.GeneratedAt) { + return errors.New("Runtime history observation is newer than result") + } + if point.ObservationCount > 0 && point.FirstObservedAt.Before(retainedStart) { + return errors.New("Runtime history observation predates retention") + } if pointIndex > 0 && series.Points[pointIndex-1].End.After(point.Start) { return errors.New("Runtime history points overlap or are out of order") } @@ -180,12 +283,31 @@ func validateResult(query Query, result Result, now time.Time) error { return nil } +func validateCoveragePoint(query Query, point CoveragePoint) error { + if !validPublicBoundary(point.Start) || !validPublicBoundary(point.End) || point.Start.Before(query.Start) || !point.End.After(point.Start) || point.End.After(query.End) || point.End.Sub(point.Start) > query.Step { + return errors.New("invalid Runtime history coverage bounds") + } + if !validCount(point.ObservationCount) || !validCount(point.ObservedCount) || !validCount(point.UnavailableCount) || uint64(point.ObservedCount)+uint64(point.UnavailableCount) != uint64(point.ObservationCount) { + return errors.New("invalid Runtime history coverage count") + } + if point.ObservationCount == 0 { + if !point.FirstObservedAt.IsZero() || !point.LastObservedAt.IsZero() { + return errors.New("empty Runtime history coverage contains observations") + } + return nil + } + if !validPublicTime(point.FirstObservedAt) || !validPublicTime(point.LastObservedAt) || point.FirstObservedAt.Before(point.Start) || point.LastObservedAt.Before(point.FirstObservedAt) || !point.LastObservedAt.Before(point.End) { + return errors.New("invalid Runtime history coverage observation bounds") + } + return nil +} + func validatePoint(query Query, startedAt time.Time, point Point) error { - if point.Start.Before(query.Start) || !point.End.After(point.Start) || point.End.After(query.End) || point.End.Sub(point.Start) > query.Step || point.End.Before(startedAt) { + if !validPublicBoundary(point.Start) || !validPublicBoundary(point.End) || point.Start.Before(query.Start) || !point.End.After(point.Start) || point.End.After(query.End) || point.End.Sub(point.Start) > query.Step || !point.End.After(startedAt) { return errors.New("invalid Runtime history point bounds") } - if point.ObservationCount < 0 || point.ObservedCount < 0 || point.UnavailableCount < 0 || point.ObservedCount+point.UnavailableCount != point.ObservationCount || - point.CPUContributorCount < 0 || point.CPUContributorCount > point.ObservedCount || point.MemoryContributorCount < 0 || point.MemoryContributorCount > point.ObservedCount { + if !validCount(point.ObservationCount) || !validCount(point.ObservedCount) || !validCount(point.UnavailableCount) || uint64(point.ObservedCount)+uint64(point.UnavailableCount) != uint64(point.ObservationCount) || + !validCount(point.CPUContributorCount) || point.CPUContributorCount > point.ObservedCount || !validCount(point.MemoryContributorCount) || point.MemoryContributorCount > point.ObservedCount { return errors.New("invalid Runtime history point coverage") } if point.ObservationCount == 0 { @@ -194,7 +316,7 @@ func validatePoint(query Query, startedAt time.Time, point Point) error { } return nil } - if point.FirstObservedAt.Before(point.Start) || point.FirstObservedAt.Before(startedAt) || point.LastObservedAt.Before(point.FirstObservedAt) || !point.LastObservedAt.Before(point.End) { + if !validPublicTime(point.FirstObservedAt) || !validPublicTime(point.LastObservedAt) || point.FirstObservedAt.Before(point.Start) || point.FirstObservedAt.Before(startedAt) || point.LastObservedAt.Before(point.FirstObservedAt) || !point.LastObservedAt.Before(point.End) { return errors.New("invalid Runtime history observation bounds") } if value := point.CPUUtilizationRatio; value != nil && (math.IsNaN(*value) || math.IsInf(*value, 0) || *value < 0) { @@ -209,7 +331,6 @@ func validatePoint(query Query, startedAt time.Time, point Point) error { if point.CPUContributorCount > 0 && point.CPUUtilizationRatio == nil && point.CPUCapacityCores == nil { return errors.New("Runtime history CPU coverage lacks a value") } - const maxSafeInteger = uint64(1<<53 - 1) if point.MemoryUsageBytes != nil && *point.MemoryUsageBytes > maxSafeInteger || point.MemoryLimitBytes != nil && (*point.MemoryLimitBytes == 0 || *point.MemoryLimitBytes > maxSafeInteger) { return errors.New("invalid Runtime history memory value") } From 17710743f736f90c0ab6478522260ab36be4899b Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 09:04:59 +0800 Subject: [PATCH 22/51] Refine Runtime live dashboard trends --- apps/web/e2e/agents-lifecycle.spec.ts | 10 ++- apps/web/src/App.tsx | 4 +- .../src/features/dashboard/DashboardView.css | 82 +++++++++++++++++++ .../features/dashboard/DashboardView.test.tsx | 44 ++++++++++ .../dashboard/RuntimeObservabilityContent.tsx | 38 ++++++--- .../dashboard/RuntimeTrendCharts.test.tsx | 16 +++- .../features/dashboard/RuntimeTrendCharts.tsx | 48 +++++++++-- 7 files changed, 215 insertions(+), 27 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 9383c0b19..2d2f13012 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2524,6 +2524,7 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard.getByRole("heading", { name: "Memory usage" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Compute uptime" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Token throughput" })).toBeVisible(); + await expect(dashboard.getByLabel("Live Runtime sampling every 30 seconds")).toBeVisible(); const liveRange = dashboard.getByRole("group", { name: "Runtime live range" }); await expect(liveRange.getByRole("button", { name: "1h" })).toHaveAttribute("aria-pressed", "true"); await liveRange.getByRole("button", { name: "15m" }).click(); @@ -2536,6 +2537,7 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard.getByLabel("Memory usage: 3 live samples")).toBeVisible(); await expect(dashboard.getByLabel("Compute uptime: 3 live samples")).toBeVisible(); await expect(dashboard.getByLabel("Token throughput: 3 live samples")).toBeVisible(); + await expect(dashboard.locator(".dashboard-runtime-sample-count")).toContainText("3 samples"); await expect(dashboard.getByText("CPU usage live trend available")).toBeAttached(); await expect(dashboard).not.toContainText("Collecting live samples"); await expect(dashboard.getByRole("table", { name: "Runtime targets" })).not.toBeVisible(); @@ -2637,9 +2639,11 @@ test("publishes Dashboard counts only after every top-level Agent and Session pa await expect(dashboard.locator(".dashboard-summary > div").filter({ hasText: "Agents" })).toContainText("3"); await expect(dashboard.locator(".dashboard-summary > div").filter({ hasText: "Sessions" })).toContainText("2"); expect(agentAfters).toEqual([null, "agent_b"]); - expect(sessionAfters).toHaveLength(4); - expect(sessionAfters.filter((after) => after === null)).toHaveLength(2); - expect(sessionAfters.filter((after) => after === "session_snapshot")).toHaveLength(2); + // Session collection loads once for the page and once per Runtime snapshot. + // Returning to Dashboard refreshes Runtime immediately instead of waiting 30 seconds. + expect(sessionAfters).toHaveLength(6); + expect(sessionAfters.filter((after) => after === null)).toHaveLength(3); + expect(sessionAfters.filter((after) => after === "session_snapshot")).toHaveLength(3); }); test("keeps the previous Dashboard result when pagination exceeds the safety limit", async ({ page, request }) => { diff --git a/apps/web/src/App.tsx b/apps/web/src/App.tsx index c0925e986..2a3462eec 100644 --- a/apps/web/src/App.tsx +++ b/apps/web/src/App.tsx @@ -973,12 +973,12 @@ export function App() { void refreshSessions(); void refreshVaults(); void refreshEnvironmentTemplates(); - void refreshRuntimeSnapshot(); - }, [refreshAgents, refreshEnvironmentTemplates, refreshRuntimeSnapshot, refreshSessions, refreshVaults]); + }, [refreshAgents, refreshEnvironmentTemplates, refreshSessions, refreshVaults]); useEffect(() => { if (view !== "dashboard") return; let timer: number | null = null; + void refreshRuntimeSnapshot(); const schedule = () => { const jitter = Math.floor(Math.random() * 5_000); timer = window.setTimeout(() => { diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index f4563cb86..38d0e0033 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -1003,6 +1003,59 @@ margin: 0; } +.dashboard-runtime-live-controls { + display: flex; + min-width: 0; + align-items: center; + justify-content: flex-end; + gap: 10px; +} + +.dashboard-runtime-live-status, +.dashboard-runtime-sample-count { + display: inline-flex; + align-items: center; + white-space: nowrap; + color: var(--fg-muted); + font-family: var(--font-mono); + font-size: 9px; + font-variant-numeric: tabular-nums; +} + +.dashboard-runtime-live-status { + gap: 6px; + color: var(--success); + font-weight: 650; +} + +.dashboard-runtime-live-status-stale { + color: var(--warning); +} + +.dashboard-runtime-live-status i { + width: 6px; + height: 6px; + background: currentColor; + border-radius: 999px; + box-shadow: 0 0 0 3px color-mix(in srgb, currentColor 13%, transparent); + animation: dashboard-runtime-live-pulse 2.4s ease-out infinite; +} + +.dashboard-runtime-live-status-stale i { + animation: none; +} + +@keyframes dashboard-runtime-live-pulse { + 0%, 55% { box-shadow: 0 0 0 3px color-mix(in srgb, currentColor 13%, transparent); } + 80%, 100% { box-shadow: 0 0 0 7px transparent; } +} + +@media (prefers-reduced-motion: reduce) { + .dashboard-runtime-live-status i { + animation: none; + } +} + .dashboard-runtime-live-toolbar h3 { font-size: 12px; font-weight: 650; @@ -1109,9 +1162,13 @@ .dashboard-runtime-trend-legend span { display: inline-flex; + max-width: 104px; min-width: 0; align-items: center; gap: 5px; + overflow: hidden; + text-overflow: ellipsis; + white-space: nowrap; } .dashboard-runtime-trend-legend i { @@ -1197,6 +1254,21 @@ vector-effect: non-scaling-stroke; } +.dashboard-runtime-trend-area { + stroke: none; + opacity: .14; +} + +.dashboard-runtime-trend-fill-purple { + fill: #b998f4; +} + +.dashboard-runtime-trend-latest { + fill: var(--surface); + stroke-width: 1.75; + vector-effect: non-scaling-stroke; +} + .dashboard-runtime-trend-dot { fill: currentColor; stroke-width: 0; @@ -1959,6 +2031,16 @@ gap: 7px; } + .dashboard-runtime-live-controls { + width: 100%; + flex-wrap: wrap; + justify-content: flex-start; + } + + .dashboard-runtime-range { + margin-left: auto; + } + .dashboard-runtime-trend-legend { max-width: 100%; justify-content: flex-start; diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 51cd7dd43..172355208 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -248,6 +248,8 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain('aria-label="Runtime live-window charts"'); expect(html).toContain("Live resource trends"); expect(html).toContain("Browser-local samples · no durable history"); + expect(html).toContain("Live · 30s"); + expect(html).toContain("1 sample"); expect(html).toContain('aria-label="Runtime live range"'); expect(html).toContain('aria-pressed="true">1h'); expect(html).toContain("CPU usage"); @@ -320,6 +322,48 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("In this snapshot"); }); + it("does not present retained Runtime samples as live after a refresh failure", () => { + const runtimeSession = session("runtime-stale", { + environment: { + type: "openai_hosted", + id: "22222222-2222-4222-8222-222222222222", + capability_directories: [], + network: { access: "enabled", allowed_domains: [] }, + packages: { npm: [], python: [], system: [] }, + files: [], + plugins: [], + skills: [], + }, + }); + const observation: RuntimeObservation = { + id: runtimeSession.id, + object: "agent.runtime_observation", + session_id: runtimeSession.id, + environment_id: "22222222-2222-4222-8222-222222222222", + mode: "openai_hosted", + provider_type: "docker", + instance: { kind: "managed_allocation", allocation_id: "33333333-3333-4333-8333-333333333333", device_id: null, connection_generation: null }, + status: "observed", + reason: null, + allocation_created_at: 1_700_000_000, + resolved_at: 1_700_000_100, + observed_at: 1_700_000_090, + started_at: 1_700_000_010, + cpu: null, + memory: null, + }; + const html = render({ + runtimeSnapshot: { sessions: [runtimeSession], observations: [observation], loadedAt: 1_700_000_100_000 }, + runtimeCollectionState: "failed", + runtimeCollectionError: "Runtime refresh failed", + runtimeCollectionHasSnapshot: true, + }); + + expect(html).toContain("Stale · retrying"); + expect(html).toContain("Runtime sampling refresh failed; showing retained samples"); + expect(html).not.toContain("Live · 30s"); + }); + it("keeps an existing snapshot visible during a refresh and disables duplicate refresh", () => { const html = render({ agents: [agent()], diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index 19f06b320..eea80f8ff 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -35,6 +35,7 @@ import { type RuntimeDashboardRow, } from "./dashboard-model"; import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; +import { RUNTIME_SNAPSHOT_REFRESH_MS } from "./runtime-snapshot"; import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; import { appendRuntimeTrendSample, @@ -299,6 +300,7 @@ export function RuntimeObservabilityContent({ () => runtimeTrendRange(trendSamples, selectedTrendRange), [selectedTrendRange, trendSamples], ); + const latestTrendSample = visibleTrendSamples.at(-1); useEffect(() => { setTrendSamples((current) => appendRuntimeTrendSample(current, snapshot)); @@ -319,17 +321,31 @@ export function RuntimeObservabilityContent({

Live resource trends

Browser-local samples · no durable history

-
- {RUNTIME_TREND_RANGES.map((range) => ( - - ))} +
+ +
{hasLine ? `${title} live trend available` : `${title} collecting live samples; ${samples.length} of 2 minimum`}{hasLine ? `${title} live trend available` : `${title} collecting live samples; ${validPoints} of 2 valid points from ${samples.length} snapshots`}
SeriesLatest valueMissing samples
Unavailable150%1
{hasLine ? `${title} live trend available` : `${title} collecting live samples; ${validPoints} of 2 valid points from ${samples.length} snapshots`}{hasLine + ? `${title} ${source} trend available` + : source === "live" + ? `${title} collecting live samples; ${validPoints} of 2 valid points from ${samples.length} snapshots` + : `${title} ${emptyMessage}; ${emptyDetail ?? `${validPoints} valid points from ${samples.length} retained buckets`}`}
SeriesLatest valueMissing samples
{hasLine @@ -242,14 +595,32 @@ export function RuntimeTrendCharts({ const tokenMaximum = Math.max(1, ...finite(charts.tokens.flatMap((series) => series.points.map((point) => point.value)))); const newest = rangeEnd ?? samples.at(-1)?.sampledAt ?? Date.now(); const oldest = rangeStart ?? samples[0]?.sampledAt ?? newest - 60 * 60 * 1_000; + const domain = useMemo(() => ({ start: oldest, end: Math.max(oldest + 1, newest) }), [oldest, newest]); + const [viewRange, setViewRange] = useState(domain); + const previousDomain = useRef(domain); const durable = source === "durable"; + useEffect(() => { + const previous = previousDomain.current; + setViewRange((current) => { + const tolerance = Math.max(1_000, (previous.end - previous.start) / 100); + const wasFullRange = current.start <= previous.start + tolerance && current.end >= previous.end - tolerance; + if (wasFullRange) return domain; + const width = Math.min(current.end - current.start, domain.end - domain.start); + const wasFollowing = current.end >= previous.end - tolerance; + const end = wasFollowing ? domain.end : clamp(current.end, domain.start + width, domain.end); + return { start: clamp(end - width, domain.start, domain.end - width), end }; + }); + previousDomain.current = domain; + }, [domain]); + return (
- `${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> - formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> - formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> - `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} /> + `${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> + formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> + formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> + `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} /> +
); } From c2b1cb8ec38ddb9d54474a4f1bbbd8f7a9963efe Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 14:32:59 +0800 Subject: [PATCH 26/51] Stabilize durable Runtime dashboard refreshes --- apps/web/src/App.tsx | 18 ++++++++----- .../src/features/dashboard/DashboardView.css | 4 +++ .../features/dashboard/DashboardView.test.tsx | 24 +++++++---------- .../src/features/dashboard/DashboardView.tsx | 2 +- .../dashboard/RuntimeObservabilityContent.tsx | 26 +++++-------------- .../dashboard/runtime-snapshot.test.ts | 15 ++++++----- .../features/dashboard/runtime-snapshot.ts | 3 --- .../runtime-observability-design.md | 5 ++-- docs/web/README.md | 22 ++++++++-------- docs/web/README.zh-CN.md | 13 +++++----- docs/web/protocol-coverage.md | 11 ++++---- .../runtime-history/clickhouse/README.md | 16 +++++++++++- 12 files changed, 83 insertions(+), 76 deletions(-) diff --git a/apps/web/src/App.tsx b/apps/web/src/App.tsx index f17133c7d..318a9ecf9 100644 --- a/apps/web/src/App.tsx +++ b/apps/web/src/App.tsx @@ -319,6 +319,7 @@ export function App() { const sessionCollectionAbortRef = useRef(null); const runtimeCollectionAbortRef = useRef(null); const runtimeCollectionRequestRef = useRef(0); + const runtimeCollectionHasSnapshotRef = useRef(false); const filteredSessionCollectionAbortRef = useRef(null); const filteredSessionCollectionRequestRef = useRef(0); const sessionAgentFilterRef = useRef(sessionAgentFilter); @@ -541,12 +542,11 @@ export function App() { timedOut = true; controller.abort(); }, RUNTIME_SNAPSHOT_TIMEOUT_MS); - setRuntimeCollectionState("connecting"); + if (!runtimeCollectionHasSnapshotRef.current) setRuntimeCollectionState("connecting"); setRuntimeCollectionError(null); try { const result = await settleCollection(() => loadRuntimeDashboardSnapshot( core, - () => sessionCollectionRevisionRef.current, controller.signal, )); if ( @@ -568,6 +568,7 @@ export function App() { return false; } setRuntimeSnapshot(result.value); + runtimeCollectionHasSnapshotRef.current = true; setRuntimeCollectionHasSnapshot(true); setRuntimeCollectionState("ready"); return true; @@ -913,10 +914,11 @@ export function App() { const refreshDashboard = useCallback(() => { void refreshAgents(); - void refreshSessions(); void refreshVaults(); void refreshEnvironmentTemplates(); - void refreshRuntimeSnapshot(); + void (async () => { + if (await refreshSessions()) await refreshRuntimeSnapshot(); + })(); const filter = sessionAgentFilterRef.current; if (filter) void refreshFilteredSessions(filter); }, [refreshAgents, refreshEnvironmentTemplates, refreshFilteredSessions, refreshRuntimeSnapshot, refreshSessions, refreshVaults]); @@ -963,6 +965,7 @@ export function App() { setAgentCollectionHasSnapshot(false); setSessionCollectionHasSnapshot(false); setRuntimeSnapshot(null); + runtimeCollectionHasSnapshotRef.current = false; setRuntimeCollectionState("connecting"); setRuntimeCollectionError(null); setRuntimeCollectionHasSnapshot(false); @@ -985,7 +988,9 @@ export function App() { useEffect(() => { if (view !== "dashboard") return; let timer: number | null = null; - void refreshRuntimeSnapshot(); + void (async () => { + if (sessionCollectionState === "ready") await refreshRuntimeSnapshot(); + })(); const schedule = () => { const jitter = Math.floor(Math.random() * 5_000); timer = window.setTimeout(() => { @@ -1002,7 +1007,7 @@ export function App() { if (timer !== null) window.clearTimeout(timer); document.removeEventListener("visibilitychange", onVisibilityChange); }; - }, [refreshRuntimeSnapshot, view]); + }, [refreshRuntimeSnapshot, sessionCollectionState, view]); useEffect(() => { filteredSessionCollectionAbortRef.current?.abort(); @@ -2085,6 +2090,7 @@ export function App() { setSessionCollectionError(null); setSessionCollectionHasSnapshot(false); setRuntimeSnapshot(null); + runtimeCollectionHasSnapshotRef.current = false; setRuntimeCollectionState("connecting"); setRuntimeCollectionError(null); setRuntimeCollectionHasSnapshot(false); diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 3d500c10c..a27154840 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -1045,6 +1045,10 @@ animation: none; } +.dashboard-runtime-live-status-durable i { + animation: none; +} + @keyframes dashboard-runtime-live-pulse { 0%, 55% { box-shadow: 0 0 0 3px color-mix(in srgb, currentColor 13%, transparent); } 80%, 100% { box-shadow: 0 0 0 7px transparent; } diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 4045eb54e..bce7423f0 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -246,24 +246,20 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("Cumulative CPU / capacity"); expect(html).toContain("1m 13s / 2 cores"); expect(html).toContain("512 MiB / 2.00 GiB"); - expect(html).toContain('aria-label="Runtime live-window charts"'); + expect(html).toContain('aria-label="Runtime durable-history charts"'); expect(html).toContain("Resource trends"); - expect(html).toContain("Browser-local samples · reset on reload"); - expect(html).toContain('aria-label="Runtime trend source"'); - expect(html).toContain('aria-pressed="true">Live'); - expect(html).toContain('aria-pressed="false" disabled="">History'); - expect(html).toContain("Live · 30s"); - expect(html).toContain("1 sample"); - expect(html).toContain('aria-label="Runtime live range"'); + expect(html).toContain("ClickHouse-backed retained samples · explicit history source"); + expect(html).not.toContain('aria-label="Runtime trend source"'); + expect(html).toContain("History · loading"); + expect(html).toContain("0 buckets"); + expect(html).toContain('aria-label="Runtime durable range"'); expect(html).toContain('aria-pressed="true">1h'); expect(html).toContain("CPU usage"); expect(html).toContain("Memory usage"); expect(html).toContain("Compute uptime"); expect(html).toContain("Token throughput"); - expect(html).toContain("150%"); - expect(html).toContain("Collecting live samples"); - expect(html).toContain("1/2 valid points · 1 snapshots · no history is synthesized"); - expect(html).toContain("CPU usage collecting live samples; 1 of 2 valid points from 1 snapshots"); + expect(html).toContain("No retained CPU samples"); + expect(html).toContain("0/2 valid points · 0 snapshots · no history is synthesized"); expect(html).toContain("Latest value"); expect(html).toContain("Missing samples"); expect(html).toContain("Runtime targets"); @@ -363,8 +359,8 @@ describe("Dashboard loaded-result presentation", () => { runtimeCollectionHasSnapshot: true, }); - expect(html).toContain("Stale · retrying"); - expect(html).toContain("Runtime sampling refresh failed; showing retained samples"); + expect(html).toContain("Runtime: Runtime refresh failed"); + expect(html).toContain("History · loading"); expect(html).not.toContain("Live · 30s"); }); diff --git a/apps/web/src/features/dashboard/DashboardView.tsx b/apps/web/src/features/dashboard/DashboardView.tsx index 1565e8acf..499311ee0 100644 --- a/apps/web/src/features/dashboard/DashboardView.tsx +++ b/apps/web/src/features/dashboard/DashboardView.tsx @@ -384,7 +384,7 @@ export function DashboardView({

Runtime monitoring

-

Current provider samples · browser-local Live · ClickHouse History when configured

+

Current provider status · ClickHouse-backed metrics history when configured

{runtimeModel ? ( diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index 8818f0bbc..25ab16d98 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -40,7 +40,6 @@ import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; import type { RuntimeTrendSource } from "./RuntimeTrendCharts"; import { RUNTIME_DURABLE_RANGES, - runtimeTrendSourceAfterHistoryUnavailable, type RuntimeDurableRange, type RuntimeDurableSnapshot, } from "./runtime-history"; @@ -312,7 +311,6 @@ export function RuntimeObservabilityContent({ const [trendSamples, setTrendSamples] = useState(() => appendRuntimeTrendSample([], snapshot)); const [selectedTrendRange, setSelectedTrendRange] = useState(RUNTIME_TREND_WINDOW_MS); const [selectedDurableRange, setSelectedDurableRange] = useState(RUNTIME_DURABLE_RANGES[0].milliseconds); - const [sourceSelection, setSourceSelection] = useState<"auto" | RuntimeTrendSource>("auto"); const [durableSnapshot, setDurableSnapshot] = useState(null); const [durableState, setDurableState] = useState<"connecting" | "ready" | "unavailable" | "failed">("connecting"); const [durableError, setDurableError] = useState(null); @@ -320,11 +318,11 @@ export function RuntimeObservabilityContent({ () => runtimeTrendRange(trendSamples, selectedTrendRange), [selectedTrendRange, trendSamples], ); - const source: RuntimeTrendSource = sourceSelection === "auto" - ? durableSnapshot === null ? "live" : "durable" - : sourceSelection; - const selectedSamples = source === "durable" && durableSnapshot !== null - ? durableSnapshot.samples + const source: RuntimeTrendSource = durableState === "unavailable" && durableSnapshot === null + ? "live" + : "durable"; + const selectedSamples = source === "durable" + ? durableSnapshot?.samples ?? [] : visibleTrendSamples; const latestTrendSample = selectedSamples.at(-1); const selectedRange = source === "durable" ? selectedDurableRange : selectedTrendRange; @@ -348,7 +346,6 @@ export function RuntimeObservabilityContent({ if (result === null) { setDurableSnapshot(null); setDurableState("unavailable"); - setSourceSelection(runtimeTrendSourceAfterHistoryUnavailable); return; } setDurableSnapshot(result); @@ -389,19 +386,8 @@ export function RuntimeObservabilityContent({

{source === "durable" ? "ClickHouse-backed retained samples · explicit history source" : "Browser-local samples · reset on reload"}

-
- - -
{ it("publishes only exact Session and observation identity sets", async () => { - const value = await loadRuntimeDashboardSnapshot(core([session], [observation]), () => 4); + const value = await loadRuntimeDashboardSnapshot(core([session], [observation])); expect(value?.sessions).toEqual([session]); expect(value?.observations).toEqual([observation]); - await expect(loadRuntimeDashboardSnapshot(core([session], []), () => 4)).rejects.toBeInstanceOf( + await expect(loadRuntimeDashboardSnapshot(core([session], []))).rejects.toBeInstanceOf( RuntimeSnapshotIncompleteError, ); }); - it("discards a candidate when local Session state changes during collection", async () => { - let revision = 1; + it("does not discard a coherent API snapshot when unrelated local Session detail changes", async () => { const changing = core([session], [observation]); changing.listRuntimeObservations = async () => { - revision += 1; return { object: "list", data: [observation], has_more: false, first_id: session.id, last_id: session.id }; }; - await expect(loadRuntimeDashboardSnapshot(changing, () => revision)).resolves.toBeNull(); + await expect(loadRuntimeDashboardSnapshot(changing)).resolves.toMatchObject({ + sessions: [session], + observations: [observation], + }); }); it("fails closed when the target budget is exceeded", async () => { - await expect(loadRuntimeDashboardSnapshot(core([session], [observation]), () => 1, undefined, 0)).rejects.toThrow( + await expect(loadRuntimeDashboardSnapshot(core([session], [observation]), undefined, 0)).rejects.toThrow( "target budget", ); }); diff --git a/apps/web/src/features/dashboard/runtime-snapshot.ts b/apps/web/src/features/dashboard/runtime-snapshot.ts index 79d794406..c317449df 100644 --- a/apps/web/src/features/dashboard/runtime-snapshot.ts +++ b/apps/web/src/features/dashboard/runtime-snapshot.ts @@ -35,18 +35,15 @@ function setsEqual(left: ReadonlySet, right: ReadonlySet): boole export async function loadRuntimeDashboardSnapshot( core: Pick, - readSessionRevision: () => number, signal?: AbortSignal, targetLimit = RUNTIME_SNAPSHOT_TARGET_LIMIT, ): Promise { - const revision = readSessionRevision(); const [sessions, observations] = await Promise.all([ listAllCollectionPages((options) => core.listSessions(options), signal), listAllCollectionPages((options) => core.listRuntimeObservations(options), signal), ]); signal?.throwIfAborted(); - if (revision !== readSessionRevision()) return null; if (sessions.length > targetLimit || observations.length > targetLimit) { throw new RuntimeSnapshotIncompleteError("Runtime snapshot exceeded the Web target budget."); } diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index d02acd8df..3c5287132 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -454,8 +454,9 @@ seams separate: 3. A feature-local state model retains `last_complete`, current refresh status, local filters, and the selected time range. It aborts an overlapping refresh and marks old data stale after a failed or incomplete refresh. -4. Presentational components render summary metrics, a browser-local live window, - the Runtime table, and an identity detail surface. The live window contains only +4. Presentational components render summary metrics, the Runtime table, an identity + detail surface, and capability-selected trends. A qualified periodic Reader is + the primary trend source. Without one, the fallback live window contains only complete snapshots collected while this Dashboard instance is mounted; it is bounded, ephemeral, and never presented as durable operator history. diff --git a/docs/web/README.md b/docs/web/README.md index 372154ee3..3a4a7a911 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -29,17 +29,17 @@ credentials or execution into the browser. ### Dashboard -Dashboard is the starting point. It summarizes the current Agent and Session results, -loads complete tenant-scoped Runtime observation snapshots, and retains a bounded -browser-local Live window for CPU, memory, compute-uptime, and token-throughput -charts without inventing missing values. The searchable, filterable, sortable, -paginated semantic table remains available on demand. Live starts when the -Dashboard opens and offers 15-minute and one-hour views. When the operator enables -qualified periodic sampling plus the ClickHouse Reader, Web automatically discovers -the capability and adds a distinct retained History source with 1-hour, 6-hour, and -24-hour ranges. A reload reconstructs History from Core; Web never queries -ClickHouse directly or merges it silently with Live points. Token throughput stays -Live-only because Runtime telemetry does not duplicate canonical Session Usage. +Dashboard is the starting point. It summarizes the current Agent and Session results +and loads complete tenant-scoped Runtime observation snapshots without inventing +missing values. The searchable, filterable, sortable, paginated semantic table +remains available on demand. When the operator enables qualified periodic sampling +plus the ClickHouse Reader, Web automatically discovers the capability and uses the +retained 1-hour, 6-hour, and 24-hour History ranges as the primary chart source. +A reload reconstructs History from Core instead of briefly publishing a new +browser-local window; Web never receives ClickHouse credentials or queries it +directly. An unconfigured deployment falls back to a bounded browser-local Live +window. Token throughput stays Live-only because Runtime telemetry does not +duplicate canonical Session Usage. When a provider reports only cumulative CPU time, Web derives interval utilization only across adjacent samples from the same verified Runtime incarnation; restarts and counter regressions create diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index d4c4b161f..581c7e29d 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -27,12 +27,13 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 ### Dashboard -Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,加载完整的租户级 Runtime -观测快照,并保留有界的浏览器本地 live window,展示 CPU、内存、compute uptime 和 -token throughput 趋势,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的 -语义表格按需展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 -Session 的流程。live window 从打开 Dashboard 后开始采集,可选择 15 分钟或 1 小时 -视图,但不是跨浏览器持久历史;持久保留仍需运维方配置独立 history 能力。若 provider +Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,并加载完整的租户级 Runtime +观测快照,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的语义表格按需 +展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。 +运维方配置周期采样和 ClickHouse Reader 后,页面优先使用可跨刷新的 1 小时、6 小时和 +24 小时持久 History,不会先发布一个新的浏览器本地窗口再切换数据源。Web 只通过 Core +读取,不接收 ClickHouse 凭据,也不会直接连接 ClickHouse。未配置 History 的部署仍回退 +到有界的浏览器本地 live window。若 provider 只报告累计 CPU 时间,Web 仅在 相邻样本属于同一已验证 Runtime incarnation 时计算区间利用率;重启或计数回退会形成 数据缺口,不会制造峰值。 diff --git a/docs/web/protocol-coverage.md b/docs/web/protocol-coverage.md index f81732071..2272ea433 100644 --- a/docs/web/protocol-coverage.md +++ b/docs/web/protocol-coverage.md @@ -433,13 +433,14 @@ upstream. at most eight rows ordered by valid Core-reported `last_active_at`. A `self_hosted` label identifies only the Session profile, not an executor connection. Row actions navigate to the exact loaded Session. -- Runtime Live charts use bounded browser-local samples from complete joined snapshots. +- Runtime charts prefer bounded durable History reads when capability discovery + proves a periodic Reader. They do not publish a transient browser-local chart + first and then switch sources after reload. CPU rate is derived only from ordered cumulative counters for the same allocation and compute incarnation; memory remains point-in-time; gaps are not interpolated. - The 15-minute and one-hour ranges are labelled Live and reset across browser - lifecycle. After capability discovery proves a periodic Reader, History performs - bounded Session-scoped reads and exposes 1-hour, 6-hour, and 24-hour retained - ranges. The two sources remain separately labelled, and any incomplete or failed + History performs bounded Session-scoped reads and exposes 1-hour, 6-hour, and + 24-hour retained ranges. An unconfigured deployment falls back to the bounded + browser-local Live window. The sources remain explicitly labelled, and any incomplete or failed multi-Session History refresh retains the previous result rather than publishing a partial replacement. Token throughput remains Live-only. - Agent and Session collection states remain independent. A failed refresh may diff --git a/services/agents-api/runtime-history/clickhouse/README.md b/services/agents-api/runtime-history/clickhouse/README.md index aa8c5ddf3..c8919bfa8 100644 --- a/services/agents-api/runtime-history/clickhouse/README.md +++ b/services/agents-api/runtime-history/clickhouse/README.md @@ -8,7 +8,9 @@ or lifecycle decisions. ## Security and ownership - The Collector writer may insert into the generic `metrics_gauge` and - `metrics_sum` tables. + `metrics_sum` tables. ClickHouse executes materialized-view queries under the + inserting identity, so that writer also needs `SELECT` on only the three + source columns consumed by each reference view. - The Agents API reader should receive `SELECT` on `runtime_history_metrics` only. It does not need access to generic telemetry, schema mutation, or another tenant selector. @@ -32,6 +34,18 @@ or lifecycle decisions. The materialized views retain only the five values needed by the history API; provider receipts, native container identities, paths, and raw errors never enter the projection. + Grant the Collector identity the minimum source-column reads required when + those views run: + + ```sql + GRANT SELECT(Attributes, MetricName, Value) + ON runtime_history.metrics_gauge TO agents_runtime_writer; + GRANT SELECT(Attributes, MetricName, Value) + ON runtime_history.metrics_sum TO agents_runtime_writer; + ``` + + Keep the existing `INSERT` grants on both generic tables. No broader table + read or projection read is required by the writer. 3. Create a read-only ClickHouse account for Core: ```sql From af8d83f04593c262648b45e2d1e9870cdcc1560c Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 14:48:59 +0800 Subject: [PATCH 27/51] Serialize Dashboard Runtime refreshes --- apps/web/e2e/agents-lifecycle.spec.ts | 25 ++++++++++++++++++++----- apps/web/src/App.tsx | 17 +++++++++++++++-- 2 files changed, 35 insertions(+), 7 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 02cbf3fa5..80c0bf1aa 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2540,6 +2540,22 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand last_id: data.at(-1)?.id ?? null, }); + await page.route("**/v1/agents/runtime-history/capabilities", (route) => route.fulfill({ + status: 200, + contentType: "application/json", + body: JSON.stringify({ + object: "agent.runtime_history_capabilities", + available: false, + reason: "not_configured", + collection_mode: null, + sample_interval_seconds: null, + retention_seconds: null, + minimum_step_seconds: null, + maximum_range_seconds: null, + maximum_points: null, + metrics: [], + }), + })); await page.goto("/"); const dashboard = page.locator(".dashboard-page"); await expect(dashboard.getByRole("heading", { name: "Dashboard", exact: true })).toBeVisible(); @@ -2785,9 +2801,8 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await page.goto("/"); const dashboard = page.locator(".dashboard-runtime-panel"); - const source = dashboard.getByRole("group", { name: "Runtime trend source" }); - await expect(source.getByRole("button", { name: "History" })).toBeEnabled(); - await expect(source.getByRole("button", { name: "History" })).toHaveAttribute("aria-pressed", "true"); + await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); + await expect(dashboard.getByLabel(/Durable · 30s; 1 Runtime targets/)).toBeVisible(); await expect(dashboard.getByLabel("Runtime durable-history charts")).toBeVisible(); await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); await expect(dashboard).toContainText("4 buckets"); @@ -2810,8 +2825,8 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa } await page.reload(); - await expect(dashboard.getByRole("group", { name: "Runtime trend source" }).getByRole("button", { name: "History" })) - .toHaveAttribute("aria-pressed", "true"); + await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); + await expect(dashboard.getByLabel(/Durable · 30s; 1 Runtime targets/)).toBeVisible(); await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); }); diff --git a/apps/web/src/App.tsx b/apps/web/src/App.tsx index 318a9ecf9..9a8381e7d 100644 --- a/apps/web/src/App.tsx +++ b/apps/web/src/App.tsx @@ -320,6 +320,7 @@ export function App() { const runtimeCollectionAbortRef = useRef(null); const runtimeCollectionRequestRef = useRef(0); const runtimeCollectionHasSnapshotRef = useRef(false); + const dashboardRefreshInFlightRef = useRef(false); const filteredSessionCollectionAbortRef = useRef(null); const filteredSessionCollectionRequestRef = useRef(0); const sessionAgentFilterRef = useRef(sessionAgentFilter); @@ -913,11 +914,17 @@ export function App() { }, [coreGeneration, refreshAgents, refreshFilteredSessions, refreshSelectedSession, refreshSessions]); const refreshDashboard = useCallback(() => { + if (dashboardRefreshInFlightRef.current) return; + dashboardRefreshInFlightRef.current = true; void refreshAgents(); void refreshVaults(); void refreshEnvironmentTemplates(); void (async () => { - if (await refreshSessions()) await refreshRuntimeSnapshot(); + try { + if (await refreshSessions()) await refreshRuntimeSnapshot(); + } finally { + dashboardRefreshInFlightRef.current = false; + } })(); const filter = sessionAgentFilterRef.current; if (filter) void refreshFilteredSessions(filter); @@ -989,7 +996,13 @@ export function App() { if (view !== "dashboard") return; let timer: number | null = null; void (async () => { - if (sessionCollectionState === "ready") await refreshRuntimeSnapshot(); + if ( + sessionCollectionState === "ready" && + !runtimeCollectionHasSnapshotRef.current && + !dashboardRefreshInFlightRef.current + ) { + await refreshRuntimeSnapshot(); + } })(); const schedule = () => { const jitter = Math.floor(Math.random() * 5_000); From c129744f40521749537419e9a66e190d72fe21bf Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 14:52:10 +0800 Subject: [PATCH 28/51] Guard Runtime refresh trigger races --- apps/web/e2e/agents-lifecycle.spec.ts | 20 ++++++++++++++++++++ apps/web/src/App.tsx | 8 ++++++-- 2 files changed, 26 insertions(+), 2 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 80c0bf1aa..6428c528f 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2392,8 +2392,28 @@ test("presents Dashboard page-chain results and System boundaries without extra "/v1/agents/sessions/session_snapshot/items", "/v1/agents/sessions/session_snapshot/turns", ]; + let delayNextSessionList = true; + let releaseSessionList: (() => void) | null = null; + let markSessionListStarted: (() => void) | null = null; + const sessionListStarted = new Promise((resolve) => { + markSessionListStarted = resolve; + }); + await page.route("**/v1/agents/sessions*", async (route) => { + if (new URL(route.request().url()).pathname !== "/v1/agents/sessions" || !delayNextSessionList) { + return route.continue(); + } + delayNextSessionList = false; + markSessionListStarted?.(); + await new Promise((resolve) => { + releaseSessionList = resolve; + }); + await route.continue(); + }); const refresh = dashboard.getByRole("button", { name: "Refresh Dashboard snapshot" }); await refresh.click(); + await sessionListStarted; + await page.evaluate(() => document.dispatchEvent(new Event("visibilitychange"))); + releaseSessionList?.(); await expect.poll(async () => { const entries = await fixtureRequests(request); return [count(entries, "/v1/agents"), count(entries, "/v1/agents/sessions")]; diff --git a/apps/web/src/App.tsx b/apps/web/src/App.tsx index 9a8381e7d..81aba4d85 100644 --- a/apps/web/src/App.tsx +++ b/apps/web/src/App.tsx @@ -1007,12 +1007,16 @@ export function App() { const schedule = () => { const jitter = Math.floor(Math.random() * 5_000); timer = window.setTimeout(() => { - if (!document.hidden) void refreshRuntimeSnapshot(); + if (!document.hidden && !dashboardRefreshInFlightRef.current) { + void refreshRuntimeSnapshot(); + } schedule(); }, RUNTIME_SNAPSHOT_REFRESH_MS + jitter); }; const onVisibilityChange = () => { - if (!document.hidden) void refreshRuntimeSnapshot(); + if (!document.hidden && !dashboardRefreshInFlightRef.current) { + void refreshRuntimeSnapshot(); + } }; document.addEventListener("visibilitychange", onVisibilityChange); schedule(); From 0361f94f03d797dcdd62643f5c8bb7cef4862d3d Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 15:13:30 +0800 Subject: [PATCH 29/51] Place timeline controls below each runtime chart --- apps/web/e2e/agents-lifecycle.spec.ts | 20 +++---- .../src/features/dashboard/DashboardView.css | 14 ++--- .../dashboard/RuntimeTrendCharts.test.tsx | 14 +++-- .../features/dashboard/RuntimeTrendCharts.tsx | 60 +++++++++---------- 4 files changed, 55 insertions(+), 53 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 6428c528f..433c59d74 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2661,8 +2661,9 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await cpuSeriesToggle.click(); await expect(cpuSeriesToggle).toHaveAttribute("aria-pressed", "true"); - const chartTimeline = dashboard.getByRole("region", { name: "Runtime chart timeline" }); - const timelineStart = chartTimeline.getByRole("slider", { name: "Timeline start" }); + await expect(dashboard.locator(".dashboard-runtime-timeline")).toHaveCount(4); + const chartTimeline = cpuCard.getByRole("region", { name: "CPU usage timeline" }); + const timelineStart = chartTimeline.getByRole("slider", { name: "CPU usage timeline start" }); const initialTimelineStart = Number(await timelineStart.getAttribute("aria-valuenow")); await timelineStart.scrollIntoViewIfNeeded(); const startHandleBox = await timelineStart.boundingBox(); @@ -2675,12 +2676,11 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await page.mouse.up(); await expect.poll(async () => Number(await timelineStart.getAttribute("aria-valuenow"))).toBeGreaterThan(initialTimelineStart); const resizedTimelineStart = Number(await timelineStart.getAttribute("aria-valuenow")); - const trendCharts = dashboard.locator(".dashboard-runtime-chart-frame svg"); - await expect(trendCharts).toHaveCount(4); - for (const chart of await trendCharts.all()) { - await expect(chart).toHaveAttribute("data-view-start", String(resizedTimelineStart)); + await expect(cpuChart).toHaveAttribute("data-view-start", String(resizedTimelineStart)); + for (const chartName of ["Memory usage", "Compute uptime", "Token throughput"]) { + await expect(dashboard.getByLabel(`${chartName}: 3 live samples`)).toHaveAttribute("data-view-start", String(initialTimelineStart)); } - const timelineWindow = chartTimeline.getByRole("button", { name: "Pan selected timeline window" }); + const timelineWindow = chartTimeline.getByRole("button", { name: "Pan CPU usage timeline window" }); const timelineWindowBox = await timelineWindow.boundingBox(); expect(timelineWindowBox).not.toBeNull(); expect(timelineWindowBox!.height).toBeGreaterThanOrEqual(24); @@ -2834,15 +2834,13 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await durableCpuChart.press("ArrowLeft"); const durableCpuCard = durableCpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); await expect(durableCpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Unavailable"); - const durableTimelineStart = dashboard.getByRole("slider", { name: "Timeline start" }); + const durableTimelineStart = durableCpuCard.getByRole("slider", { name: "CPU usage timeline start" }); const durableInitialStart = Number(await durableTimelineStart.getAttribute("aria-valuenow")); await durableTimelineStart.focus(); await durableTimelineStart.press("ArrowRight"); await expect.poll(async () => Number(await durableTimelineStart.getAttribute("aria-valuenow"))).toBeGreaterThan(durableInitialStart); const durableViewStart = await durableTimelineStart.getAttribute("aria-valuenow"); - for (const chart of await dashboard.locator(".dashboard-runtime-chart-frame svg").all()) { - await expect(chart).toHaveAttribute("data-view-start", durableViewStart!); - } + await expect(durableCpuChart).toHaveAttribute("data-view-start", durableViewStart!); await page.reload(); await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 666220a0f..6a0d0dabd 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -1466,11 +1466,11 @@ .dashboard-runtime-timeline { min-width: 0; - padding: 12px 15px 10px; - grid-column: 1 / -1; - background: var(--surface); - border: 1px solid var(--line); - border-radius: 9px; + padding: 10px 15px 9px; + margin: 2px -15px -10px; + background: color-mix(in srgb, var(--surface-subtle) 54%, var(--surface)); + border-top: 1px solid var(--line); + border-radius: 0 0 8px 8px; } .dashboard-runtime-timeline > header, @@ -1490,7 +1490,7 @@ } .dashboard-runtime-timeline > header strong { - font-size: 11px; + font-size: 9px; } .dashboard-runtime-timeline > header span, @@ -1526,7 +1526,7 @@ .dashboard-runtime-timeline-track { height: 36px; - margin: 8px 7px 0; + margin: 6px 7px 0; position: relative; cursor: pointer; touch-action: none; diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index 82f7c54f2..d37872a34 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -44,7 +44,7 @@ describe("Runtime live-window chart accessibility", () => { expect(html).not.toContain("Collecting live samples"); }); - it("exposes interactive series, point selection, and a draggable shared timeline", () => { + it("exposes interactive series, point selection, and a draggable timeline below every chart", () => { const html = renderToStaticMarkup( , ); @@ -52,10 +52,14 @@ describe("Runtime live-window chart accessibility", () => { expect(html).toContain('role="application"'); expect(html).toContain("Move the pointer over the plot for exact values. Click to pin a time."); expect(html).toContain('aria-label="Hide Runtime worker series"'); - expect(html).toContain('aria-label="Timeline start"'); - expect(html).toContain('aria-label="Timeline end"'); - expect(html).toContain("Drag either handle to zoom · drag the selected window to pan"); - expect(html).toContain('aria-label="Pan selected timeline window"'); + expect(html.match(/class="dashboard-runtime-timeline"/g)).toHaveLength(4); + for (const title of ["CPU usage", "Memory usage", "Compute uptime", "Token throughput"]) { + expect(html).toContain(`aria-label="${title} timeline"`); + expect(html).toContain(`aria-label="${title} timeline start"`); + expect(html).toContain(`aria-label="${title} timeline end"`); + expect(html).toContain(`aria-label="Pan ${title} timeline window"`); + } + expect(html).toContain("Drag handles to zoom · window to pan"); }); it("maps pointer positions through horizontal SVG letterboxing", () => { diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index b2da78a74..2bd0b097b 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -123,11 +123,13 @@ function TimeRangeNavigator({ value, onChange, source, + title, }: { domain: TimeWindow; value: TimeWindow; onChange: (next: TimeWindow) => void; source: RuntimeTrendSource; + title: string; }) { const trackRef = useRef(null); const dragRef = useRef(null); @@ -207,11 +209,11 @@ function TimeRangeNavigator({ const fullRange = value.start <= domain.start && value.end >= domain.end; return ( -
+
- Timeline - Drag either handle to zoom · drag the selected window to pan + Time range + Drag handles to zoom · window to pan
@@ -237,7 +239,7 @@ function TimeRangeNavigator({ type="button" className="dashboard-runtime-timeline-selection" style={{ left: `${startPercent}%`, width: `${Math.max(0, endPercent - startPercent)}%` }} - aria-label="Pan selected timeline window" + aria-label={`Pan ${title} timeline window`} title="Drag to pan; use arrow keys for precise movement" disabled={fullRange} onKeyDown={panWindow} @@ -248,7 +250,7 @@ function TimeRangeNavigator({ className="dashboard-runtime-timeline-handle dashboard-runtime-timeline-handle-start" style={{ left: `${startPercent}%` }} role="slider" - aria-label="Timeline start" + aria-label={`${title} timeline start`} aria-valuemin={domain.start} aria-valuemax={value.end - minimumWindow} aria-valuenow={Math.round(value.start)} @@ -261,7 +263,7 @@ function TimeRangeNavigator({ className="dashboard-runtime-timeline-handle dashboard-runtime-timeline-handle-end" style={{ left: `${endPercent}%` }} role="slider" - aria-label="Timeline end" + aria-label={`${title} timeline end`} aria-valuemin={value.start + minimumWindow} aria-valuemax={domain.end} aria-valuenow={Math.round(value.end)} @@ -287,7 +289,6 @@ function TrendChart({ rangeStart, rangeEnd, source, - viewRange, bands = [], ticks = [1, .66, .33, 0], emptyMessage = "Collecting live samples", @@ -302,7 +303,6 @@ function TrendChart({ rangeStart: number; rangeEnd: number; source: RuntimeTrendSource; - viewRange: TimeWindow; bands?: TrendBand[]; ticks?: number[]; emptyMessage?: string; @@ -315,6 +315,9 @@ function TrendChart({ const [hoveredAt, setHoveredAt] = useState(null); const [pinnedAt, setPinnedAt] = useState(null); const [svgSize, setSvgSize] = useState({ width: WIDTH, height: HEIGHT }); + const domain = useMemo(() => ({ start: rangeStart, end: Math.max(rangeStart + 1, rangeEnd) }), [rangeEnd, rangeStart]); + const [viewRange, setViewRange] = useState(domain); + const previousDomain = useRef(domain); const start = Math.max(rangeStart, viewRange.start); const end = Math.min(Math.max(rangeStart + 1, rangeEnd), viewRange.end); const range = Math.max(1, end - start); @@ -359,6 +362,20 @@ function TrendChart({ return () => observer.disconnect(); }, []); + useEffect(() => { + const previous = previousDomain.current; + setViewRange((current) => { + const tolerance = Math.max(1_000, (previous.end - previous.start) / 100); + const wasFullRange = current.start <= previous.start + tolerance && current.end >= previous.end - tolerance; + if (wasFullRange) return domain; + const width = Math.min(current.end - current.start, domain.end - domain.start); + const wasFollowing = current.end >= previous.end - tolerance; + const nextEnd = wasFollowing ? domain.end : clamp(current.end, domain.start + width, domain.end); + return { start: clamp(nextEnd - width, domain.start, domain.end - width), end: nextEnd }; + }); + previousDomain.current = domain; + }, [domain]); + useEffect(() => { if (hoveredAt !== null && (hoveredAt < start || hoveredAt > end)) setHoveredAt(null); if (pinnedAt !== null && (pinnedAt < start || pinnedAt > end)) setPinnedAt(null); @@ -507,6 +524,7 @@ function TrendChart({ )} {!hasLine ?
{visibleSeries.length === 0 ? "All series hidden" : emptyMessage}{visibleSeries.length === 0 ? "Use the legend to show a series" : emptyDetail ?? `${validPoints}/2 valid points · ${visibleSamples.length} snapshots · no history is synthesized`}
: null}
+ "); }); - it("renders smooth honest paths and a memory area only after two real samples", () => { + it("renders uPlot chart mounts and reports trends only after two real samples", () => { const html = renderToStaticMarkup( , ); - expect(html).toContain("dashboard-runtime-trend-line"); - expect(html).toContain(" C "); - expect(html).toContain("dashboard-runtime-trend-area dashboard-runtime-trend-fill-purple"); - expect(html).toContain("dashboard-runtime-trend-latest"); + expect(html.match(/data-chart-engine="uplot"/g)).toHaveLength(4); expect(html).not.toContain("Collecting live samples"); }); - it("exposes interactive series, point selection, and a draggable timeline below every chart", () => { + it("exposes interactive series, point selection, and Grafana-style in-plot range selection", () => { const html = renderToStaticMarkup( , ); expect(html).toContain('role="application"'); - expect(html).toContain("Move the pointer over the plot for exact values. Click to pin a time."); + expect(html).toContain("Drag horizontally to select and zoom a time range."); expect(html).toContain('aria-label="Hide Runtime worker series"'); - expect(html.match(/class="dashboard-runtime-timeline"/g)).toHaveLength(4); - for (const title of ["CPU usage", "Memory usage", "Compute uptime", "Token throughput"]) { - expect(html).toContain(`aria-label="${title} timeline"`); - expect(html).toContain(`aria-label="${title} timeline start"`); - expect(html).toContain(`aria-label="${title} timeline end"`); - expect(html).toContain(`aria-label="Pan ${title} timeline window"`); - } - expect(html).toContain("Drag handles to zoom · window to pan"); - }); - - it("maps pointer positions through horizontal SVG letterboxing", () => { - // A 900 × 220 CSS box renders the 640 × 220 viewBox with a 130px horizontal inset. - expect(runtimeChartViewBoxX(130 + 52, 900, 220)).toBe(52); - expect(runtimeChartViewBoxX(130 + 624, 900, 220)).toBe(624); - expect(runtimeChartRenderedX(52, 900, 220)).toBe(130 + 52); - expect(runtimeChartRenderedX(624, 900, 220)).toBe(130 + 624); + const instructionId = html.match(/aria-describedby="([^"]+-instructions)"/)?.[1]; + expect(instructionId).toBeTruthy(); + expect(html).toContain(`id="${instructionId}"`); + expect(html).not.toContain(">Reset zoom"); + expect(html).not.toContain("dashboard-runtime-timeline"); }); it("names durable charts and buckets without claiming they are live", () => { diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index 2bd0b097b..ef1e6800f 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -5,8 +5,9 @@ import { useRef, useState, type KeyboardEvent as ReactKeyboardEvent, - type PointerEvent as ReactPointerEvent, } from "react"; +import uPlot from "uplot"; +import "uplot/dist/uPlot.min.css"; import { formatDashboardBytes, formatDashboardDuration, formatDashboardTokens } from "./dashboard-model"; import { tokenThroughput, type RuntimeTrendSample } from "./runtime-trends"; @@ -32,21 +33,11 @@ interface TrendBand { export type RuntimeTrendSource = "live" | "durable"; -const WIDTH = 640; -const HEIGHT = 220; -const PLOT = { left: 52, right: 16, top: 22, bottom: 34 }; - interface TimeWindow { start: number; end: number; } -type TimelineDrag = { - kind: "start" | "end" | "window"; - originClientX: number; - originWindow: TimeWindow; -}; - function clamp(value: number, minimum: number, maximum: number): number { return Math.min(maximum, Math.max(minimum, value)); } @@ -59,224 +50,16 @@ function timeLabel(value: number): string { return new Date(value).toLocaleTimeString([], { hour: "2-digit", minute: "2-digit", second: "2-digit" }); } -function segments(points: readonly TrendPoint[]): TrendPoint[][] { - const result: TrendPoint[][] = []; - let current: TrendPoint[] = []; - for (const point of points) { - if (point.value === null || !Number.isFinite(point.value)) { - if (current.length > 0) result.push(current); - current = []; - } else current.push(point); - } - if (current.length > 0) result.push(current); - return result; -} - -function linePath(points: readonly TrendPoint[], x: (value: number) => number, y: (value: number) => number): string { - if (points.length === 0) return ""; - const first = points[0]!; - let result = `M ${x(first.sampledAt)} ${y(first.value ?? 0)}`; - for (let index = 1; index < points.length; index += 1) { - const previous = points[index - 1]!; - const current = points[index]!; - const previousX = x(previous.sampledAt); - const currentX = x(current.sampledAt); - const controlOffset = (currentX - previousX) / 3; - result += ` C ${previousX + controlOffset} ${y(previous.value ?? 0)} ${currentX - controlOffset} ${y(current.value ?? 0)} ${currentX} ${y(current.value ?? 0)}`; - } - return result; -} - -function nearestTime(times: readonly number[], target: number): number | null { - if (times.length === 0) return null; - let nearest = times[0]!; - let distance = Math.abs(nearest - target); - for (let index = 1; index < times.length; index += 1) { - const candidate = times[index]!; - const candidateDistance = Math.abs(candidate - target); - if (candidateDistance < distance) { - nearest = candidate; - distance = candidateDistance; - } - } - return nearest; -} - -export function runtimeChartViewBoxX(clientOffsetX: number, renderedWidth: number, renderedHeight: number): number { - if (renderedWidth <= 0 || renderedHeight <= 0) return PLOT.left; - const scale = Math.min(renderedWidth / WIDTH, renderedHeight / HEIGHT); - const contentWidth = WIDTH * scale; - const horizontalInset = (renderedWidth - contentWidth) / 2; - return clamp((clientOffsetX - horizontalInset) / scale, PLOT.left, WIDTH - PLOT.right); -} - -export function runtimeChartRenderedX(viewBoxX: number, renderedWidth: number, renderedHeight: number): number { - if (renderedWidth <= 0 || renderedHeight <= 0) return 0; - const scale = Math.min(renderedWidth / WIDTH, renderedHeight / HEIGHT); - const contentWidth = WIDTH * scale; - const horizontalInset = (renderedWidth - contentWidth) / 2; - return horizontalInset + clamp(viewBoxX, 0, WIDTH) * scale; -} - -function TimeRangeNavigator({ - domain, - value, - onChange, - source, - title, -}: { - domain: TimeWindow; - value: TimeWindow; - onChange: (next: TimeWindow) => void; - source: RuntimeTrendSource; - title: string; -}) { - const trackRef = useRef(null); - const dragRef = useRef(null); - const [dragging, setDragging] = useState(null); - const span = Math.max(1, domain.end - domain.start); - const step = Math.max(1_000, Math.round(span / 100)); - const minimumWindow = Math.min(span, Math.max(step, 30_000, Math.ceil(span * .1))); - const position = (time: number) => clamp((time - domain.start) / span * 100, 0, 100); - - const timeFromClientX = (clientX: number): number => { - const rect = trackRef.current?.getBoundingClientRect(); - if (!rect || rect.width <= 0) return domain.start; - return domain.start + clamp((clientX - rect.left) / rect.width, 0, 1) * span; - }; - - const updateDrag = (clientX: number) => { - const drag = dragRef.current; - if (!drag) return; - if (drag.kind === "start") { - onChange({ start: clamp(timeFromClientX(clientX), domain.start, value.end - minimumWindow), end: value.end }); - return; - } - if (drag.kind === "end") { - onChange({ start: value.start, end: clamp(timeFromClientX(clientX), value.start + minimumWindow, domain.end) }); - return; - } - const rect = trackRef.current?.getBoundingClientRect(); - if (!rect || rect.width <= 0) return; - const delta = (clientX - drag.originClientX) / rect.width * span; - const width = drag.originWindow.end - drag.originWindow.start; - const start = clamp(drag.originWindow.start + delta, domain.start, domain.end - width); - onChange({ start, end: start + width }); - }; - - const beginDrag = (event: ReactPointerEvent, kind: TimelineDrag["kind"], moveImmediately = false) => { - event.preventDefault(); - event.stopPropagation(); - dragRef.current = { kind, originClientX: event.clientX, originWindow: value }; - setDragging(kind); - event.currentTarget.setPointerCapture(event.pointerId); - if (moveImmediately) updateDrag(event.clientX); - }; - - const endDrag = (event: ReactPointerEvent) => { - if (event.currentTarget.hasPointerCapture(event.pointerId)) event.currentTarget.releasePointerCapture(event.pointerId); - dragRef.current = null; - setDragging(null); - }; - - const adjustHandle = (kind: "start" | "end", event: ReactKeyboardEvent) => { - let delta = 0; - if (event.key === "ArrowLeft" || event.key === "ArrowDown") delta = -step; - else if (event.key === "ArrowRight" || event.key === "ArrowUp") delta = step; - else if (event.key === "Home") delta = kind === "start" ? domain.start - value.start : value.start + minimumWindow - value.end; - else if (event.key === "End") delta = kind === "start" ? value.end - minimumWindow - value.start : domain.end - value.end; - else return; - event.preventDefault(); - if (kind === "start") onChange({ start: clamp(value.start + delta, domain.start, value.end - minimumWindow), end: value.end }); - else onChange({ start: value.start, end: clamp(value.end + delta, value.start + minimumWindow, domain.end) }); - }; - - const panWindow = (event: ReactKeyboardEvent) => { - const width = value.end - value.start; - let start = value.start; - if (event.key === "ArrowLeft" || event.key === "ArrowDown") start -= step; - else if (event.key === "ArrowRight" || event.key === "ArrowUp") start += step; - else if (event.key === "Home") start = domain.start; - else if (event.key === "End") start = domain.end - width; - else return; - event.preventDefault(); - const clampedStart = clamp(start, domain.start, domain.end - width); - onChange({ start: clampedStart, end: clampedStart + width }); - }; - - const startPercent = position(value.start); - const endPercent = position(value.end); - const fullRange = value.start <= domain.start && value.end >= domain.end; +const toneColors: Record = { + orange: "#f59e52", + green: "#50d5a0", + blue: "#78a7ff", + purple: "#b998f4", +}; - return ( -
-
-
- Time range - Drag handles to zoom · window to pan -
-
- - — - - -
-
-
{ - const time = timeFromClientX(event.clientX); - beginDrag(event, Math.abs(time - value.start) <= Math.abs(time - value.end) ? "start" : "end", true); - }} - onPointerMove={(event) => updateDrag(event.clientX)} - onPointerUp={endDrag} - onPointerCancel={endDrag} - onLostPointerCapture={() => { dragRef.current = null; setDragging(null); }} - > -
-
- -
- ); +function withAlpha(hex: string, alpha: number): string { + const value = Number.parseInt(hex.slice(1), 16); + return `rgba(${value >> 16}, ${(value >> 8) & 255}, ${value & 255}, ${alpha})`; } function TrendChart({ @@ -308,105 +91,247 @@ function TrendChart({ emptyMessage?: string; emptyDetail?: string; }) { - const clipId = useId().replaceAll(":", ""); - const instructionId = `${clipId}-instructions`; - const svgRef = useRef(null); + const instructionId = `${useId().replaceAll(":", "")}-instructions`; + const mountRef = useRef(null); + const plotRef = useRef(null); + const pinnedRef = useRef(false); + const zoomedRef = useRef(false); const [hiddenSeries, setHiddenSeries] = useState>(() => new Set()); - const [hoveredAt, setHoveredAt] = useState(null); - const [pinnedAt, setPinnedAt] = useState(null); - const [svgSize, setSvgSize] = useState({ width: WIDTH, height: HEIGHT }); + const [tooltip, setTooltip] = useState<{ idx: number; left: number; pinned: boolean } | null>(null); + const [zoomed, setZoomed] = useState(false); + const [theme, setTheme] = useState(""); const domain = useMemo(() => ({ start: rangeStart, end: Math.max(rangeStart + 1, rangeEnd) }), [rangeEnd, rangeStart]); + const domainRef = useRef(domain); + const hiddenSeriesRef = useRef(hiddenSeries); + domainRef.current = domain; + hiddenSeriesRef.current = hiddenSeries; const [viewRange, setViewRange] = useState(domain); - const previousDomain = useRef(domain); - const start = Math.max(rangeStart, viewRange.start); - const end = Math.min(Math.max(rangeStart + 1, rangeEnd), viewRange.end); - const range = Math.max(1, end - start); - const plotWidth = WIDTH - PLOT.left - PLOT.right; - const plotHeight = HEIGHT - PLOT.top - PLOT.bottom; - const yMaximum = Math.max(1, maximum); - const x = (sampledAt: number) => PLOT.left + (sampledAt - start) / range * plotWidth; - const y = (value: number) => PLOT.top + (1 - Math.min(yMaximum, Math.max(0, value)) / yMaximum) * plotHeight; - const visibleSeries = series.filter((entry) => !hiddenSeries.has(entry.id)).map((entry) => ({ - ...entry, - points: entry.points.filter((point) => point.sampledAt >= start && point.sampledAt <= end), - })); - const visibleSamples = samples.filter((sample) => sample.sampledAt >= start && sample.sampledAt <= end); - const hasLine = visibleSeries.some((entry) => segments(entry.points).some((segment) => segment.length >= 2)); + const visibleSeries = series.filter((entry) => !hiddenSeries.has(entry.id)); + const hasLine = visibleSeries.some((entry) => entry.points.some((point, index) => ( + point.value !== null && Number.isFinite(point.value) + && index > 0 + && entry.points[index - 1]?.value !== null + && Number.isFinite(entry.points[index - 1]?.value) + ))); const validPoints = Math.max(0, ...visibleSeries.map((entry) => entry.points.filter((point) => ( point.value !== null && Number.isFinite(point.value) )).length)); - const xTicks = [start, start + range / 2, end]; - const interactiveTimes = [...new Set(visibleSeries.flatMap((entry) => entry.points - .map((point) => point.sampledAt)))].sort((left, right) => left - right); - const selectedAt = pinnedAt ?? hoveredAt; - const selectedValues = selectedAt === null ? [] : visibleSeries.map((entry) => ({ - ...entry, - value: entry.points.find((point) => point.sampledAt === selectedAt)?.value ?? null, - })); - const selectedX = selectedAt === null ? null : x(selectedAt); - const tooltipLeft = selectedX === null - ? 50 - : runtimeChartRenderedX(selectedX, svgSize.width, svgSize.height) / Math.max(1, svgSize.width) * 100; + const times = useMemo(() => [...new Set(series.flatMap((entry) => entry.points.map((point) => point.sampledAt)))].sort((left, right) => left - right), [series]); + const chartData = useMemo(() => [ + times.map((value) => value / 1_000), + ...series.map((entry) => { + const values = new Map(entry.points.map((point) => [point.sampledAt, point.value])); + return times.map((value) => values.get(value) ?? null); + }), + ], [series, times]); + const dataRef = useRef(chartData); + const formatRef = useRef(formatValue); + const maximumRef = useRef(maximum); + dataRef.current = chartData; + formatRef.current = formatValue; + maximumRef.current = maximum; + const seriesKey = series.map((entry) => `${entry.id}:${entry.label}:${entry.tone}:${entry.fill ? 1 : 0}`).join("|"); + const bandsKey = bands.map((band) => `${band.from}:${band.to}:${band.tone}`).join("|"); + const ticksKey = ticks.join(":"); + const selectedTimestamp = tooltip === null ? null : chartData[0][tooltip.idx]; + const selectedAt = selectedTimestamp === null || selectedTimestamp === undefined ? null : selectedTimestamp * 1_000; + const selectedValues = tooltip === null ? [] : visibleSeries.map((entry) => { + const seriesIndex = series.findIndex((candidate) => candidate.id === entry.id); + return { ...entry, value: chartData[seriesIndex + 1]?.[tooltip.idx] ?? null }; + }); useEffect(() => { - const svg = svgRef.current; - if (!svg) return; - const updateSize = () => { - const rect = svg.getBoundingClientRect(); - if (rect.width > 0 && rect.height > 0) setSvgSize({ width: rect.width, height: rect.height }); - }; - updateSize(); - if (typeof ResizeObserver === "undefined") return; - const observer = new ResizeObserver(updateSize); - observer.observe(svg); + const update = () => setTheme(document.documentElement.dataset.theme ?? "light"); + update(); + const observer = new MutationObserver(update); + observer.observe(document.documentElement, { attributes: true, attributeFilter: ["data-theme"] }); return () => observer.disconnect(); }, []); useEffect(() => { - const previous = previousDomain.current; - setViewRange((current) => { - const tolerance = Math.max(1_000, (previous.end - previous.start) / 100); - const wasFullRange = current.start <= previous.start + tolerance && current.end >= previous.end - tolerance; - if (wasFullRange) return domain; - const width = Math.min(current.end - current.start, domain.end - domain.start); - const wasFollowing = current.end >= previous.end - tolerance; - const nextEnd = wasFollowing ? domain.end : clamp(current.end, domain.start + width, domain.end); - return { start: clamp(nextEnd - width, domain.start, domain.end - width), end: nextEnd }; + const mount = mountRef.current; + if (!mount || theme === "") return; + const styles = getComputedStyle(mount); + const color = (name: string, fallback: string) => styles.getPropertyValue(name).trim() || fallback; + const axisColor = color("--fg-muted", theme === "dark" ? "#a8adb8" : "#6f7480"); + const gridColor = color("--line", theme === "dark" ? "#30343b" : "#e3e5e8"); + const surfaceColor = color("--surface", theme === "dark" ? "#17191d" : "#ffffff"); + const yMaximum = Math.max(1, maximum); + const bandColors: Record = { + safe: withAlpha("#50d5a0", .07), + warning: withAlpha("#f59e52", .07), + danger: withAlpha("#ef6a72", .07), + }; + const options: uPlot.Options = { + width: Math.max(320, mount.clientWidth), + height: 220, + padding: [10, 8, 0, 0], + legend: { show: false }, + scales: { + x: { time: true }, + y: { auto: false, range: [0, yMaximum] }, + }, + axes: [ + { + stroke: axisColor, + grid: { stroke: gridColor, width: 1 }, + ticks: { stroke: gridColor, width: 1 }, + values: (_plot, values) => values.map((value) => timeLabel(value * 1_000)), + font: "8px ui-monospace, SFMono-Regular, Menlo, monospace", + size: 28, + }, + { + stroke: axisColor, + grid: { stroke: gridColor, width: 1 }, + ticks: { stroke: gridColor, width: 1 }, + splits: () => ticks.map((tick) => Math.max(1, maximumRef.current) * tick).sort((left, right) => left - right), + values: (_plot, values) => values.map((value) => formatRef.current(value)), + font: "8px ui-monospace, SFMono-Regular, Menlo, monospace", + size: 52, + }, + ], + cursor: { + x: true, + y: true, + lock: false, + drag: { x: true, y: false, setScale: true, dist: 8 }, + points: { size: 7, width: 2, fill: surfaceColor }, + }, + select: { show: true, left: 0, top: 0, width: 0, height: 0 }, + series: [ + {}, + ...series.map((entry): uPlot.Series => ({ + label: entry.label, + show: !hiddenSeriesRef.current.has(entry.id), + stroke: toneColors[entry.tone], + fill: entry.fill ? withAlpha(toneColors[entry.tone], .13) : undefined, + width: 2, + spanGaps: false, + points: { show: false }, + })), + ], + hooks: { + drawClear: bands.length === 0 ? [] : [(plot) => { + for (const band of bands) { + const top = plot.valToPos(band.to, "y", true); + const bottom = plot.valToPos(band.from, "y", true); + plot.ctx.fillStyle = bandColors[band.tone]; + plot.ctx.fillRect(plot.bbox.left, top, plot.bbox.width, Math.max(0, bottom - top)); + } + }], + setCursor: [(plot) => { + const idx = plot.cursor.idx; + if (idx === null || idx === undefined || pinnedRef.current) return; + const left = ((plot.cursor.left ?? 0) + plot.bbox.left / uPlot.pxRatio) / Math.max(1, plot.width) * 100; + setTooltip({ idx, left: clamp(left, 18, 82), pinned: pinnedRef.current }); + }], + setScale: [(plot, scaleKey) => { + if (scaleKey !== "x") return; + const scale = plot.scales.x; + if (!scale || scale.min === undefined || scale.max === undefined) return; + const currentDomain = domainRef.current; + const tolerance = Math.max(1, (currentDomain.end - currentDomain.start) / 100_000); + const isZoomed = Math.abs(scale.min * 1_000 - currentDomain.start) > tolerance || Math.abs(scale.max * 1_000 - currentDomain.end) > tolerance; + zoomedRef.current = isZoomed; + setZoomed(isZoomed); + setViewRange({ start: scale.min * 1_000, end: scale.max * 1_000 }); + }], + }, + }; + const plot = new uPlot(options, dataRef.current, mount); + plotRef.current = plot; + plot.setScale("x", { min: domain.start / 1_000, max: domain.end / 1_000 }); + let pointerStart = 0; + let moved = false; + const pointerDown = (event: PointerEvent) => { pointerStart = event.clientX; moved = false; }; + const pointerMove = (event: PointerEvent) => { + if (moved || Math.abs(event.clientX - pointerStart) < 8) return; + moved = true; + pinnedRef.current = false; + setTooltip(null); + }; + const click = () => { + if (moved || plot.cursor.idx === null || plot.cursor.idx === undefined) return; + pinnedRef.current = !pinnedRef.current; + setTooltip((current) => current === null ? null : { ...current, pinned: pinnedRef.current }); + }; + const leave = () => { if (!pinnedRef.current) setTooltip(null); }; + const reset = (event: MouseEvent) => { + event.preventDefault(); + pinnedRef.current = false; + setTooltip(null); + const currentDomain = domainRef.current; + plot.setScale("x", { min: currentDomain.start / 1_000, max: currentDomain.end / 1_000 }); + }; + plot.over.addEventListener("pointerdown", pointerDown); + plot.over.addEventListener("pointermove", pointerMove); + plot.over.addEventListener("click", click); + plot.over.addEventListener("mouseleave", leave); + plot.over.addEventListener("dblclick", reset); + const resize = new ResizeObserver(() => { + const width = mount.clientWidth; + if (width > 0 && width !== plot.width) plot.setSize({ width, height: 220 }); }); - previousDomain.current = domain; - }, [domain]); + resize.observe(mount); + return () => { + resize.disconnect(); + plot.destroy(); + if (plotRef.current === plot) plotRef.current = null; + }; + // Data updates are applied without rebuilding so a selected time range remains stable. + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [bandsKey, seriesKey, source, theme, ticksKey]); useEffect(() => { - if (hoveredAt !== null && (hoveredAt < start || hoveredAt > end)) setHoveredAt(null); - if (pinnedAt !== null && (pinnedAt < start || pinnedAt > end)) setPinnedAt(null); - }, [end, hoveredAt, pinnedAt, start]); + const plot = plotRef.current; + if (!plot) return; + plot.setData(chartData, false); + plot.setScale("y", { min: 0, max: Math.max(1, maximum) }); + if (!zoomedRef.current) plot.setScale("x", { min: domain.start / 1_000, max: domain.end / 1_000 }); + }, [chartData, domain, maximum]); - const selectFromPointer = (event: { currentTarget: SVGSVGElement; clientX: number }): number | null => { - const rect = event.currentTarget.getBoundingClientRect(); - if (rect.width <= 0 || rect.height <= 0) return null; - const viewBoxX = runtimeChartViewBoxX(event.clientX - rect.left, rect.width, rect.height); - return nearestTime(interactiveTimes, start + (viewBoxX - PLOT.left) / plotWidth * range); + useEffect(() => { + const plot = plotRef.current; + if (!plot) return; + series.forEach((entry, index) => plot.setSeries(index + 1, { show: !hiddenSeries.has(entry.id) })); + }, [hiddenSeries, series]); + + const resetZoom = () => { + zoomedRef.current = false; + setZoomed(false); + setViewRange(domain); + plotRef.current?.setScale("x", { min: domain.start / 1_000, max: domain.end / 1_000 }); }; - const handleKeyboard = (event: ReactKeyboardEvent) => { + const handleKeyboard = (event: ReactKeyboardEvent) => { if (event.key === "Escape") { - setPinnedAt(null); - setHoveredAt(null); + pinnedRef.current = false; + setTooltip(null); return; } - if (interactiveTimes.length === 0) return; - const current = selectedAt === null ? interactiveTimes.length - 1 : interactiveTimes.indexOf(selectedAt); + const visibleIndices = times.flatMap((time, index) => time >= viewRange.start && time <= viewRange.end ? [index] : []); + if (visibleIndices.length === 0) return; + const currentPosition = tooltip === null ? visibleIndices.length - 1 : visibleIndices.indexOf(tooltip.idx); if (event.key === "ArrowLeft" || event.key === "ArrowRight") { event.preventDefault(); const offset = event.key === "ArrowLeft" ? -1 : 1; - const next = interactiveTimes[clamp(current + offset, 0, interactiveTimes.length - 1)]!; - setPinnedAt(next); - setHoveredAt(next); + const position = currentPosition < 0 + ? visibleIndices.length - 1 + : clamp(currentPosition + offset, 0, visibleIndices.length - 1); + const idx = visibleIndices[position]!; + const plot = plotRef.current; + const plotLeft = plot?.valToPos(times[idx]! / 1_000, "x") ?? 0; + const left = plot === null + ? 50 + : (plotLeft + plot.bbox.left / uPlot.pxRatio) / Math.max(1, plot.width) * 100; + pinnedRef.current = true; + if (plot) plot.setCursor({ left: plotLeft, top: 0 }); + setTooltip({ idx, left: clamp(left, 18, 82), pinned: true }); } else if (event.key === "Enter" || event.key === " ") { event.preventDefault(); - const next = selectedAt ?? interactiveTimes.at(-1)!; - setPinnedAt(pinnedAt === next ? null : next); - setHoveredAt(next); + pinnedRef.current = !pinnedRef.current; + const idx = tooltip?.idx ?? visibleIndices.at(-1)!; + setTooltip({ idx, left: tooltip?.left ?? 50, pinned: pinnedRef.current }); } }; @@ -437,94 +362,40 @@ function TrendChart({ >{entry.label} ); })} + {zoomed ? : null} -
-

Move the pointer over the plot for exact values. Click to pin a time. Use Left and Right arrows to move the pinned selection, and Escape to clear it.

- +

Move the pointer over the plot for exact values. Drag horizontally to select and zoom a time range. Double-click or use Reset zoom to restore the full range. Click to pin a time. Use Left and Right arrows to move the pinned selection, and Escape to clear it.

+
{ - const nearest = selectFromPointer(event); - if (nearest !== null) setHoveredAt(nearest); - }} - onPointerLeave={() => { if (pinnedAt === null) setHoveredAt(null); }} - onClick={(event) => { - const nearest = selectFromPointer(event); - if (nearest !== null) { - setHoveredAt(nearest); - setPinnedAt((current) => current === nearest ? null : nearest); - } - }} - onKeyDown={handleKeyboard} - > - - {bands.map((band) => ( - - ))} - {ticks.map((tick) => { - const value = yMaximum * tick; - return ( - - - {formatValue(value)} - - ); - })} - {visibleSamples.length > 0 ? xTicks.map((tick, index) => ( - {timeLabel(tick)} - )) : null} - - {visibleSeries.flatMap((entry) => segments(entry.points).flatMap((segment, index) => { - const key = `${entry.id}:${index}`; - if (segment.length === 1) { - return ; - } - const path = linePath(segment, x, y); - const last = segment.at(-1)!; - return [ - entry.fill ? ( - - ) : null, - , - , - ]; - }))} - {selectedX === null ? null : ( - <> - - {selectedValues.map((entry) => entry.value === null ? null : ( - - ))} - - )} - - + data-chart-engine="uplot" + data-view-start={Math.round(viewRange.start)} + data-view-end={Math.round(viewRange.end)} + data-selected-at={selectedAt === null ? undefined : Math.round(selectedAt)} + /> {selectedAt === null ? null : (
-
{pinnedAt === null ? Hover : Pinned}
+
{tooltip?.pinned ? Pinned : Hover}
{selectedValues.map((entry) => (
{entry.label}{entry.value === null ? "Unavailable" : formatValue(entry.value)}
))}
)} - {!hasLine ?
{visibleSeries.length === 0 ? "All series hidden" : emptyMessage}{visibleSeries.length === 0 ? "Use the legend to show a series" : emptyDetail ?? `${validPoints}/2 valid points · ${visibleSamples.length} snapshots · no history is synthesized`}
: null} + {!hasLine ?
{visibleSeries.length === 0 ? "All series hidden" : emptyMessage}{visibleSeries.length === 0 ? "Use the legend to show a series" : emptyDetail ?? `${validPoints}/2 valid points · ${samples.length} snapshots · no history is synthesized`}
: null}
-
{hasLine ? `${title} ${source} trend available` @@ -595,32 +613,14 @@ export function RuntimeTrendCharts({ const tokenMaximum = Math.max(1, ...finite(charts.tokens.flatMap((series) => series.points.map((point) => point.value)))); const newest = rangeEnd ?? samples.at(-1)?.sampledAt ?? Date.now(); const oldest = rangeStart ?? samples[0]?.sampledAt ?? newest - 60 * 60 * 1_000; - const domain = useMemo(() => ({ start: oldest, end: Math.max(oldest + 1, newest) }), [oldest, newest]); - const [viewRange, setViewRange] = useState(domain); - const previousDomain = useRef(domain); const durable = source === "durable"; - useEffect(() => { - const previous = previousDomain.current; - setViewRange((current) => { - const tolerance = Math.max(1_000, (previous.end - previous.start) / 100); - const wasFullRange = current.start <= previous.start + tolerance && current.end >= previous.end - tolerance; - if (wasFullRange) return domain; - const width = Math.min(current.end - current.start, domain.end - domain.start); - const wasFollowing = current.end >= previous.end - tolerance; - const end = wasFollowing ? domain.end : clamp(current.end, domain.start + width, domain.end); - return { start: clamp(end - width, domain.start, domain.end - width), end }; - }); - previousDomain.current = domain; - }, [domain]); - return (
- `${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> - formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> - formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> - `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} viewRange={viewRange} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} /> - + `${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> + formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> + formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> + `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} />
); } From e7c6f7a9085cdc7e32996fa0d0625d99da490fea Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 15:59:32 +0800 Subject: [PATCH 30/51] Adopt Grafana-style runtime chart zoom --- apps/web/e2e/agents-lifecycle.spec.ts | 70 +- apps/web/package.json | 3 +- .../src/features/dashboard/DashboardView.css | 224 ++----- .../dashboard/RuntimeTrendCharts.test.tsx | 34 +- .../features/dashboard/RuntimeTrendCharts.tsx | 617 +++++++----------- pnpm-lock.yaml | 8 + 6 files changed, 348 insertions(+), 608 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 433c59d74..2e9ace74d 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2646,6 +2646,10 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard).not.toContainText("Collecting live samples"); const cpuChart = dashboard.getByLabel("CPU usage: 3 live samples"); const cpuCard = cpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); + await cpuChart.focus(); + await cpuChart.press("Enter"); + await expect(cpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Pinned"); + await cpuChart.press("Escape"); await cpuChart.hover({ position: { x: 260, y: 90 } }); await expect(cpuCard.locator(".dashboard-runtime-trend-tooltip")).toBeVisible(); await cpuChart.click({ position: { x: 260, y: 90 } }); @@ -2661,35 +2665,36 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await cpuSeriesToggle.click(); await expect(cpuSeriesToggle).toHaveAttribute("aria-pressed", "true"); - await expect(dashboard.locator(".dashboard-runtime-timeline")).toHaveCount(4); - const chartTimeline = cpuCard.getByRole("region", { name: "CPU usage timeline" }); - const timelineStart = chartTimeline.getByRole("slider", { name: "CPU usage timeline start" }); - const initialTimelineStart = Number(await timelineStart.getAttribute("aria-valuenow")); - await timelineStart.scrollIntoViewIfNeeded(); - const startHandleBox = await timelineStart.boundingBox(); - expect(startHandleBox).not.toBeNull(); - expect(startHandleBox!.width).toBeGreaterThanOrEqual(24); - expect(startHandleBox!.height).toBeGreaterThanOrEqual(24); - await page.mouse.move(startHandleBox!.x + startHandleBox!.width / 2, startHandleBox!.y + startHandleBox!.height / 2); + await expect(dashboard.locator('[data-chart-engine="uplot"]')).toHaveCount(4); + await expect(dashboard.locator(".dashboard-runtime-timeline")).toHaveCount(0); + const initialViewStart = Number(await cpuChart.getAttribute("data-view-start")); + const initialViewEnd = Number(await cpuChart.getAttribute("data-view-end")); + const plot = cpuCard.locator(".u-over"); + const plotBox = await plot.boundingBox(); + expect(plotBox).not.toBeNull(); + await page.mouse.move(plotBox!.x + plotBox!.width * .55, plotBox!.y + plotBox!.height / 2); await page.mouse.down(); - await page.mouse.move(startHandleBox!.x + 90, startHandleBox!.y + startHandleBox!.height / 2, { steps: 8 }); + await page.mouse.move(plotBox!.x + plotBox!.width, plotBox!.y + plotBox!.height / 2, { steps: 8 }); await page.mouse.up(); - await expect.poll(async () => Number(await timelineStart.getAttribute("aria-valuenow"))).toBeGreaterThan(initialTimelineStart); - const resizedTimelineStart = Number(await timelineStart.getAttribute("aria-valuenow")); - await expect(cpuChart).toHaveAttribute("data-view-start", String(resizedTimelineStart)); + await expect.poll(async () => Number(await cpuChart.getAttribute("data-view-start"))).toBeGreaterThan(initialViewStart); + await expect.poll(async () => Number(await cpuChart.getAttribute("data-view-end"))).toBeLessThanOrEqual(initialViewEnd); + const zoomedViewStart = Number(await cpuChart.getAttribute("data-view-start")); + const zoomedViewEnd = Number(await cpuChart.getAttribute("data-view-end")); + await cpuChart.focus(); + await cpuChart.press("Escape"); + await cpuChart.press("ArrowRight"); + const keyboardSelectedAt = Number(await cpuChart.getAttribute("data-selected-at")); + expect(keyboardSelectedAt).toBeGreaterThanOrEqual(zoomedViewStart); + expect(keyboardSelectedAt).toBeLessThanOrEqual(zoomedViewEnd); for (const chartName of ["Memory usage", "Compute uptime", "Token throughput"]) { - await expect(dashboard.getByLabel(`${chartName}: 3 live samples`)).toHaveAttribute("data-view-start", String(initialTimelineStart)); + await expect(dashboard.getByLabel(`${chartName}: 3 live samples`)).toHaveAttribute("data-view-start", String(initialViewStart)); + await expect(dashboard.getByLabel(`${chartName}: 3 live samples`)).toHaveAttribute("data-view-end", String(initialViewEnd)); } - const timelineWindow = chartTimeline.getByRole("button", { name: "Pan CPU usage timeline window" }); - const timelineWindowBox = await timelineWindow.boundingBox(); - expect(timelineWindowBox).not.toBeNull(); - expect(timelineWindowBox!.height).toBeGreaterThanOrEqual(24); - await timelineWindow.focus(); - await timelineWindow.press("ArrowLeft"); - await expect.poll(async () => Number(await timelineStart.getAttribute("aria-valuenow"))).toBeLessThan(resizedTimelineStart); - await expect(chartTimeline.getByRole("button", { name: "Reset" })).toBeEnabled(); - await chartTimeline.getByRole("button", { name: "Reset" }).click(); - await expect(timelineStart).toHaveAttribute("aria-valuenow", String(initialTimelineStart)); + const resetZoom = cpuCard.getByRole("button", { name: "Reset zoom" }); + await expect(resetZoom).toBeEnabled(); + await resetZoom.click(); + await expect(cpuChart).toHaveAttribute("data-view-start", String(initialViewStart)); + await expect(cpuChart).toHaveAttribute("data-view-end", String(initialViewEnd)); await expect(dashboard.getByRole("table", { name: "Runtime targets" })).not.toBeVisible(); for (const close of await page.getByRole("button", { name: "Close notification" }).all()) await close.click(); await expect(page.getByRole("button", { name: "Close notification" })).toHaveCount(0); @@ -2834,13 +2839,14 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await durableCpuChart.press("ArrowLeft"); const durableCpuCard = durableCpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); await expect(durableCpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Unavailable"); - const durableTimelineStart = durableCpuCard.getByRole("slider", { name: "CPU usage timeline start" }); - const durableInitialStart = Number(await durableTimelineStart.getAttribute("aria-valuenow")); - await durableTimelineStart.focus(); - await durableTimelineStart.press("ArrowRight"); - await expect.poll(async () => Number(await durableTimelineStart.getAttribute("aria-valuenow"))).toBeGreaterThan(durableInitialStart); - const durableViewStart = await durableTimelineStart.getAttribute("aria-valuenow"); - await expect(durableCpuChart).toHaveAttribute("data-view-start", durableViewStart!); + const durableInitialStart = Number(await durableCpuChart.getAttribute("data-view-start")); + const durablePlotBox = await durableCpuCard.locator(".u-over").boundingBox(); + expect(durablePlotBox).not.toBeNull(); + await page.mouse.move(durablePlotBox!.x + durablePlotBox!.width * .2, durablePlotBox!.y + durablePlotBox!.height / 2); + await page.mouse.down(); + await page.mouse.move(durablePlotBox!.x + durablePlotBox!.width * .8, durablePlotBox!.y + durablePlotBox!.height / 2, { steps: 8 }); + await page.mouse.up(); + await expect.poll(async () => Number(await durableCpuChart.getAttribute("data-view-start"))).toBeGreaterThan(durableInitialStart); await page.reload(); await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); diff --git a/apps/web/package.json b/apps/web/package.json index 2a7f1da60..47db10c42 100644 --- a/apps/web/package.json +++ b/apps/web/package.json @@ -15,7 +15,8 @@ "lucide-react": "^1.22.0", "react": "^19.2.7", "react-dom": "^19.2.7", - "react-markdown": "^10.1.0" + "react-markdown": "^10.1.0", + "uplot": "^1.6.32" }, "devDependencies": { "@types/react": "^19.2.17", diff --git a/apps/web/src/features/dashboard/DashboardView.css b/apps/web/src/features/dashboard/DashboardView.css index 6a0d0dabd..994880b84 100644 --- a/apps/web/src/features/dashboard/DashboardView.css +++ b/apps/web/src/features/dashboard/DashboardView.css @@ -1224,6 +1224,18 @@ text-decoration: line-through; } +.dashboard-runtime-trend-legend .dashboard-runtime-chart-reset { + padding: 2px 6px; + color: var(--accent); + border: 1px solid color-mix(in srgb, var(--accent) 30%, var(--line)); +} + +.dashboard-runtime-trend-legend .dashboard-runtime-chart-reset:disabled { + color: var(--fg-muted); + cursor: default; + opacity: .42; +} + .dashboard-runtime-trend-legend i, .dashboard-runtime-trend-tooltip i { width: 11px; @@ -1280,6 +1292,40 @@ border-radius: 4px; } +.dashboard-runtime-uplot { + width: 100%; + height: 220px; + color: var(--fg-muted); + outline: none; + touch-action: pan-y; +} + +.dashboard-runtime-uplot:focus-visible { + outline: 2px solid color-mix(in srgb, var(--accent) 48%, transparent); + outline-offset: 2px; + border-radius: 4px; +} + +.dashboard-runtime-uplot .uplot { + width: 100% !important; +} + +.dashboard-runtime-uplot .u-over { + cursor: crosshair; + touch-action: pan-y; +} + +.dashboard-runtime-uplot .u-select { + background: color-mix(in srgb, var(--accent) 16%, transparent); + border-right: 1px dashed color-mix(in srgb, var(--accent) 78%, transparent); + border-left: 1px dashed color-mix(in srgb, var(--accent) 78%, transparent); +} + +.dashboard-runtime-uplot .u-cursor-x, +.dashboard-runtime-uplot .u-cursor-y { + border-color: color-mix(in srgb, var(--fg) 48%, transparent); +} + .dashboard-runtime-trend-gridline { stroke: color-mix(in srgb, var(--line) 88%, transparent); stroke-width: 1; @@ -1464,167 +1510,6 @@ border: 0; } -.dashboard-runtime-timeline { - min-width: 0; - padding: 10px 15px 9px; - margin: 2px -15px -10px; - background: color-mix(in srgb, var(--surface-subtle) 54%, var(--surface)); - border-top: 1px solid var(--line); - border-radius: 0 0 8px 8px; -} - -.dashboard-runtime-timeline > header, -.dashboard-runtime-timeline > header > div { - display: flex; - align-items: center; -} - -.dashboard-runtime-timeline > header { - justify-content: space-between; - gap: 16px; -} - -.dashboard-runtime-timeline > header > div { - min-width: 0; - gap: 7px; -} - -.dashboard-runtime-timeline > header strong { - font-size: 9px; -} - -.dashboard-runtime-timeline > header span, -.dashboard-runtime-timeline > header time, -.dashboard-runtime-timeline-bounds { - color: var(--fg-muted); - font-family: var(--font-mono); - font-size: 8px; - font-variant-numeric: tabular-nums; -} - -.dashboard-runtime-timeline > header button { - padding: 3px 7px; - color: var(--fg-muted); - font-family: var(--font-mono); - font-size: 8px; - background: var(--surface-subtle); - border: 1px solid var(--line); - border-radius: 5px; - cursor: pointer; -} - -.dashboard-runtime-timeline > header button:hover:not(:disabled), -.dashboard-runtime-timeline > header button:focus-visible { - color: var(--fg); - border-color: color-mix(in srgb, var(--accent) 45%, var(--line)); -} - -.dashboard-runtime-timeline > header button:disabled { - cursor: default; - opacity: .45; -} - -.dashboard-runtime-timeline-track { - height: 36px; - margin: 6px 7px 0; - position: relative; - cursor: pointer; - touch-action: none; - user-select: none; -} - -.dashboard-runtime-timeline-rail { - height: 6px; - position: absolute; - top: 15px; - border-radius: 999px; -} - -.dashboard-runtime-timeline-rail { - right: 0; - left: 0; - background: color-mix(in srgb, var(--line) 78%, transparent); -} - -.dashboard-runtime-timeline-selection { - height: 28px; - position: absolute; - z-index: 1; - top: 4px; - min-width: 1px; - padding: 0; - background: transparent; - border: 0; - cursor: grab; -} - -.dashboard-runtime-timeline-selection::before { - height: 6px; - position: absolute; - top: 11px; - right: 0; - left: 0; - content: ""; - background: color-mix(in srgb, var(--accent) 58%, var(--surface)); - border-radius: 999px; -} - -.dashboard-runtime-timeline-selection:disabled { - cursor: default; -} - -.dashboard-runtime-timeline-selection:focus-visible { - outline: 2px solid color-mix(in srgb, var(--accent) 42%, transparent); - outline-offset: 3px; -} - -.dashboard-runtime-timeline-track.is-dragging .dashboard-runtime-timeline-selection { - cursor: grabbing; -} - -.dashboard-runtime-timeline-handle { - width: 28px; - height: 28px; - padding: 0; - position: absolute; - z-index: 2; - top: 4px; - background: var(--surface); - border: 2px solid var(--accent); - border-radius: 5px; - box-shadow: var(--shadow-control); - cursor: ew-resize; - transform: translateX(-50%); -} - -.dashboard-runtime-timeline-handle::after { - width: 2px; - height: 10px; - position: absolute; - top: 7px; - left: 9px; - content: ""; - background: color-mix(in srgb, var(--accent) 70%, var(--fg)); - box-shadow: 5px 0 0 color-mix(in srgb, var(--accent) 70%, var(--fg)); -} - -.dashboard-runtime-timeline-handle:hover, -.dashboard-runtime-timeline-handle:focus-visible { - background: color-mix(in srgb, var(--accent) 10%, var(--surface)); - outline: 2px solid color-mix(in srgb, var(--accent) 35%, transparent); - outline-offset: 2px; -} - -.dashboard-runtime-timeline-bounds { - display: flex; - justify-content: space-between; - padding: 0 2px; -} - -.dashboard-runtime-timeline-bounds span { - text-transform: uppercase; - letter-spacing: .04em; -} .dashboard-runtime-explorer { min-width: 0; @@ -2364,23 +2249,6 @@ justify-content: flex-start; } - .dashboard-runtime-timeline > header { - align-items: flex-start; - flex-direction: column; - gap: 6px; - } - - .dashboard-runtime-timeline > header > div:first-child { - align-items: flex-start; - flex-direction: column; - gap: 1px; - } - - .dashboard-runtime-timeline > header > div:last-child { - width: 100%; - justify-content: space-between; - } - .dashboard-runtime-visual-card { min-height: 224px; } diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index d37872a34..1d44e0bab 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -1,7 +1,7 @@ import { renderToStaticMarkup } from "react-dom/server"; import { describe, expect, it } from "vitest"; -import { RuntimeTrendCharts, runtimeChartRenderedX, runtimeChartViewBoxX } from "./RuntimeTrendCharts"; +import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; import type { RuntimeTrendSample } from "./runtime-trends"; function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample { @@ -32,42 +32,28 @@ describe("Runtime live-window chart accessibility", () => { expect(html).not.toContain("Runtime worker
50%1
{hasLine ? `${title} ${source} trend available` diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index d46c38103..ebd742003 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -38,6 +38,9 @@ importers: react-markdown: specifier: ^10.1.0 version: 10.1.0(@types/react@19.3.0)(react@19.3.0) + uplot: + specifier: ^1.6.32 + version: 1.6.32 devDependencies: '@types/react': specifier: ^19.2.17 @@ -1415,6 +1418,9 @@ packages: resolution: {integrity: sha512-pjy2bYhSsufwWlKwPc+l3cN7+wuJlK6uz0YdJEOlQDbl6jo/YlPi4mb8agUkVC8BF7V8NuzeyPNqRksA3hztKQ==} engines: {node: '>= 0.8'} + uplot@1.6.32: + resolution: {integrity: sha512-KIMVnG68zvu5XXUbC4LQEPnhwOxBuLyW1AHtpm6IKTXImkbLgkMy+jabjLgSLMasNuGGzQm/ep3tOkyTxpiQIw==} + vary@1.1.2: resolution: {integrity: sha512-BNGbWLfd0eUPabhkXUVm0j8uuvREyTh5ovRa/dyow/BqAbZJyC+5fU+IzQOzmAKzYqYRAISoRhdQr3eIZ/PXqg==} engines: {node: '>= 0.8'} @@ -2853,6 +2859,8 @@ snapshots: unpipe@1.0.0: {} + uplot@1.6.32: {} + vary@1.1.2: {} vfile-message@4.0.3: From b32a8c42e8c70089fdc089abc9bdd32fa61ec903 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 16:33:19 +0800 Subject: [PATCH 31/51] Show sparse runtime samples explicitly --- .../dashboard/RuntimeTrendCharts.test.tsx | 30 ++++++- .../features/dashboard/RuntimeTrendCharts.tsx | 78 ++++++++++++++++--- 2 files changed, 96 insertions(+), 12 deletions(-) diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index 1d44e0bab..c63382c37 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -1,7 +1,7 @@ import { renderToStaticMarkup } from "react-dom/server"; import { describe, expect, it } from "vitest"; -import { RuntimeTrendCharts } from "./RuntimeTrendCharts"; +import { RuntimeTrendCharts, runtimeChartCaption, runtimeChartShowsSparsePoints } from "./RuntimeTrendCharts"; import type { RuntimeTrendSample } from "./runtime-trends"; function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample { @@ -23,6 +23,25 @@ function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample } describe("Runtime live-window chart accessibility", () => { + it("shows isolated or sparse values as points without inventing continuity", () => { + expect(runtimeChartShowsSparsePoints([null, 512, null])).toBe(true); + expect(runtimeChartShowsSparsePoints([512, 768])).toBe(true); + expect(runtimeChartShowsSparsePoints(Array.from({ length: 13 }, (_, index) => index))).toBe(false); + expect(runtimeChartShowsSparsePoints([null, null])).toBe(false); + }); + + it("announces hidden series instead of claiming retained data is empty", () => { + expect(runtimeChartCaption({ + title: "Memory usage", + source: "durable", + hasLine: false, + allSeriesHidden: true, + validPoints: 0, + sampleCount: 24, + emptyMessage: "No complete retained memory samples", + })).toBe("Memory usage all series hidden; use the legend to show a series"); + }); + it("reports a current gap as unavailable instead of announcing a stale value as latest", () => { const html = renderToStaticMarkup( , @@ -65,4 +84,13 @@ describe("Runtime live-window chart accessibility", () => { expect(html).toContain('aria-label="CPU usage: 2 retained buckets"'); expect(html).not.toContain('aria-label="CPU usage: 2 live samples"'); }); + + it("announces an isolated durable value as sparse rather than empty", () => { + const html = renderToStaticMarkup( + , + ); + + expect(html).toContain("Memory usage durable trend has 1 sparse valid point; a line requires consecutive buckets"); + expect(html).not.toContain("Memory usage No complete retained memory samples"); + }); }); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index ef1e6800f..8c216695d 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -22,7 +22,6 @@ interface TrendSeries { label: string; tone: "orange" | "green" | "blue" | "purple"; points: TrendPoint[]; - fill?: boolean; } interface TrendBand { @@ -62,6 +61,51 @@ function withAlpha(hex: string, alpha: number): string { return `rgba(${value >> 16}, ${(value >> 8) & 255}, ${value & 255}, ${alpha})`; } +export function runtimeChartShowsSparsePoints(values: readonly (number | null | undefined)[]): boolean { + let valid = 0; + let previousValid = false; + let connected = false; + for (const value of values) { + const currentValid = value !== null && value !== undefined && Number.isFinite(value); + if (currentValid) { + valid += 1; + connected ||= previousValid; + } + previousValid = currentValid; + } + return valid > 0 && (valid <= 12 || !connected); +} + +export function runtimeChartCaption({ + title, + source, + hasLine, + allSeriesHidden, + validPoints, + sampleCount, + emptyMessage, + emptyDetail, +}: { + title: string; + source: RuntimeTrendSource; + hasLine: boolean; + allSeriesHidden: boolean; + validPoints: number; + sampleCount: number; + emptyMessage: string; + emptyDetail?: string; +}): string { + if (allSeriesHidden) return `${title} all series hidden; use the legend to show a series`; + if (hasLine) return `${title} ${source} trend available`; + if (validPoints > 0) { + return `${title} ${source} trend has ${validPoints} sparse valid point${validPoints === 1 ? "" : "s"}; a line requires consecutive buckets`; + } + if (source === "live") { + return `${title} collecting live samples; ${validPoints} of 2 valid points from ${sampleCount} snapshots`; + } + return `${title} ${emptyMessage}; ${emptyDetail ?? `${validPoints} valid points from ${sampleCount} retained buckets`}`; +} + function TrendChart({ title, subtitle, @@ -107,6 +151,7 @@ function TrendChart({ hiddenSeriesRef.current = hiddenSeries; const [viewRange, setViewRange] = useState(domain); const visibleSeries = series.filter((entry) => !hiddenSeries.has(entry.id)); + const allSeriesHidden = series.length > 0 && visibleSeries.length === 0; const hasLine = visibleSeries.some((entry) => entry.points.some((point, index) => ( point.value !== null && Number.isFinite(point.value) && index > 0 @@ -130,7 +175,7 @@ function TrendChart({ dataRef.current = chartData; formatRef.current = formatValue; maximumRef.current = maximum; - const seriesKey = series.map((entry) => `${entry.id}:${entry.label}:${entry.tone}:${entry.fill ? 1 : 0}`).join("|"); + const seriesKey = series.map((entry) => `${entry.id}:${entry.label}:${entry.tone}`).join("|"); const bandsKey = bands.map((band) => `${band.from}:${band.to}:${band.tone}`).join("|"); const ticksKey = ticks.join(":"); const selectedTimestamp = tooltip === null ? null : chartData[0][tooltip.idx]; @@ -204,10 +249,16 @@ function TrendChart({ label: entry.label, show: !hiddenSeriesRef.current.has(entry.id), stroke: toneColors[entry.tone], - fill: entry.fill ? withAlpha(toneColors[entry.tone], .13) : undefined, width: 2, spanGaps: false, - points: { show: false }, + points: { + show: (plot, seriesIndex, first, last) => runtimeChartShowsSparsePoints( + Array.from(plot.data[seriesIndex] ?? []).slice(first, last + 1), + ), + size: 6, + width: 2, + fill: surfaceColor, + }, })), ], hooks: { @@ -394,14 +445,19 @@ function TrendChart({ ))} )} - {!hasLine ?
{visibleSeries.length === 0 ? "All series hidden" : emptyMessage}{visibleSeries.length === 0 ? "Use the legend to show a series" : emptyDetail ?? `${validPoints}/2 valid points · ${samples.length} snapshots · no history is synthesized`}
: null} + {!hasLine ?
{allSeriesHidden ? "All series hidden" : validPoints > 0 ? "Sparse samples" : emptyMessage}{allSeriesHidden ? "Use the legend to show a series" : validPoints > 0 ? `${validPoints} valid point${validPoints === 1 ? "" : "s"} · a line requires consecutive buckets` : emptyDetail ?? `${validPoints}/2 valid points · ${samples.length} snapshots · no history is synthesized`}
: null} - + {series.map((entry) => { @@ -468,7 +524,7 @@ export function RuntimeTrendCharts({ return { cpu, memory: [ - { id: "used", label: "used", tone: "purple", points: memoryUsed, fill: true }, + { id: "used", label: "used", tone: "purple", points: memoryUsed }, { id: "limit", label: "configured limit", tone: "green", points: memoryLimit }, ] satisfies TrendSeries[], uptime, From ab60de2109896e615947fa92e7cc80e1979ac4ab Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 17:31:58 +0800 Subject: [PATCH 32/51] Redraw updated runtime chart data --- apps/web/e2e/agents-lifecycle.spec.ts | 28 +++++++++++++++++++ .../features/dashboard/RuntimeTrendCharts.tsx | 16 +++++++++-- 2 files changed, 42 insertions(+), 2 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 2e9ace74d..40b887e4b 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2644,6 +2644,27 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(dashboard.locator(".dashboard-runtime-sample-count")).toContainText("3 samples"); await expect(dashboard.getByText("CPU usage live trend available")).toBeAttached(); await expect(dashboard).not.toContainText("Collecting live samples"); + const memoryCard = dashboard.getByRole("region", { name: "Memory usage live chart" }); + const memoryStrokePixelCount = () => memoryCard.locator("canvas").evaluate((canvas: HTMLCanvasElement) => { + const context = canvas.getContext("2d"); + if (!context) return 0; + const { data, width, height } = context.getImageData(0, 0, canvas.width, canvas.height); + let pixels = 0; + for (let y = 0; y < height; y += 1) { + for (let x = 0; x < width; x += 1) { + const offset = (y * width + x) * 4; + const red = data[offset] ?? 0; + const green = data[offset + 1] ?? 0; + const blue = data[offset + 2] ?? 0; + const alpha = data[offset + 3] ?? 0; + if (alpha > 128 && Math.abs(red - 185) < 20 && Math.abs(green - 152) < 20 && Math.abs(blue - 244) < 20) { + pixels += 1; + } + } + } + return pixels; + }); + expect(await memoryStrokePixelCount()).toBeGreaterThan(0); const cpuChart = dashboard.getByLabel("CPU usage: 3 live samples"); const cpuCard = cpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); await cpuChart.focus(); @@ -2847,6 +2868,13 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await page.mouse.move(durablePlotBox!.x + durablePlotBox!.width * .8, durablePlotBox!.y + durablePlotBox!.height / 2, { steps: 8 }); await page.mouse.up(); await expect.poll(async () => Number(await durableCpuChart.getAttribute("data-view-start"))).toBeGreaterThan(durableInitialStart); + const durableZoomStart = Number(await durableCpuChart.getAttribute("data-view-start")); + const durableZoomEnd = Number(await durableCpuChart.getAttribute("data-view-end")); + await expect(durableCpuCard.getByRole("button", { name: "Reset zoom" })).toBeVisible(); + await page.getByRole("button", { name: "Refresh Dashboard snapshot" }).click(); + const refreshedDurableCpuChart = dashboard.getByLabel("CPU usage: 4 retained buckets"); + await expect(refreshedDurableCpuChart).toHaveAttribute("data-view-start", String(durableZoomStart)); + await expect(refreshedDurableCpuChart).toHaveAttribute("data-view-end", String(durableZoomEnd)); await page.reload(); await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index 8c216695d..7e6bdf8e6 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -336,10 +336,22 @@ function TrendChart({ useEffect(() => { const plot = plotRef.current; if (!plot) return; + const currentX = plot.scales.x; + const preserveZoom = zoomed && currentX !== undefined && + Number.isFinite(currentX.min) && Number.isFinite(currentX.max); + const xMinimum = preserveZoom ? currentX!.min! : domain.start / 1_000; + const xMaximum = preserveZoom ? currentX!.max! : domain.end / 1_000; plot.setData(chartData, false); plot.setScale("y", { min: 0, max: Math.max(1, maximum) }); - if (!zoomedRef.current) plot.setScale("x", { min: domain.start / 1_000, max: domain.end / 1_000 }); - }, [chartData, domain, maximum]); + // Reapplying x recalculates uPlot's visible data indices after setData, + // while preserving an active drag-selected range. + plot.setScale("x", { min: xMinimum, max: xMaximum }); + // setData(..., false) preserves the selected time range, but uPlot does not + // commit a draw when the explicit scales are unchanged. Memory keeps stable + // series ids across loading and loaded states, so force the refreshed paths + // onto the canvas even when its 0..limit scale remains identical. + plot.redraw(false); + }, [chartData, domain, maximum, zoomed]); useEffect(() => { const plot = plotRef.current; From 8635eadfaddc9eb79e6de92e030ff84eddd2a121 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 18:06:52 +0800 Subject: [PATCH 33/51] Fix dense runtime metric chart scaling --- apps/web/e2e/agents-lifecycle.spec.ts | 52 +++++++++++++------ .../features/dashboard/RuntimeTrendCharts.tsx | 22 ++++---- 2 files changed, 47 insertions(+), 27 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index 40b887e4b..db4678581 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2805,9 +2805,8 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa const url = new URL(route.request().url()); const start = Number(url.searchParams.get("start")); const end = Number(url.searchParams.get("end")); - const firstStart = start + 30; - const secondStart = start + 60; - const gapStart = Math.floor((start + end) / 2); + const firstStart = start; + const gapStart = end - 60; const lastStart = end - 30; const startedAt = start; const point = (pointStart: number, ratio: number, memory: number) => ({ @@ -2820,12 +2819,12 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa start: gapStart, end: gapStart + 30, first_observed_at: gapStart + 10, last_observed_at: gapStart + 20, observation_count: 1, observed_count: 0, unavailable_count: 1, cpu: null, memory: null, }; - const points = [ - point(firstStart, .25, 536_870_912), - point(secondStart, .35, 671_088_640), - gap, - point(lastStart, .5, 805_306_368), - ]; + const points = Array.from({ length: Math.floor((end - start) / 30) }, (_, index) => { + const pointStart = start + index * 30; + return pointStart === gapStart + ? gap + : point(pointStart, .25 + index / 1_000, 536_870_912 + index * 1_048_576); + }); await route.fulfill({ status: 200, contentType: "application/json", @@ -2834,7 +2833,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa requested_range: { start, end }, resolution_seconds: 30, generated_at: end, coverage: { retained_start: start, first_sample_at: firstStart + 10, last_sample_at: lastStart + 20, - sample_count: 4, expected_sample_count: Math.floor((end - start) / 30), + sample_count: points.length, expected_sample_count: Math.floor((end - start) / 30), buckets: points.map(({ cpu: _cpu, memory: _memory, ...coverage }) => coverage), }, series: [{ @@ -2851,15 +2850,38 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await expect(dashboard.getByLabel(/Durable · 30s; 1 Runtime targets/)).toBeVisible(); await expect(dashboard.getByLabel("Runtime durable-history charts")).toBeVisible(); await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); - await expect(dashboard).toContainText("4 buckets"); - await expect(dashboard).toContainText("4/120 observations"); + await expect(dashboard).toContainText("120 buckets"); + await expect(dashboard).toContainText("120/120 observations"); await expect(dashboard.getByText("Live-only metric", { exact: true })).toBeVisible(); - const durableCpuChart = dashboard.getByLabel("CPU usage: 4 retained buckets"); + const durableCpuChart = dashboard.getByLabel("CPU usage: 120 retained buckets"); await expect(dashboard.getByRole("region", { name: "CPU usage durable history chart" })).toBeVisible(); + const durableCpuCard = durableCpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); await durableCpuChart.focus(); await durableCpuChart.press("ArrowLeft"); - const durableCpuCard = durableCpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); await expect(durableCpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Unavailable"); + const durableMemoryCard = dashboard.getByRole("region", { name: "Memory usage durable history chart" }); + const durableMemorySpan = await durableMemoryCard.locator("canvas").evaluate((canvas: HTMLCanvasElement) => { + const context = canvas.getContext("2d"); + if (!context) return 0; + const { data, width, height } = context.getImageData(0, 0, canvas.width, canvas.height); + let minimumX = width; + let maximumX = -1; + for (let y = 0; y < height; y += 1) { + for (let x = 0; x < width; x += 1) { + const offset = (y * width + x) * 4; + const red = data[offset] ?? 0; + const green = data[offset + 1] ?? 0; + const blue = data[offset + 2] ?? 0; + const alpha = data[offset + 3] ?? 0; + if (alpha > 128 && Math.abs(red - 185) < 20 && Math.abs(green - 152) < 20 && Math.abs(blue - 244) < 20) { + minimumX = Math.min(minimumX, x); + maximumX = Math.max(maximumX, x); + } + } + } + return maximumX < minimumX ? 0 : maximumX - minimumX; + }); + expect(durableMemorySpan).toBeGreaterThan(100); const durableInitialStart = Number(await durableCpuChart.getAttribute("data-view-start")); const durablePlotBox = await durableCpuCard.locator(".u-over").boundingBox(); expect(durablePlotBox).not.toBeNull(); @@ -2872,7 +2894,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa const durableZoomEnd = Number(await durableCpuChart.getAttribute("data-view-end")); await expect(durableCpuCard.getByRole("button", { name: "Reset zoom" })).toBeVisible(); await page.getByRole("button", { name: "Refresh Dashboard snapshot" }).click(); - const refreshedDurableCpuChart = dashboard.getByLabel("CPU usage: 4 retained buckets"); + const refreshedDurableCpuChart = dashboard.getByLabel("CPU usage: 120 retained buckets"); await expect(refreshedDurableCpuChart).toHaveAttribute("data-view-start", String(durableZoomStart)); await expect(refreshedDurableCpuChart).toHaveAttribute("data-view-end", String(durableZoomEnd)); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index 7e6bdf8e6..a2d4034c7 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -201,7 +201,6 @@ function TrendChart({ const axisColor = color("--fg-muted", theme === "dark" ? "#a8adb8" : "#6f7480"); const gridColor = color("--line", theme === "dark" ? "#30343b" : "#e3e5e8"); const surfaceColor = color("--surface", theme === "dark" ? "#17191d" : "#ffffff"); - const yMaximum = Math.max(1, maximum); const bandColors: Record = { safe: withAlpha("#50d5a0", .07), warning: withAlpha("#f59e52", .07), @@ -214,7 +213,10 @@ function TrendChart({ legend: { show: false }, scales: { x: { time: true }, - y: { auto: false, range: [0, yMaximum] }, + // Stable series can create the plot before their first data response. + // Read the current maximum whenever uPlot ranges the y scale so the + // initial 0..1 fallback does not survive a later data update. + y: { range: () => [0, Math.max(1, maximumRef.current)] }, }, axes: [ { @@ -341,16 +343,12 @@ function TrendChart({ Number.isFinite(currentX.min) && Number.isFinite(currentX.max); const xMinimum = preserveZoom ? currentX!.min! : domain.start / 1_000; const xMaximum = preserveZoom ? currentX!.max! : domain.end / 1_000; - plot.setData(chartData, false); - plot.setScale("y", { min: 0, max: Math.max(1, maximum) }); - // Reapplying x recalculates uPlot's visible data indices after setData, - // while preserving an active drag-selected range. - plot.setScale("x", { min: xMinimum, max: xMaximum }); - // setData(..., false) preserves the selected time range, but uPlot does not - // commit a draw when the explicit scales are unchanged. Memory keeps stable - // series ids across loading and loaded states, so force the refreshed paths - // onto the canvas even when its 0..limit scale remains identical. - plot.redraw(false); + plot.batch(() => { + plot.setData(chartData, false); + // Reapplying x recalculates visible indices and the auto y scale after + // setData, while preserving an active drag-selected range. + plot.setScale("x", { min: xMinimum, max: xMaximum }); + }); }, [chartData, domain, maximum, zoomed]); useEffect(() => { From 5a26e1df768bbdebebcc5b52862832eaefbee902 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 18:11:34 +0800 Subject: [PATCH 34/51] Keep runtime history continuous by allocation --- .../dashboard/RuntimeTrendCharts.test.tsx | 2 +- .../features/dashboard/RuntimeTrendCharts.tsx | 10 ++-- .../src/features/dashboard/runtime-history.ts | 15 ++---- .../features/dashboard/runtime-trends.test.ts | 36 +++++++++---- .../src/features/dashboard/runtime-trends.ts | 52 ++++++++++--------- contracts/agents-api/runtime-history-api.md | 21 ++++---- .../runtime-observability-design.md | 32 ++++++------ contracts/agents-api/v1/runtime_history.go | 4 +- .../src/runtime-history-projection.ts | 2 +- .../agents-client/src/runtime-history.test.ts | 15 ++---- packages/agents-client/src/types.ts | 2 +- .../internal/api/runtime_history_test.go | 5 +- .../clickhousereader/acceptance_test.go | 32 ++++++------ .../runtimehistory/clickhousereader/reader.go | 28 ++++------ .../clickhousereader/reader_test.go | 31 +++++++++++ .../runtimehistory/storeresolver/resolver.go | 4 +- .../internal/runtimehistory/types.go | 7 +-- 17 files changed, 164 insertions(+), 134 deletions(-) diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index c63382c37..09ce5d4c0 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -8,7 +8,7 @@ function sample(sampledAt: number, cpuRatio: number | null): RuntimeTrendSample return { sampledAt, targets: [{ - sessionId: "session-1", + seriesId: "session-1:allocation-1", label: "Runtime worker", cpuRatio, uptimeSeconds: 120, diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index a2d4034c7..02186cd9d 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -486,7 +486,7 @@ function targetIds(samples: readonly RuntimeTrendSample[], field: "cpuRatio" | " for (const sample of samples) { for (const target of sample.targets) { const value = target[field]; - if (value !== null) latest.set(target.sessionId, value); + if (value !== null) latest.set(target.seriesId, value); } } return [...latest.entries()].sort((left, right) => right[1] - left[1]).slice(0, 3).map(([id]) => id); @@ -494,7 +494,7 @@ function targetIds(samples: readonly RuntimeTrendSample[], field: "cpuRatio" | " function targetLabel(samples: readonly RuntimeTrendSample[], id: string): string { for (let index = samples.length - 1; index >= 0; index -= 1) { - const target = samples[index]?.targets.find((candidate) => candidate.sessionId === id); + const target = samples[index]?.targets.find((candidate) => candidate.seriesId === id); if (target) return target.label; } return "Runtime"; @@ -520,7 +520,7 @@ export function RuntimeTrendCharts({ id, label: targetLabel(samples, id), tone: tones[index] ?? "blue", - points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.sessionId === id)?.cpuRatio ?? null })).map((point) => ({ ...point, value: point.value === null ? null : point.value * 100 })), + points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.seriesId === id)?.cpuRatio ?? null })).map((point) => ({ ...point, value: point.value === null ? null : point.value * 100 })), })); const memoryUsed = samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.memoryUsageBytes })); const memoryLimit = samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.memoryLimitBytes })); @@ -528,7 +528,7 @@ export function RuntimeTrendCharts({ id, label: targetLabel(samples, id), tone: tones[(index + 2) % tones.length] ?? "blue", - points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.sessionId === id)?.uptimeSeconds ?? null })), + points: samples.map((sample) => ({ sampledAt: sample.sampledAt, value: sample.targets.find((target) => target.seriesId === id)?.uptimeSeconds ?? null })), })); const throughput = tokenThroughput(samples); return { @@ -556,7 +556,7 @@ export function RuntimeTrendCharts({
`${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> - formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> + formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} />
); diff --git a/apps/web/src/features/dashboard/runtime-history.ts b/apps/web/src/features/dashboard/runtime-history.ts index 4eab4def2..dc30421f4 100644 --- a/apps/web/src/features/dashboard/runtime-history.ts +++ b/apps/web/src/features/dashboard/runtime-history.ts @@ -84,15 +84,6 @@ interface MutableBucket { memory: Map; } -function incarnationLabel(title: string, history: RuntimeHistory, startedAt: number): string { - const incarnationCount = new Set(history.series.map((series) => ( - `${series.allocation_id}:${series.started_at.seconds}:${series.started_at.nanoseconds}` - ))).size; - if (incarnationCount <= 1) return title; - const time = new Date(startedAt * 1_000).toLocaleTimeString([], { hour: "2-digit", minute: "2-digit" }); - return `${title} · ${time}`; -} - export function runtimeDurableTrendSamples( sessions: readonly AgentSession[], histories: readonly RuntimeHistory[], @@ -112,14 +103,14 @@ export function runtimeDurableTrendSamples( for (const coverage of history.coverage.buckets) bucket(coverage.end * 1_000); for (const series of history.series) { const startedAt = series.started_at.seconds + series.started_at.nanoseconds / 1_000_000_000; - const targetID = `${history.session_id}:${series.allocation_id}:${series.started_at.seconds}:${series.started_at.nanoseconds}`; - const label = incarnationLabel(titles.get(history.session_id) ?? "Runtime", history, series.started_at.seconds); + const targetID = `${history.session_id}:${series.allocation_id}`; + const label = titles.get(history.session_id) ?? "Runtime"; for (const point of series.points) { const value = bucket(point.end * 1_000); const observedAt = point.last_observed_at; const uptime = observedAt === null || observedAt < startedAt ? null : observedAt - startedAt; value.targets.set(targetID, { - sessionId: targetID, + seriesId: targetID, label, cpuRatio: point.cpu?.utilization_ratio ?? null, uptimeSeconds: uptime, diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts index fa33ec427..683da9bc7 100644 --- a/apps/web/src/features/dashboard/runtime-trends.test.ts +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -140,7 +140,7 @@ describe("Runtime live-window trends", () => { ]); }); - it("derives real CPU utilization from cumulative samples within one Runtime incarnation", () => { + it("derives real CPU utilization from cumulative samples within one allocation", () => { const cumulative = (at: number, usage: number) => snapshot(at, { cpuRatio: null, cpuUsageCores: null, @@ -154,12 +154,27 @@ describe("Runtime live-window trends", () => { expect(samples.map((sample) => sample.targets[0]?.cpuRatio ?? null)).toEqual([null, .5, .5]); }); - it("does not derive CPU across restarts, allocation changes, or counter regressions", () => { + it("uses allocation identity for Live chart series while ignoring start-time jitter", () => { + const first = runtimeTrendSample(snapshot(60_000, { cpuRatio: .25, startedAt: 0 })); + const jittered = runtimeTrendSample(snapshot(120_000, { cpuRatio: .5, startedAt: 1 })); + const replaced = runtimeTrendSample(snapshot(180_000, { + cpuRatio: .5, + startedAt: 1, + allocationId: "44444444-4444-4444-8444-444444444444", + })); + expect(first.targets[0]?.seriesId).toBe(jittered.targets[0]?.seriesId); + expect(replaced.targets[0]?.seriesId).not.toBe(first.targets[0]?.seriesId); + }); + + it("keeps one allocation across start changes but resets CPU on allocation changes or counter regressions", () => { const base = snapshot(60_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 100, startedAt: 0, }); + const continued = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), snapshot(120_000, { + cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 1, + })); + expect(continued.at(-1)?.targets[0]?.cpuRatio ?? null).toBe(.5); for (const next of [ - snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 1 }), snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 0, allocationId: "44444444-4444-4444-8444-444444444444" }), snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 10, startedAt: 0 }), ]) { @@ -168,16 +183,15 @@ describe("Runtime live-window trends", () => { } }); - it("does not connect directly reported CPU across Runtime incarnation fences", () => { + it("keeps directly reported CPU continuous across start changes but not stale observations", () => { const base = snapshot(60_000, { cpuRatio: .25, startedAt: 0 }); - for (const next of [ - snapshot(120_000, { cpuRatio: .5, startedAt: 1 }), - snapshot(120_000, { cpuRatio: .5, startedAt: 0, allocationId: "44444444-4444-4444-8444-444444444444" }), + const continued = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), snapshot(120_000, { cpuRatio: .5, startedAt: 1 })); + expect(continued.at(-1)?.targets[0]?.cpuRatio ?? null).toBe(.5); + const stale = appendRuntimeTrendSample( + appendRuntimeTrendSample([], base), snapshot(120_000, { cpuRatio: .5, startedAt: 0, observedAt: 60 }), - ]) { - const samples = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), next); - expect(samples.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); - } + ); + expect(stale.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); }); it("rejects non-finite CPU ratios produced by finite provider inputs", () => { diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts index 76667bf4f..c99ae9c89 100644 --- a/apps/web/src/features/dashboard/runtime-trends.ts +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -13,7 +13,7 @@ export const RUNTIME_TREND_RANGES = [ export type RuntimeTrendRange = typeof RUNTIME_TREND_RANGES[number]["milliseconds"]; export interface RuntimeTrendTarget { - sessionId: string; + seriesId: string; label: string; cpuRatio: number | null; uptimeSeconds: number | null; @@ -21,7 +21,7 @@ export interface RuntimeTrendTarget { export interface RuntimeTrendCPUCandidate extends RuntimeTrendTarget { observedAt: number | null; - incarnationKey: string | null; + allocationKey: string | null; usageSecondsTotal: number | null; capacityCores: number | null; reportedRatio: number | null; @@ -75,10 +75,12 @@ function reportedCpuRatio(observation: RuntimeObservation): number | null { : null; } -function incarnationKey(observation: RuntimeObservation): string | null { +function allocationKey(observation: RuntimeObservation): string | null { if (observation.status !== "observed") return null; - const startedAt = safeInteger(observation.started_at); - return startedAt === null ? null : `${observation.instance.kind}:${observation.instance.allocation_id}:${startedAt}`; + const allocationId = observation.instance.allocation_id; + return typeof allocationId === "string" && allocationId.length > 0 + ? `${observation.instance.kind}:${allocationId}` + : null; } function uptimeSeconds(observation: RuntimeObservation): number | null { @@ -105,12 +107,14 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT const observed = snapshot.observations.flatMap((observation) => { const session = sessions.get(observation.session_id); if (!session || observation.status !== "observed") return []; + const key = allocationKey(observation); + if (key === null) return []; return [{ - sessionId: observation.session_id, + seriesId: `${observation.session_id}:${key}`, label: sessionTitle(session), cpuRatio: reportedCpuRatio(observation), observedAt: safeInteger(observation.observed_at), - incarnationKey: incarnationKey(observation), + allocationKey: allocationKey(observation), usageSecondsTotal: finiteNonNegative(observation.cpu?.usage_seconds_total), capacityCores: finiteNonNegative(observation.cpu?.capacity_cores), memoryUsageBytes: finiteNonNegative(observation.memory?.usage_bytes), @@ -122,14 +126,14 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT ...observed.filter((target) => target.cpuRatio !== null) .sort((left, right) => (right.cpuRatio ?? 0) - (left.cpuRatio ?? 0)) .slice(0, RUNTIME_TREND_SERIES_LIMIT) - .map((target) => target.sessionId), + .map((target) => target.seriesId), ...observed.filter((target) => target.uptimeSeconds !== null) .sort((left, right) => (right.uptimeSeconds ?? 0) - (left.uptimeSeconds ?? 0)) .slice(0, RUNTIME_TREND_SERIES_LIMIT) - .map((target) => target.sessionId), + .map((target) => target.seriesId), ]); - const targets = observed.filter((target) => targetIds.has(target.sessionId)).map((target) => ({ - sessionId: target.sessionId, + const targets = observed.filter((target) => targetIds.has(target.seriesId)).map((target) => ({ + seriesId: target.seriesId, label: target.label, cpuRatio: target.cpuRatio, uptimeSeconds: target.uptimeSeconds, @@ -142,16 +146,16 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT targets, cpuCandidates: observed.flatMap((target): RuntimeTrendCPUCandidate[] => ( target.cpuRatio !== null || ( - target.observedAt !== null && target.incarnationKey !== null && + target.observedAt !== null && target.allocationKey !== null && target.usageSecondsTotal !== null && target.capacityCores !== null && target.capacityCores > 0 ) ? [{ - sessionId: target.sessionId, + seriesId: target.seriesId, label: target.label, cpuRatio: target.cpuRatio, uptimeSeconds: target.uptimeSeconds, observedAt: target.observedAt, - incarnationKey: target.incarnationKey, + allocationKey: target.allocationKey, usageSecondsTotal: target.usageSecondsTotal, capacityCores: target.capacityCores, reportedRatio: target.cpuRatio, @@ -171,20 +175,20 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT } function cpuRatios(previous: RuntimeTrendSample, next: RuntimeTrendSample): Map { - const previousCandidates = new Map(previous.cpuCandidates.map((candidate) => [candidate.sessionId, candidate])); + const previousCandidates = new Map(previous.cpuCandidates.map((candidate) => [candidate.seriesId, candidate])); const ratios = new Map(); for (const current of next.cpuCandidates) { - const prior = previousCandidates.get(current.sessionId); + const prior = previousCandidates.get(current.seriesId); if (current.reportedRatio !== null && !prior) { - ratios.set(current.sessionId, current.reportedRatio); + ratios.set(current.seriesId, current.reportedRatio); continue; } if ( - !prior || prior.incarnationKey === null || prior.incarnationKey !== current.incarnationKey || + !prior || prior.allocationKey === null || prior.allocationKey !== current.allocationKey || prior.observedAt === null || current.observedAt === null || current.observedAt <= prior.observedAt ) continue; if (current.reportedRatio !== null) { - ratios.set(current.sessionId, current.reportedRatio); + ratios.set(current.seriesId, current.reportedRatio); continue; } if ( @@ -195,7 +199,7 @@ function cpuRatios(previous: RuntimeTrendSample, next: RuntimeTrendSample): Map< const usageCores = (current.usageSecondsTotal - prior.usageSecondsTotal) / (current.observedAt - prior.observedAt); const ratio = finiteNonNegative(usageCores / current.capacityCores); - if (ratio !== null) ratios.set(current.sessionId, ratio); + if (ratio !== null) ratios.set(current.seriesId, ratio); } return ratios; } @@ -206,9 +210,9 @@ function applyCPURatios(sample: RuntimeTrendSample, ratios: ReadonlyMap ({ ...target, cpuRatio: null })); const cpu = sample.cpuCandidates.flatMap((candidate): RuntimeTrendTarget[] => { - const ratio = ratios.get(candidate.sessionId); + const ratio = ratios.get(candidate.seriesId); return ratio === undefined ? [] : [{ - sessionId: candidate.sessionId, + seriesId: candidate.seriesId, label: candidate.label, cpuRatio: ratio, uptimeSeconds: candidate.uptimeSeconds, @@ -216,9 +220,9 @@ function applyCPURatios(sample: RuntimeTrendSample, ratios: ReadonlyMap (right.cpuRatio ?? 0) - (left.cpuRatio ?? 0)) .slice(0, RUNTIME_TREND_SERIES_LIMIT); const selected = new Map( - uptime.map((target) => [target.sessionId, target]), + uptime.map((target) => [target.seriesId, target]), ); - for (const target of cpu) selected.set(target.sessionId, target); + for (const target of cpu) selected.set(target.seriesId, target); sample.targets = [...selected.values()]; } diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md index f795ae8ea..e5b79aa78 100644 --- a/contracts/agents-api/runtime-history-api.md +++ b/contracts/agents-api/runtime-history-api.md @@ -83,13 +83,16 @@ query inputs. ``` `coverage` describes all resolved samples, including unavailable observations that -cannot safely be attached to one compute incarnation. `retained_start` is the +cannot safely be attached to one allocation series. `retained_start` is the latest of the requested start, configured retention boundary, and backend-reported retention boundary. Expected coverage is calculated only over that retained period and only from qualified periodic cadence. -Resource `series` are split by `allocation_id` plus the lossless compute -`started_at` fence: +Resource `series` are keyed only by `allocation_id`. This keeps one continuous +Dashboard lifecycle when a provider pauses, restores, restarts, or replaces its +underlying compute without replacing the durable allocation. `started_at` remains +the earliest retained provider start estimate for compatible uptime display; it is +not series identity: ```json { @@ -101,11 +104,11 @@ Resource `series` are split by `allocation_id` plus the lossless compute ``` The two integers are JSON-safe, nonnegative Unix seconds and a 0–999,999,999 -nanosecond remainder. They are identity, not merely a display timestamp. CPU -utilization is derived only from ordered cumulative counters inside that fence; -successive counter intervals are assigned to the bucket containing their right -endpoint and combined by CPU-capacity time. Memory values are the last observed -values in a bucket. Every point contains +nanosecond remainder. CPU utilization is derived from ordered cumulative counters +inside the allocation. A counter regression resets the baseline, so no interval +is derived across a compute replacement. Successive valid counter intervals are +assigned to the bucket containing their right endpoint and combined by CPU-capacity +time. Memory values are the last observed values in a bucket. Every point contains observation and contributor counts. Missing values are null and gaps remain gaps. Numeric zero is retained as an observed value. The endpoint does not return token history: token throughput remains sourced from canonical Session Usage and is @@ -149,7 +152,7 @@ interface AgentCore { ``` The client validates exact fields, capability consistency, requested-range echo, -Session identity, half-open bucket ordering, coverage totals, incarnation identity, +Session identity, half-open bucket ordering, coverage totals, allocation identity, contributor counts, nullability, finite numbers, and response size. Unknown fields or malformed data reject the entire response with a 502 client projection error. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 3c5287132..3e792aae8 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -168,14 +168,14 @@ CPU percentage is derived from the delta between two cumulative samples and thei observation times. A single sample cannot truthfully supply CPU percentage. The API projection may additionally expose `usage_cores` and `utilization_ratio` -only when the service has two ordered samples for the same Runtime incarnation. +only when the service has two ordered samples for the same Runtime allocation. A future process-local observation cache may keep the previous cumulative value for this calculation. Phase 2 intentionally leaves both derived fields null because it has only one provider sample per request. The browser-local live window therefore derives interval utilization from adjacent cumulative samples only when -Session, allocation, compute `started_at`, and provider observation order still -match. Restart, replacement, counter regression, missing capacity, or cache loss -creates a gap; none changes the cumulative source measurement or lifecycle state. +Session, allocation, and provider observation order still match. Counter regression, +missing capacity, or cache loss creates a gap; none changes the cumulative source +measurement or lifecycle state. ## 8. Duration semantics @@ -330,9 +330,9 @@ Core-owned tenant, Session, Environment, allocation, mode, provider type, status, safe reason, collection source (`on_read` or `periodic`), and Core resolved/observed timestamps are metric attributes. The explicit nanosecond timestamps preserve the record join key when a backend's generic OTLP tables -store metric event time at lower precision. CPU, capacity, and memory points -are exported only when the sample also carries the compute `started_at` fence; -that fence is included as an attribute on every such point. Provider keys, +store metric event time at lower precision. CPU, capacity, and memory points are +exported only when the sample also carries a provider start estimate; it is +included for compatible display but is not series identity. Provider keys, provider receipts, native container/pod/instance identifiers, raw errors, paths, and credentials are not attributes. Missing measurements produce no value point; they are represented only by the explicit @@ -373,8 +373,9 @@ before calling a Reader. Reader queries always carry tenant, Session, and Environment scope plus a bounded start, exclusive end, server-selected step, and total point budget. Provider-native identity is never a query input. -Reader results remain divided by allocation and the lossless compute `started_at` -seconds-plus-nanoseconds fence. +Reader results remain divided by allocation. Provider `started_at` values are +retained only as compatible display metadata and never split one durable allocation +into multiple Dashboard series. Every bucket reports explicit observation coverage and nullable CPU/memory values. CPU utilization may be derived only from ordered cumulative counters inside one fence; successive intervals are assigned to the bucket containing @@ -394,7 +395,8 @@ This internal boundary, the Session-scoped public extension, capability discover strict client, and production ClickHouse reference Reader are implemented. The Reader queries only the specialized projection, always includes tenant, Session, Environment, bounded time, and `collection_source = 'periodic'` predicates, and -aggregates resource points by allocation plus lossless incarnation fence. Durable +aggregates resource points by allocation, resetting CPU derivation after a +cumulative-counter regression. Durable Web ranges remain gated on real retention, isolation, restart, and incarnation acceptance. @@ -551,13 +553,13 @@ Implemented for the browser-local current-snapshot live window. observation results. It is disabled by default, drops on queue saturation, and cannot fail the current-observation request path. - Implemented: optional server-only OTLP/HTTP protobuf transport for the six - documented Runtime instruments, including allocation and compute-incarnation - fencing attributes. Configuration is strict and secrets never reach Web. + documented Runtime instruments, including allocation identity and provider + start metadata. Configuration is strict and secrets never reach Web. - Implemented: optional execution-owner singleton sampling across all nondeleted managed Sessions. Keyset scans, provider concurrency, source deadlines, and non-overlapping sweeps are bounded; collection source is exported explicitly. - Implemented: backend-neutral `runtimehistory` types and service validation. - Tenant/Session/Environment scope precedes every Reader query; incarnation, + Tenant/Session/Environment scope precedes every Reader query; allocation, coverage, nullability, ordering, range and total-point invariants are enforced. - Implemented: safe public capability discovery, bounded Session-scoped history query routes, generated OpenAPI schemas, and strict `packages/agents-client` @@ -567,8 +569,8 @@ Implemented for the browser-local current-snapshot live window. tenant/Session/Environment/periodic-source query predicates. No backend is a Core execution dependency. - Qualified: real OTLP Collector-to-ClickHouse acceptance covers tenant isolation, - periodic-only public reads, Reader restart persistence, and multiple Runtime - incarnations without cross-fence CPU derivation. + periodic-only public reads, Reader restart persistence, continuous allocation + series, and CPU baseline reset after cumulative-counter regression. - Implemented: Core Web discovers capabilities, reloads bounded Session histories with bounded concurrency, and exposes explicit Live versus History sources with 1h, 6h, and 24h Durable ranges. Token throughput remains honestly Live-only. diff --git a/contracts/agents-api/v1/runtime_history.go b/contracts/agents-api/v1/runtime_history.go index 741c1ba04..c8f11ffeb 100644 --- a/contracts/agents-api/v1/runtime_history.go +++ b/contracts/agents-api/v1/runtime_history.go @@ -56,8 +56,8 @@ type RuntimeHistorySeries struct { Points []RuntimeHistoryPoint `json:"points" binding:"required" validate:"max=10000"` } -// RuntimeHistoryTime is a lossless JSON-safe timestamp used as an incarnation -// identity fence. Other public history timestamps are display/query seconds. +// RuntimeHistoryTime is a lossless JSON-safe provider start estimate retained +// for compatible uptime display. AllocationID is the series identity. type RuntimeHistoryTime struct { Seconds int64 `json:"seconds" binding:"required" minimum:"0" maximum:"9007199254740991"` Nanoseconds int `json:"nanoseconds" binding:"required" minimum:"0" maximum:"999999999"` diff --git a/packages/agents-client/src/runtime-history-projection.ts b/packages/agents-client/src/runtime-history-projection.ts index 38c79df43..9fd253ad5 100644 --- a/packages/agents-client/src/runtime-history-projection.ts +++ b/packages/agents-client/src/runtime-history-projection.ts @@ -273,7 +273,7 @@ export function projectRuntimeHistory( )); if ( new Set(series.map((entry) => entry.environment_id)).size > 1 || - new Set(series.map((entry) => `${entry.allocation_id}\u0000${entry.started_at.seconds}\u0000${entry.started_at.nanoseconds}`)).size !== series.length || + new Set(series.map((entry) => entry.allocation_id)).size !== series.length || buckets.length + series.reduce((sum, entry) => sum + entry.points.length, 0) > maximumTotalPoints || series.some((entry) => entry.points.some((point) => point.first_observed_at !== null && point.first_observed_at < retainedStart)) || series.some((entry) => entry.points.some((point) => point.last_observed_at !== null && point.last_observed_at > generatedAt)) diff --git a/packages/agents-client/src/runtime-history.test.ts b/packages/agents-client/src/runtime-history.test.ts index 948933d50..982966143 100644 --- a/packages/agents-client/src/runtime-history.test.ts +++ b/packages/agents-client/src/runtime-history.test.ts @@ -101,10 +101,6 @@ describe("Runtime history client", () => { const calls: FetchCall[] = []; const controller = new AbortController(); const response = history(); - const incarnations = response.series as Array>; - const secondIncarnation = structuredClone(incarnations[0]!); - secondIncarnation.started_at = { seconds: 900, nanoseconds: 1 }; - incarnations.push(secondIncarnation); const value = await clientFor(response, calls).retrieveRuntimeHistory(sessionId.toUpperCase(), { start: 1000, end: 1120, @@ -117,10 +113,7 @@ describe("Runtime history client", () => { expect(calls[0]?.init?.signal).toBe(controller.signal); expect(value.series[0]?.points[0]?.cpu?.utilization_ratio).toBe(0); expect(value.series[0]?.points[0]?.memory?.usage_bytes).toBe(0); - expect(value.series.map((series) => series.started_at)).toEqual([ - { seconds: 900, nanoseconds: 0 }, - { seconds: 900, nanoseconds: 1 }, - ]); + expect(value.series.map((series) => series.started_at)).toEqual([{ seconds: 900, nanoseconds: 0 }]); expect(value.coverage.sample_count).toBe(2); }); @@ -223,9 +216,11 @@ describe("Runtime history client", () => { const point = (value.series as Array<{ points: Array> }>)[0]!.points[0]!; point.memory = { contributor_count: 1, usage_bytes: Number.MAX_SAFE_INTEGER + 1, limit_bytes: 2048 }; }], - ["duplicate incarnation", (value: Record) => { + ["duplicate allocation with another start estimate", (value: Record) => { const series = value.series as Array>; - series.push(structuredClone(series[0]!)); + const duplicate = structuredClone(series[0]!); + duplicate.started_at = { seconds: 901, nanoseconds: 0 }; + series.push(duplicate); }], ["mixed Environment scope", (value: Record) => { const series = value.series as Array>; diff --git a/packages/agents-client/src/types.ts b/packages/agents-client/src/types.ts index 44d2f2598..0a0ea3a2d 100644 --- a/packages/agents-client/src/types.ts +++ b/packages/agents-client/src/types.ts @@ -890,7 +890,7 @@ export interface RuntimeHistoryPoint extends RuntimeHistoryCoveragePoint { export interface RuntimeHistorySeries { environment_id: string; allocation_id: string; - /** Lossless compute-incarnation identity fence. */ + /** Earliest retained provider start estimate for compatible uptime display. */ started_at: RuntimeHistoryTime; provider_type: string; points: RuntimeHistoryPoint[]; diff --git a/services/agents-api/internal/api/runtime_history_test.go b/services/agents-api/internal/api/runtime_history_test.go index 3ff2532eb..b9705985f 100644 --- a/services/agents-api/internal/api/runtime_history_test.go +++ b/services/agents-api/internal/api/runtime_history_test.go @@ -116,15 +116,12 @@ func TestRuntimeHistoryRouteBindsAuthenticatedSessionAndPreservesCoverage(t *tes Coverage: []runtimehistory.CoveragePoint{coveragePoint}, Series: []runtimehistory.Series{{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker", Points: []runtimehistory.Point{resourcePoint}}}, } - secondIncarnation := service.response.Series[0] - secondIncarnation.StartedAt = secondIncarnation.StartedAt.Add(time.Nanosecond) - service.response.Series = append(service.response.Series, secondIncarnation) response := runtimeObservationRequest(handler, "/v1/agents/sessions/"+sessionID+"/runtime-history?start="+timeString(start)+"&end="+timeString(now)+"&max_points=60") if response.Code != http.StatusOK { t.Fatalf("history returned %d: %s", response.Code, response.Body) } var value v1.RuntimeHistory - if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history" || value.Source != "durable" || value.SessionID != sessionID || value.ResolutionSeconds != 60 || value.Coverage.SampleCount != 1 || value.Coverage.ExpectedSampleCount != 120 || len(value.Coverage.Buckets) != 1 || len(value.Series) != 2 || len(value.Series[0].Points) != 1 || value.Series[0].StartedAt.Seconds != value.Series[1].StartedAt.Seconds || value.Series[0].StartedAt.Nanoseconds == value.Series[1].StartedAt.Nanoseconds { + if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history" || value.Source != "durable" || value.SessionID != sessionID || value.ResolutionSeconds != 60 || value.Coverage.SampleCount != 1 || value.Coverage.ExpectedSampleCount != 120 || len(value.Coverage.Buckets) != 1 || len(value.Series) != 1 || len(value.Series[0].Points) != 1 { t.Fatalf("invalid history response: %s", response.Body) } point := value.Series[0].Points[0] diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go b/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go index c1c6f141a..b4d38e831 100644 --- a/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go +++ b/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go @@ -87,7 +87,7 @@ func TestClickHouseCollectorAcceptance(t *testing.T) { deadline := time.Now().Add(20 * time.Second) for { result, err = reader.Query(t.Context(), query) - if err == nil && coverageCount(result.Coverage) == 5 && len(result.Series) == 2 { + if err == nil && coverageCount(result.Coverage) == 5 && len(result.Series) == 1 { break } if time.Now().After(deadline) { @@ -98,7 +98,7 @@ func TestClickHouseCollectorAcceptance(t *testing.T) { if err := reader.Close(); err != nil { t.Fatal(err) } - assertAcceptanceResult(t, result, firstStartedAt, secondStartedAt) + assertAcceptanceResult(t, result, firstStartedAt) // A fresh Reader proves history is backend-owned, not process memory. reader = open() @@ -107,7 +107,7 @@ func TestClickHouseCollectorAcceptance(t *testing.T) { if err != nil { t.Fatal(err) } - if coverageCount(reloaded.Coverage) != 5 || len(reloaded.Series) != 2 { + if coverageCount(reloaded.Coverage) != 5 || len(reloaded.Series) != 1 { t.Fatalf("history did not survive Reader restart: %+v", reloaded) } } @@ -134,28 +134,26 @@ func coverageCount(values []runtimehistory.CoveragePoint) int { return total } -func assertAcceptanceResult(t *testing.T, result runtimehistory.Result, firstStartedAt, secondStartedAt time.Time) { +func assertAcceptanceResult(t *testing.T, result runtimehistory.Result, firstStartedAt time.Time) { t.Helper() if coverageCount(result.Coverage) != 5 { t.Fatalf("cross-tenant sample leaked into coverage: %+v", result.Coverage) } - if len(result.Series) != 2 || !result.Series[0].StartedAt.Equal(firstStartedAt) || !result.Series[1].StartedAt.Equal(secondStartedAt) { - t.Fatalf("incarnations were not kept separate and lossless: %+v", result.Series) + if len(result.Series) != 1 || !result.Series[0].StartedAt.Equal(firstStartedAt) { + t.Fatalf("allocation history was split by compute start estimates: %+v", result.Series) } wantRatios := []float64{.5, 1.0 / 6.0} - for index, series := range result.Series { - found := false - for _, point := range series.Points { - if point.CPUUtilizationRatio != nil { - if math.Abs(*point.CPUUtilizationRatio-wantRatios[index]) > 1e-9 { - t.Fatalf("unexpected CPU ratio for incarnation %d: %+v", index, point) - } - found = true + found := 0 + for _, point := range result.Series[0].Points { + if point.CPUUtilizationRatio != nil { + if found >= len(wantRatios) || math.Abs(*point.CPUUtilizationRatio-wantRatios[found]) > 1e-9 { + t.Fatalf("unexpected CPU ratio after allocation counter reset: %+v", point) } + found++ } - if !found { - t.Fatalf("incarnation %d has no derived CPU point: %+v", index, series) - } + } + if found != len(wantRatios) { + t.Fatalf("allocation lost CPU segments across counter reset: %+v", result.Series[0]) } } diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader.go index 367720497..fc9a6084d 100644 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go +++ b/services/agents-api/internal/runtimehistory/clickhousereader/reader.go @@ -54,7 +54,7 @@ GROUP BY provider_type, status, metric_name -ORDER BY resolved_at_unix_nano, allocation_id, started_at_unix_nano, metric_name +ORDER BY resolved_at_unix_nano, allocation_id, metric_name LIMIT {row_limit:UInt64}` type Config struct { @@ -279,17 +279,19 @@ func readSamples(resultRows rows, rawStart, end time.Time) ([]*rawSample, int, e if err := validateProjectedRow(parsedStatus, metric, minimum, observedAt, startedAt); err != nil { return nil, rowCount, err } - startedKey := "" - if startedAt != nil { - startedKey = fmt.Sprintf("%d", startedNano.value) - } - key := fmt.Sprintf("%d\x00%s\x00%s", resolvedNano, allocation, startedKey) + // Allocation identity is the durable Dashboard series boundary. Provider + // start-time estimates are sample metadata and may jitter between reads. + key := fmt.Sprintf("%d\x00%s", resolvedNano, allocation) sample := byKey[key] if sample == nil { sample = &rawSample{resolvedAt: resolvedAt, observedAt: observedAt, allocation: allocation, startedAt: startedAt, provider: provider, status: parsedStatus, metrics: map[string]float64{}} byKey[key] = sample } else if sample.provider != provider || sample.status != parsedStatus || !sameOptionalTime(sample.observedAt, observedAt) { return nil, rowCount, errors.New("conflicting Runtime history sample identity") + } else if startedAt != nil && (sample.startedAt == nil || startedAt.Before(*sample.startedAt)) { + // Retain the earliest estimate for compatible display metadata without + // allowing it to split one allocation into multiple series. + sample.startedAt = startedAt } if previous, ok := sample.metrics[metric]; ok && previous != minimum { return nil, rowCount, errors.New("conflicting Runtime history metric value") @@ -311,9 +313,6 @@ func readSamples(resultRows rows, rawStart, end time.Time) ([]*rawSample, int, e } sort.Slice(result, func(i, j int) bool { if result[i].resolvedAt.Equal(result[j].resolvedAt) { - if result[i].allocation == result[j].allocation { - return timeValue(result[i].startedAt) < timeValue(result[j].startedAt) - } return result[i].allocation < result[j].allocation } return result[i].resolvedAt.Before(result[j].resolvedAt) @@ -344,13 +343,6 @@ func sameOptionalTime(left, right *time.Time) bool { return left == nil && right == nil || left != nil && right != nil && left.Equal(*right) } -func timeValue(value *time.Time) int64 { - if value == nil { - return -1 - } - return value.UnixNano() -} - type coverageAggregate struct { first, last time.Time observations int @@ -397,13 +389,15 @@ func aggregate(query runtimehistory.Query, generatedAt time.Time, raw []*rawSamp if allocationErr != nil || parsedAllocation == uuid.Nil || parsedAllocation.String() != sample.allocation { return runtimehistory.Result{}, errors.New("invalid Runtime history allocation") } - key := sample.allocation + "\x00" + sample.startedAt.Format(time.RFC3339Nano) + key := sample.allocation value := series[key] if value == nil { value = &seriesAggregate{allocation: sample.allocation, startedAt: *sample.startedAt, provider: sample.provider} series[key] = value } else if value.provider != sample.provider { return runtimehistory.Result{}, errors.New("conflicting Runtime history provider") + } else if sample.startedAt.Before(value.startedAt) { + value.startedAt = *sample.startedAt } value.samples = append(value.samples, sample) } diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go index 69abcdbd0..3721a59d6 100644 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go +++ b/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go @@ -156,6 +156,37 @@ func TestReaderQueriesMandatoryScopeAndAggregatesFencedHistory(t *testing.T) { } } +func TestReaderKeepsStartedAtJitterInOneAllocationSeries(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + firstStarted := start.Add(-time.Minute).UnixNano() + secondStarted := start.Add(-time.Minute + 1500*time.Millisecond).UnixNano() + baseline := start.Add(-10 * time.Second).UnixNano() + observed := start.Add(10 * time.Second).UnixNano() + client := &fakeClient{rows: &fakeRows{values: [][]any{ + metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.SampleName, 1), + metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.CPUUsageName, 1), + metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.CPUCapacityName, 2), + metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.SampleName, 1), + metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.CPUUsageName, 3), + metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.CPUCapacityName, 2), + }}} + reader, err := newReader(client, testCapabilities(), time.Second) + if err != nil { + t.Fatal(err) + } + reader.now = func() time.Time { return start.Add(time.Minute) } + result, err := reader.Query(t.Context(), testQuery(start)) + if err != nil { + t.Fatal(err) + } + if len(result.Series) != 1 || !result.Series[0].StartedAt.Equal(time.Unix(0, firstStarted)) { + t.Fatalf("start-time jitter split one allocation: %+v", result.Series) + } + if len(result.Series[0].Points) != 1 || result.Series[0].Points[0].CPUUtilizationRatio == nil || *result.Series[0].Points[0].CPUUtilizationRatio != .05 { + t.Fatalf("allocation CPU interval was not retained: %+v", result.Series[0].Points) + } +} + func runtimehistoryResponseValidation(query runtimehistory.Query, result runtimehistory.Result) error { response := runtimehistory.Response{ Capabilities: testCapabilities(), Scope: query.Scope, diff --git a/services/agents-api/internal/runtimehistory/storeresolver/resolver.go b/services/agents-api/internal/runtimehistory/storeresolver/resolver.go index e18ee3912..171743ec0 100644 --- a/services/agents-api/internal/runtimehistory/storeresolver/resolver.go +++ b/services/agents-api/internal/runtimehistory/storeresolver/resolver.go @@ -25,8 +25,8 @@ func NewResolver(value environmentStore) (*Resolver, error) { // ResolveRuntimeHistoryScope authorizes the Session through the Core store and // returns only durable Core identity. It intentionally does not resolve a -// current allocation: retained history may contain earlier allocations or -// multiple compute incarnations that the Reader must keep separately fenced. +// current allocation: retained history may contain earlier allocations. The +// Reader keeps each durable allocation as one continuous series. func (r *Resolver) ResolveRuntimeHistoryScope(ctx context.Context, tenantID, sessionID string) (runtimehistory.Scope, error) { environment, err := r.store.GetSessionEnvironment(ctx, tenantID, sessionID) if err != nil { diff --git a/services/agents-api/internal/runtimehistory/types.go b/services/agents-api/internal/runtimehistory/types.go index 55b83caa5..b3414c3cc 100644 --- a/services/agents-api/internal/runtimehistory/types.go +++ b/services/agents-api/internal/runtimehistory/types.go @@ -161,8 +161,9 @@ type Series struct { } // Point represents one server-selected bucket. CPUUtilizationRatio is derived -// only from ordered cumulative counters within this Series' allocation and -// StartedAt fence. Memory values are the final observed values in the bucket. +// only from ordered cumulative counters within this Series' allocation. A +// counter regression resets the rate baseline. Memory values are the final +// observed values in the bucket. type Point struct { Start, End time.Time FirstObservedAt, LastObservedAt time.Time @@ -255,7 +256,7 @@ func validateResult(query Query, result Result, now time.Time) error { if series.Scope != query.Scope || !canonicalUUID(series.AllocationID) || !validPublicTime(series.StartedAt) || series.StartedAt.After(result.GeneratedAt) || !series.StartedAt.Before(query.End) || !providerTypePattern.MatchString(series.ProviderType) { return errors.New("invalid Runtime history series identity") } - key := series.AllocationID + "\x00" + series.StartedAt.UTC().Format(time.RFC3339Nano) + key := series.AllocationID if seriesKeys[key] { return errors.New("duplicate Runtime history series") } From 1d6cf998159f34c934ff7ea82807038851d60fb9 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 19:16:40 +0800 Subject: [PATCH 35/51] Persist canonical Session token throughput --- apps/web/e2e/agents-lifecycle.spec.ts | 11 +- .../features/dashboard/RuntimeTrendCharts.tsx | 2 +- .../dashboard/runtime-history.test.ts | 36 ++++- .../src/features/dashboard/runtime-history.ts | 19 ++- .../features/dashboard/runtime-trends.test.ts | 6 +- .../src/features/dashboard/runtime-trends.ts | 46 ++++-- contracts/agents-api/openapi.yaml | 38 ++++- contracts/agents-api/runtime-history-api.md | 18 ++- .../runtime-observability-design.md | 22 ++- contracts/agents-api/runtime-observability.md | 4 +- contracts/agents-api/v1/runtime_history.go | 29 ++-- .../src/runtime-history-projection.ts | 42 +++++- .../agents-client/src/runtime-history.test.ts | 27 +++- packages/agents-client/src/types.ts | 11 +- .../agents-api/cmd/server/runtime_history.go | 2 +- .../internal/api/runtime_history.go | 9 +- .../internal/api/runtime_history_test.go | 9 +- .../runtimehistory/clickhousereader/reader.go | 137 +++++++++++++++++- .../clickhousereader/reader_test.go | 60 +++++++- .../internal/runtimehistory/service.go | 1 + .../internal/runtimehistory/service_test.go | 20 ++- .../internal/runtimehistory/types.go | 34 ++++- .../internal/runtimeobs/exporter.go | 10 ++ .../internal/runtimeobs/identity.go | 10 ++ .../runtimeobs/otlpexporter/exporter.go | 18 +++ .../runtimeobs/otlpexporter/exporter_test.go | 23 +++ .../runtimeobs/storeresolver/resolver.go | 55 +++++++ .../runtimeobs/storeresolver/resolver_test.go | 38 ++++- .../clickhouse/002_session_token_usage.sql | 31 ++++ .../runtime-history/clickhouse/README.md | 7 +- 30 files changed, 701 insertions(+), 74 deletions(-) create mode 100644 services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index db4678581..d5b41eab8 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2798,7 +2798,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa body: JSON.stringify({ object: "agent.runtime_history_capabilities", available: true, reason: null, collection_mode: "periodic", sample_interval_seconds: 30, retention_seconds: 604_800, minimum_step_seconds: 30, - maximum_range_seconds: 86_400, maximum_points: 1_000, metrics: ["cpu", "memory"], + maximum_range_seconds: 86_400, maximum_points: 1_000, metrics: ["cpu", "memory", "tokens"], }), })); await page.route(`**/v1/agents/sessions/${sessionId}/runtime-history*`, async (route) => { @@ -2840,6 +2840,13 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa environment_id: environmentId, allocation_id: allocationId, started_at: { seconds: startedAt, nanoseconds: 123_456_789 }, provider_type: "docker", points, }], + token_usage: points.map((_, index) => ({ + start: start + index * 30, + end: start + (index + 1) * 30, + sampled_at: start + index * 30 + 20, + input_tokens: 10_000 + index * 300, + output_tokens: 2_000 + index * 60, + })), }), }); }); @@ -2852,7 +2859,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); await expect(dashboard).toContainText("120 buckets"); await expect(dashboard).toContainText("120/120 observations"); - await expect(dashboard.getByText("Live-only metric", { exact: true })).toBeVisible(); + await expect(dashboard.getByText("Token throughput durable trend available")).toBeAttached(); const durableCpuChart = dashboard.getByLabel("CPU usage: 120 retained buckets"); await expect(dashboard.getByRole("region", { name: "CPU usage durable history chart" })).toBeVisible(); const durableCpuCard = durableCpuChart.locator("xpath=ancestor::section[contains(@class, 'dashboard-runtime-trend-card')]"); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index a2d4034c7..70312bc64 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -557,7 +557,7 @@ export function RuntimeTrendCharts({ `${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> - `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "Live-only metric" : undefined} emptyDetail={durable ? "Runtime history does not duplicate canonical token usage" : undefined} /> + `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained token samples" : undefined} /> ); } diff --git a/apps/web/src/features/dashboard/runtime-history.test.ts b/apps/web/src/features/dashboard/runtime-history.test.ts index 42d5614bb..fb4ec06f6 100644 --- a/apps/web/src/features/dashboard/runtime-history.test.ts +++ b/apps/web/src/features/dashboard/runtime-history.test.ts @@ -63,7 +63,7 @@ const capabilities: RuntimeHistoryCapabilities = { minimum_step_seconds: 30, maximum_range_seconds: 86_400, maximum_points: 1_000, - metrics: ["cpu", "memory"], + metrics: ["cpu", "memory", "tokens"], }; function history(overrides: Partial = {}): RuntimeHistory { @@ -105,6 +105,10 @@ function history(overrides: Partial = {}): RuntimeHistory { memory: { contributor_count: 1, usage_bytes: 768, limit_bytes: 1_024 }, }], }], + token_usage: [ + { start: 100, end: 130, sampled_at: 110, input_tokens: 10, output_tokens: 5 }, + { start: 130, end: 160, sampled_at: 140, input_tokens: 40, output_tokens: 15 }, + ], ...overrides, }; } @@ -116,7 +120,7 @@ describe("Runtime Durable Dashboard history", () => { expect(runtimeTrendSourceAfterHistoryUnavailable("live")).toBe("live"); }); - it("projects persisted buckets without synthesizing token history", () => { + it("projects persisted Runtime and canonical token history", () => { const samples = runtimeDurableTrendSamples([session], [history()]); expect(samples).toHaveLength(2); expect(samples[0]).toMatchObject({ @@ -128,6 +132,7 @@ describe("Runtime Durable Dashboard history", () => { targets: [{ label: "Durable worker", cpuRatio: .25, uptimeSeconds: 99.5 }], }); expect(samples[1]?.targets[0]?.uptimeSeconds).toBe(129.5); + expect(samples[1]).toMatchObject({ inputTokensPerMinute: 60, outputTokensPerMinute: 20 }); }); it("keeps aggregate memory absent when any queried target has no memory value", () => { @@ -141,6 +146,33 @@ describe("Runtime Durable Dashboard history", () => { expect(samples.every((sample) => sample.memoryUsageBytes === null && sample.memoryLimitBytes === null)).toBe(true); }); + it("keeps aggregate token throughput absent when any queried Session lacks usage", () => { + const second = { ...session, id: "44444444-4444-4444-8444-444444444444" } as AgentSession; + const secondHistory = history({ session_id: second.id, token_usage: [] }); + const samples = runtimeDurableTrendSamples([session, second], [history(), secondHistory]); + expect(samples.every((sample) => sample.inputTokensPerMinute === null && sample.outputTokensPerMinute === null)).toBe(true); + }); + + it("derives each Session token rate from its actual sample interval", () => { + const samples = runtimeDurableTrendSamples([session], [history({ + token_usage: [ + { start: 100, end: 130, sampled_at: 105, input_tokens: 10, output_tokens: 5 }, + { start: 130, end: 160, sampled_at: 150, input_tokens: 40, output_tokens: 20 }, + ], + })]); + expect(samples[1]).toMatchObject({ inputTokensPerMinute: 40, outputTokensPerMinute: 20 }); + }); + + it("keeps a token counter regression as a gap instead of inventing throughput", () => { + const samples = runtimeDurableTrendSamples([session], [history({ + token_usage: [ + { start: 100, end: 130, sampled_at: 110, input_tokens: 100, output_tokens: 20 }, + { start: 130, end: 160, sampled_at: 140, input_tokens: 90, output_tokens: 30 }, + ], + })]); + expect(samples[1]).toMatchObject({ inputTokensPerMinute: null, outputTokensPerMinute: 20 }); + }); + it("loads capability-gated Session histories with a bounded common range", async () => { const calls: Array<{ sessionID: string; start: number; end: number; maxPoints?: number }> = []; const core: Pick = { diff --git a/apps/web/src/features/dashboard/runtime-history.ts b/apps/web/src/features/dashboard/runtime-history.ts index 4eab4def2..36290cc10 100644 --- a/apps/web/src/features/dashboard/runtime-history.ts +++ b/apps/web/src/features/dashboard/runtime-history.ts @@ -6,7 +6,7 @@ import type { } from "@agents-core-web/agents-client"; import type { RuntimeDashboardSnapshot } from "./runtime-snapshot"; -import type { RuntimeTrendSample, RuntimeTrendTarget } from "./runtime-trends"; +import { deriveTokenThroughput, type RuntimeTrendSample, type RuntimeTrendTarget } from "./runtime-trends"; export const RUNTIME_DURABLE_TARGET_LIMIT = 24; export const RUNTIME_DURABLE_MAX_POINTS = 120; @@ -82,6 +82,7 @@ interface MutableBucket { sampledAt: number; targets: Map; memory: Map; + tokens: Map; } function incarnationLabel(title: string, history: RuntimeHistory, startedAt: number): string { @@ -102,7 +103,7 @@ export function runtimeDurableTrendSamples( const bucket = (sampledAt: number): MutableBucket => { let value = buckets.get(sampledAt); if (!value) { - value = { sampledAt, targets: new Map(), memory: new Map() }; + value = { sampledAt, targets: new Map(), memory: new Map(), tokens: new Map() }; buckets.set(sampledAt, value); } return value; @@ -110,6 +111,13 @@ export function runtimeDurableTrendSamples( for (const history of histories) { for (const coverage of history.coverage.buckets) bucket(coverage.end * 1_000); + for (const usage of history.token_usage) { + bucket(usage.end * 1_000).tokens.set(history.session_id, { + sampledAt: usage.sampled_at * 1_000, + inputTokens: usage.input_tokens, + outputTokens: usage.output_tokens, + }); + } for (const series of history.series) { const startedAt = series.started_at.seconds + series.started_at.nanoseconds / 1_000_000_000; const targetID = `${history.session_id}:${series.allocation_id}:${series.started_at.seconds}:${series.started_at.nanoseconds}`; @@ -136,7 +144,7 @@ export function runtimeDurableTrendSamples( } } - return [...buckets.values()].sort((left, right) => left.sampledAt - right.sampledAt).map((value) => { + const samples = [...buckets.values()].sort((left, right) => left.sampledAt - right.sampledAt).map((value) => { const completeMemory = sessions.length > 0 && value.memory.size === sessions.length; return { sampledAt: value.sampledAt, @@ -148,11 +156,14 @@ export function runtimeDurableTrendSamples( memoryLimitBytes: completeMemory ? [...value.memory.values()].reduce((total, current) => total + current.limit, 0) : null, - tokenTotals: [], + tokenTotals: value.tokens.size === sessions.length + ? [...value.tokens.entries()].map(([sessionId, usage]) => ({ sessionId, ...usage })) + : [], inputTokensPerMinute: null, outputTokensPerMinute: null, }; }); + return deriveTokenThroughput(samples); } export async function loadRuntimeDurableSnapshot( diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts index fa33ec427..f2cd78a1f 100644 --- a/apps/web/src/features/dashboard/runtime-trends.test.ts +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -202,7 +202,7 @@ describe("Runtime live-window trends", () => { expect(samples.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); }); - it("matches token counters by Session without creating churn spikes", () => { + it("keeps Session-set churn as a gap instead of publishing partial throughput", () => { const first = snapshot(60_000, { input: 100, output: 20 }); const second = snapshot(120_000, { input: 220, output: 50 }); second.sessions.push({ @@ -216,8 +216,8 @@ describe("Runtime live-window trends", () => { samples = appendRuntimeTrendSample(samples, third); expect(tokenThroughput(samples)).toEqual([ { sampledAt: 60_000, inputPerMinute: null, outputPerMinute: null }, - { sampledAt: 120_000, inputPerMinute: 120, outputPerMinute: 30 }, - { sampledAt: 180_000, inputPerMinute: 30, outputPerMinute: null }, + { sampledAt: 120_000, inputPerMinute: null, outputPerMinute: null }, + { sampledAt: 180_000, inputPerMinute: null, outputPerMinute: null }, ]); }); diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts index 76667bf4f..43154954b 100644 --- a/apps/web/src/features/dashboard/runtime-trends.ts +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -40,6 +40,7 @@ export interface RuntimeTrendSample { export interface RuntimeTrendTokenTotal { sessionId: string; + sampledAt: number; inputTokens: number; outputTokens: number; } @@ -90,13 +91,13 @@ function uptimeSeconds(observation: RuntimeObservation): number | null { : null; } -function tokenTotals(sessions: readonly AgentSession[]): RuntimeTrendTokenTotal[] { +function tokenTotals(sessions: readonly AgentSession[], sampledAt: number): RuntimeTrendTokenTotal[] { return sessions.flatMap((session): RuntimeTrendTokenTotal[] => { const inputTokens = safeInteger(session.usage?.input_tokens); const outputTokens = safeInteger(session.usage?.output_tokens); return inputTokens === null || outputTokens === null ? [] - : [{ sessionId: session.id, inputTokens, outputTokens }]; + : [{ sessionId: session.id, sampledAt, inputTokens, outputTokens }]; }); } @@ -164,7 +165,7 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT memoryLimitBytes: pairedMemory.length === 0 ? null : pairedMemory.reduce((total, target) => total + (target.memoryLimitBytes ?? 0), 0), - tokenTotals: tokenTotals(snapshot.sessions), + tokenTotals: tokenTotals(snapshot.sessions, snapshot.loadedAt), inputTokensPerMinute: null, outputTokensPerMinute: null, }; @@ -226,27 +227,46 @@ function tokenRate( previous: RuntimeTrendSample, next: RuntimeTrendSample, ): Pick { - const elapsedMinutes = (next.sampledAt - previous.sampledAt) / 60_000; - if (elapsedMinutes <= 0) return { inputTokensPerMinute: null, outputTokensPerMinute: null }; const previousTotals = new Map(previous.tokenTotals.map((total) => [total.sessionId, total])); + if (previousTotals.size === 0 || previousTotals.size !== next.tokenTotals.length || + next.tokenTotals.some((current) => !previousTotals.has(current.sessionId))) { + return { inputTokensPerMinute: null, outputTokensPerMinute: null }; + } const pairs = next.tokenTotals.flatMap((current): Array<[RuntimeTrendTokenTotal, RuntimeTrendTokenTotal]> => { const prior = previousTotals.get(current.sessionId); return prior ? [[prior, current]] : []; }); - const inputDeltas = pairs.map(([prior, current]) => current.inputTokens - prior.inputTokens); - const outputDeltas = pairs.map(([prior, current]) => current.outputTokens - prior.outputTokens); - const inputDelta = inputDeltas.length > 0 && inputDeltas.every((delta) => delta >= 0) - ? inputDeltas.reduce((total, delta) => total + delta, 0) + const inputRates = pairs.map(([prior, current]) => { + const elapsedMinutes = (current.sampledAt - prior.sampledAt) / 60_000; + const delta = current.inputTokens - prior.inputTokens; + return elapsedMinutes > 0 && delta >= 0 ? delta / elapsedMinutes : null; + }); + const outputRates = pairs.map(([prior, current]) => { + const elapsedMinutes = (current.sampledAt - prior.sampledAt) / 60_000; + const delta = current.outputTokens - prior.outputTokens; + return elapsedMinutes > 0 && delta >= 0 ? delta / elapsedMinutes : null; + }); + const inputRate = inputRates.length > 0 && inputRates.every((rate) => rate !== null) + ? inputRates.reduce((total, rate) => total + (rate ?? 0), 0) : null; - const outputDelta = outputDeltas.length > 0 && outputDeltas.every((delta) => delta >= 0) - ? outputDeltas.reduce((total, delta) => total + delta, 0) + const outputRate = outputRates.length > 0 && outputRates.every((rate) => rate !== null) + ? outputRates.reduce((total, rate) => total + (rate ?? 0), 0) : null; return { - inputTokensPerMinute: inputDelta === null ? null : inputDelta / elapsedMinutes, - outputTokensPerMinute: outputDelta === null ? null : outputDelta / elapsedMinutes, + inputTokensPerMinute: inputRate, + outputTokensPerMinute: outputRate, }; } +export function deriveTokenThroughput(samples: readonly RuntimeTrendSample[]): RuntimeTrendSample[] { + return samples.map((sample, index) => { + const current = { ...sample }; + const previous = samples[index - 1]; + if (previous) Object.assign(current, tokenRate(previous, sample)); + return current; + }); +} + export function appendRuntimeTrendSample( current: readonly RuntimeTrendSample[], snapshot: RuntimeDashboardSnapshot, diff --git a/contracts/agents-api/openapi.yaml b/contracts/agents-api/openapi.yaml index 4802fece9..4999a105c 100644 --- a/contracts/agents-api/openapi.yaml +++ b/contracts/agents-api/openapi.yaml @@ -1096,6 +1096,11 @@ definitions: enum: - durable type: string + token_usage: + items: + $ref: '#/definitions/v1.RuntimeHistoryTokenUsagePoint' + maxItems: 10000 + type: array required: - coverage - generated_at @@ -1105,6 +1110,7 @@ definitions: - series - session_id - source + - token_usage type: object v1.RuntimeHistoryCPU: properties: @@ -1150,8 +1156,9 @@ definitions: enum: - cpu - memory + - tokens type: string - maxItems: 2 + maxItems: 3 type: array minimum_step_seconds: maximum: 9007199254740991 @@ -1393,6 +1400,35 @@ definitions: - nanoseconds - seconds type: object + v1.RuntimeHistoryTokenUsagePoint: + properties: + end: + maximum: 9007199254740991 + minimum: 1 + type: integer + input_tokens: + maximum: 9007199254740991 + minimum: 0 + type: integer + output_tokens: + maximum: 9007199254740991 + minimum: 0 + type: integer + sampled_at: + maximum: 9007199254740991 + minimum: 0 + type: integer + start: + maximum: 9007199254740991 + minimum: 0 + type: integer + required: + - end + - input_tokens + - output_tokens + - sampled_at + - start + type: object v1.RuntimeInstance: properties: allocation_id: diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md index f795ae8ea..bb5abc128 100644 --- a/contracts/agents-api/runtime-history-api.md +++ b/contracts/agents-api/runtime-history-api.md @@ -37,7 +37,7 @@ The route accepts no query parameters. An unconfigured Core returns: `available=true` requires `collection_mode=periodic`, a qualified positive sample interval, a validated Reader, retention and query bounds, and at least one of -`cpu` or `memory`. A Reader backed only by request-triggered samples returns +`cpu`, `memory`, or `tokens`. A Reader backed only by request-triggered samples returns `reason=periodic_collection_required`; Web must not call that data Durable. Malformed capability configuration fails closed as unconfigured. Backend names, URLs, credentials, table names, and tenant data are never capability fields. @@ -78,7 +78,8 @@ query inputs. "expected_sample_count": 120, "buckets": [] }, - "series": [] + "series": [], + "token_usage": [] } ``` @@ -107,9 +108,14 @@ successive counter intervals are assigned to the bucket containing their right endpoint and combined by CPU-capacity time. Memory values are the last observed values in a bucket. Every point contains observation and contributor counts. Missing values are null and gaps remain gaps. -Numeric zero is retained as an observed value. The endpoint does not return token -history: token throughput remains sourced from canonical Session Usage and is -Live-only until a separate durable usage contract exists. +Numeric zero is retained as an observed value. + +`token_usage` is Session-scoped rather than allocation-scoped. Each point is the +last cumulative canonical Session Usage snapshot sampled in that bucket and +contains `input_tokens`, `output_tokens`, and `sampled_at`. Web derives throughput +only from adjacent nondecreasing cumulative points. A missing measurement or a +counter regression produces a gap; it is never filled with zero. These counters +are measured model usage, not price, cost, or billing records. Core Web queries each current managed Session through this boundary with bounded concurrency and an all-or-nothing target budget. It offers 1h, 6h, and 24h History @@ -150,7 +156,7 @@ interface AgentCore { The client validates exact fields, capability consistency, requested-range echo, Session identity, half-open bucket ordering, coverage totals, incarnation identity, -contributor counts, nullability, finite numbers, and response size. Unknown fields +contributor counts, token usage ordering, nullability, finite numbers, and response size. Unknown fields or malformed data reject the entire response with a 502 client projection error. ## Explicit boundaries diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 3c5287132..7e018798d 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -234,6 +234,8 @@ Recommended instruments are: - `agents.runtime.cpu.capacity` cores; - `agents.runtime.memory.usage` bytes; - `agents.runtime.memory.limit` bytes; +- `agents.session.tokens.input` cumulative measured tokens; +- `agents.session.tokens.output` cumulative measured tokens; - `agents.runtime.sample` success/unavailable count; and - `agents.runtime.sample.duration` seconds. @@ -323,6 +325,8 @@ The OTLP request uses standard protobuf metrics and these instruments: | `agents.runtime.cpu.capacity` | gauge, cores | configured provider capacity | | `agents.runtime.memory.usage` | gauge, bytes | provider memory usage | | `agents.runtime.memory.limit` | gauge, bytes | configured provider limit | +| `agents.session.tokens.input` | gauge, tokens | canonical cumulative Session Usage | +| `agents.session.tokens.output` | gauge, tokens | canonical cumulative Session Usage | | `agents.runtime.sample` | monotonic delta sum | one validated result, including unavailable/unsupported | | `agents.runtime.sample.duration` | delta histogram, seconds | bounded provider read duration | @@ -479,10 +483,18 @@ persisted by Core. ## 12. Token usage boundary -Runtime observations do not duplicate token usage. Web uses the existing canonical -Session Usage snapshot and joins it to Runtime rows by `session_id`. The Dashboard -shows both total and coverage, such as `1.84M reported by 7/8 Sessions`. Missing or -incomplete native usage remains unknown. +Provider Runtime sources do not own or report token usage. During a periodic +history sweep, the Core resolver reads the existing canonical cumulative Session +Usage snapshot from the execution store alongside Runtime identity. The exporter +emits Session-scoped input/output token gauges with the same Session and sampling +time, independently of Docker, microsandbox, Kubernetes, or another provider. +ClickHouse retains those cumulative points separately from allocation/incarnation +series. Web derives throughput from adjacent nondecreasing points. Missing or +incomplete native usage and counter regressions remain gaps, never zero. + +The current snapshot API still does not duplicate Usage fields: Web joins its +existing Session collection by exact `session_id`. The Dashboard reports measured +coverage and never estimates missing usage. Cost and billing stay outside this Core API. A product may join billing in its own authorized backend, never by exposing product credentials to Core Web. @@ -571,7 +583,7 @@ Implemented for the browser-local current-snapshot live window. incarnations without cross-fence CPU derivation. - Implemented: Core Web discovers capabilities, reloads bounded Session histories with bounded concurrency, and exposes explicit Live versus History sources with - 1h, 6h, and 24h Durable ranges. Token throughput remains honestly Live-only. + 1h, 6h, and 24h Durable ranges, including canonical Session token throughput. - Not implemented: exporter queue/drop/error coverage telemetry. ### Phase 5: additional sources diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 57cf77f52..76018fb8d 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -115,7 +115,9 @@ an operator must configure the Reader and qualified periodic collection. ## First-phase boundary The current-snapshot implementation adds no migration, metrics backend, token -duplication, Kubernetes/E2B source, or telemetry-driven lifecycle action. The +duplication, Kubernetes/E2B source, or telemetry-driven lifecycle action. Optional +Durable history separately samples canonical Session Usage into its configured +metrics backend; provider sources still do not own token accounting. The internal source interface admits Docker and microsandbox without changing Session attribution or the existing sandbox lifecycle interface. diff --git a/contracts/agents-api/v1/runtime_history.go b/contracts/agents-api/v1/runtime_history.go index 741c1ba04..0c89b5619 100644 --- a/contracts/agents-api/v1/runtime_history.go +++ b/contracts/agents-api/v1/runtime_history.go @@ -10,18 +10,29 @@ type RuntimeHistoryCapabilities struct { MinimumStepSeconds *int64 `json:"minimum_step_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` MaximumRangeSeconds *int64 `json:"maximum_range_seconds" extensions:"x-nullable" binding:"required" minimum:"1" maximum:"9007199254740991"` MaximumPoints *int `json:"maximum_points" extensions:"x-nullable" binding:"required" minimum:"2" maximum:"10000"` - Metrics []string `json:"metrics" binding:"required" validate:"max=2" enums:"cpu,memory"` + Metrics []string `json:"metrics" binding:"required" validate:"max=3" enums:"cpu,memory,tokens"` } type RuntimeHistory struct { - Object string `json:"object" enums:"agent.runtime_history" binding:"required"` - Source string `json:"source" enums:"durable" binding:"required"` - SessionID string `json:"session_id" binding:"required" format:"uuid"` - RequestedRange RuntimeHistoryRange `json:"requested_range" binding:"required"` - ResolutionSeconds int64 `json:"resolution_seconds" binding:"required" minimum:"1" maximum:"9007199254740991"` - GeneratedAt int64 `json:"generated_at" binding:"required" minimum:"0" maximum:"9007199254740991"` - Coverage RuntimeHistoryCoverage `json:"coverage" binding:"required"` - Series []RuntimeHistorySeries `json:"series" binding:"required" validate:"max=1000"` + Object string `json:"object" enums:"agent.runtime_history" binding:"required"` + Source string `json:"source" enums:"durable" binding:"required"` + SessionID string `json:"session_id" binding:"required" format:"uuid"` + RequestedRange RuntimeHistoryRange `json:"requested_range" binding:"required"` + ResolutionSeconds int64 `json:"resolution_seconds" binding:"required" minimum:"1" maximum:"9007199254740991"` + GeneratedAt int64 `json:"generated_at" binding:"required" minimum:"0" maximum:"9007199254740991"` + Coverage RuntimeHistoryCoverage `json:"coverage" binding:"required"` + Series []RuntimeHistorySeries `json:"series" binding:"required" validate:"max=1000"` + TokenUsage []RuntimeHistoryTokenUsagePoint `json:"token_usage" binding:"required" validate:"max=10000"` +} + +// RuntimeHistoryTokenUsagePoint is the final cumulative canonical Session +// Usage snapshot observed in one server-selected bucket. +type RuntimeHistoryTokenUsagePoint struct { + Start int64 `json:"start" binding:"required" minimum:"0" maximum:"9007199254740991"` + End int64 `json:"end" binding:"required" minimum:"1" maximum:"9007199254740991"` + SampledAt int64 `json:"sampled_at" binding:"required" minimum:"0" maximum:"9007199254740991"` + InputTokens uint64 `json:"input_tokens" binding:"required" minimum:"0" maximum:"9007199254740991"` + OutputTokens uint64 `json:"output_tokens" binding:"required" minimum:"0" maximum:"9007199254740991"` } type RuntimeHistoryRange struct { diff --git a/packages/agents-client/src/runtime-history-projection.ts b/packages/agents-client/src/runtime-history-projection.ts index 38c79df43..dfe9d6662 100644 --- a/packages/agents-client/src/runtime-history-projection.ts +++ b/packages/agents-client/src/runtime-history-projection.ts @@ -8,6 +8,7 @@ import type { RuntimeHistoryPoint, RuntimeHistoryQuery, RuntimeHistorySeries, + RuntimeHistoryTokenUsagePoint, } from "./types"; type InvalidRuntimeHistory = (message?: string) => never; @@ -17,7 +18,7 @@ const capabilityFields = new Set([ "retention_seconds", "minimum_step_seconds", "maximum_range_seconds", "maximum_points", "metrics", ]); const historyFields = new Set([ - "object", "source", "session_id", "requested_range", "resolution_seconds", "generated_at", "coverage", "series", + "object", "source", "session_id", "requested_range", "resolution_seconds", "generated_at", "coverage", "series", "token_usage", ]); const rangeFields = new Set(["start", "end"]); const coverageFields = new Set([ @@ -31,8 +32,9 @@ const timeFields = new Set(["seconds", "nanoseconds"]); const pointFields = new Set([...coveragePointFields, "cpu", "memory"]); const cpuFields = new Set(["contributor_count", "utilization_ratio", "capacity_cores"]); const memoryFields = new Set(["contributor_count", "usage_bytes", "limit_bytes"]); +const tokenUsageFields = new Set(["start", "end", "sampled_at", "input_tokens", "output_tokens"]); const providerTypePattern = /^[a-z][a-z0-9_]{0,31}$/; -const metrics = new Set(["cpu", "memory"]); +const metrics = new Set(["cpu", "memory", "tokens"]); const reasons = new Set(["not_configured", "periodic_collection_required"]); const maximumSeries = 1_000; const maximumTotalPoints = 100_000; @@ -185,10 +187,34 @@ function projectPoint( return { ...coverage, cpu, memory }; } -function ordered(points: readonly RuntimeHistoryCoveragePoint[]): boolean { +function ordered(points: readonly { start: number; end: number }[]): boolean { return points.every((point, index) => index === 0 || points[index - 1]!.end <= point.start); } +function projectTokenUsagePoint( + value: unknown, + requestedStart: number, + requestedEnd: number, + resolution: number, + generatedAt: number, + invalid: InvalidRuntimeHistory, +): RuntimeHistoryTokenUsagePoint { + if ( + !isRecord(value) || !exactFields(value, tokenUsageFields) || + !isNonnegativeInteger(value.start) || !positiveInteger(value.end) || + value.start < requestedStart || value.end <= value.start || value.end > requestedEnd || value.end - value.start > resolution || + !isNonnegativeInteger(value.sampled_at) || value.sampled_at < value.start || value.sampled_at >= value.end || value.sampled_at > generatedAt || + !isNonnegativeInteger(value.input_tokens) || !isNonnegativeInteger(value.output_tokens) + ) return invalid(); + return { + start: value.start, + end: value.end, + sampled_at: value.sampled_at, + input_tokens: value.input_tokens, + output_tokens: value.output_tokens, + }; +} + function projectSeries( value: unknown, requestedStart: number, @@ -236,7 +262,7 @@ export function projectRuntimeHistory( value.requested_range.start !== requested.start || value.requested_range.end !== requested.end || !positiveInteger(value.resolution_seconds) || !isNonnegativeInteger(value.generated_at) || !isRecord(value.coverage) || !exactFields(value.coverage, coverageFields) || !Array.isArray(value.coverage.buckets) || - !Array.isArray(value.series) + !Array.isArray(value.series) || !Array.isArray(value.token_usage) ) return invalid(); const sessionId = canonicalUuid(value.session_id); if (sessionId === null) return invalid(); @@ -271,10 +297,15 @@ export function projectRuntimeHistory( const series = value.series.map((entry) => projectSeries( entry, requested.start, requested.end, value.resolution_seconds as number, generatedAt, maximumPoints, invalid, )); + const tokenUsage = value.token_usage.map((entry) => projectTokenUsagePoint( + entry, requested.start, requested.end, value.resolution_seconds as number, generatedAt, invalid, + )); if ( new Set(series.map((entry) => entry.environment_id)).size > 1 || new Set(series.map((entry) => `${entry.allocation_id}\u0000${entry.started_at.seconds}\u0000${entry.started_at.nanoseconds}`)).size !== series.length || - buckets.length + series.reduce((sum, entry) => sum + entry.points.length, 0) > maximumTotalPoints || + tokenUsage.length > maximumPoints || !ordered(tokenUsage) || + buckets.length + tokenUsage.length + series.reduce((sum, entry) => sum + entry.points.length, 0) > maximumTotalPoints || + tokenUsage.some((point) => point.sampled_at < retainedStart) || series.some((entry) => entry.points.some((point) => point.first_observed_at !== null && point.first_observed_at < retainedStart)) || series.some((entry) => entry.points.some((point) => point.last_observed_at !== null && point.last_observed_at > generatedAt)) ) return invalid(); @@ -294,5 +325,6 @@ export function projectRuntimeHistory( buckets, }, series, + token_usage: tokenUsage, }; } diff --git a/packages/agents-client/src/runtime-history.test.ts b/packages/agents-client/src/runtime-history.test.ts index 948933d50..2da6b6d09 100644 --- a/packages/agents-client/src/runtime-history.test.ts +++ b/packages/agents-client/src/runtime-history.test.ts @@ -36,7 +36,7 @@ function capabilities(overrides: Record = {}): Record = {}): Record { { seconds: 900, nanoseconds: 1 }, ]); expect(value.coverage.sample_count).toBe(2); + expect(value.token_usage[1]).toMatchObject({ input_tokens: 160, output_tokens: 50 }); }); it("rejects invalid requests before fetch", async () => { @@ -211,6 +216,17 @@ describe("Runtime history client", () => { coverage.sample_count = 1; coverage.buckets = [coverage.buckets[1]!]; }], + ["token sample predates retention", (value: Record) => { + const coverage = value.coverage as { retained_start: number; first_sample_at: number; last_sample_at: number; sample_count: number; buckets: Array> }; + coverage.retained_start = 1020; + coverage.first_sample_at = 1070; + coverage.last_sample_at = 1070; + coverage.sample_count = 1; + coverage.buckets = [coverage.buckets[1]!]; + const series = (value.series as Array<{ points: Array> }>)[0]!; + series.points = series.points.slice(1); + (value.token_usage as Array>)[0]!.sampled_at = 1010; + }], ["overlapping buckets", (value: Record) => { const buckets = (value.coverage as { buckets: Array> }).buckets; buckets[1]!.start = 1050; @@ -234,6 +250,15 @@ describe("Runtime history client", () => { foreign.allocation_id = "55555555-5555-4555-8555-555555555555"; series.push(foreign); }], + ["unsafe token usage", (value: Record) => { + (value.token_usage as Array>)[0]!.input_tokens = Number.MAX_SAFE_INTEGER + 1; + }], + ["token sample outside bucket", (value: Record) => { + (value.token_usage as Array>)[0]!.sampled_at = 1060; + }], + ["overlapping token buckets", (value: Record) => { + (value.token_usage as Array>)[1]!.start = 1050; + }], ] as const) { it(`rejects ${name} in history`, async () => { const value = history(); diff --git a/packages/agents-client/src/types.ts b/packages/agents-client/src/types.ts index 44d2f2598..605e42d7d 100644 --- a/packages/agents-client/src/types.ts +++ b/packages/agents-client/src/types.ts @@ -831,7 +831,7 @@ export interface RuntimeObservationList extends ListPage { export type RuntimeHistoryCollectionMode = "on_read" | "periodic"; export type RuntimeHistoryCapabilityReason = "not_configured" | "periodic_collection_required"; -export type RuntimeHistoryMetric = "cpu" | "memory"; +export type RuntimeHistoryMetric = "cpu" | "memory" | "tokens"; export interface RuntimeHistoryCapabilities { object: "agent.runtime_history_capabilities"; @@ -910,6 +910,14 @@ export interface RuntimeHistoryCoverage { buckets: RuntimeHistoryCoveragePoint[]; } +export interface RuntimeHistoryTokenUsagePoint { + start: number; + end: number; + sampled_at: number; + input_tokens: number; + output_tokens: number; +} + export interface RuntimeHistory { object: "agent.runtime_history"; source: "durable"; @@ -919,6 +927,7 @@ export interface RuntimeHistory { generated_at: number; coverage: RuntimeHistoryCoverage; series: RuntimeHistorySeries[]; + token_usage: RuntimeHistoryTokenUsagePoint[]; } export interface AgentCore { diff --git a/services/agents-api/cmd/server/runtime_history.go b/services/agents-api/cmd/server/runtime_history.go index 985d8697d..d75f2bde8 100644 --- a/services/agents-api/cmd/server/runtime_history.go +++ b/services/agents-api/cmd/server/runtime_history.go @@ -129,7 +129,7 @@ func runtimeHistory(ctx context.Context) (runtimeHistorySetup, error) { Capabilities: runtimehistory.Capabilities{ CollectionMode: mode, SampleInterval: interval, Retention: 7 * 24 * time.Hour, MinimumStep: minimumStep, MaximumRange: 24 * time.Hour, MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, - Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory}, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory, runtimehistory.MetricTokens}, }, }) if err != nil { diff --git a/services/agents-api/internal/api/runtime_history.go b/services/agents-api/internal/api/runtime_history.go index b99c7a596..e21caa5a4 100644 --- a/services/agents-api/internal/api/runtime_history.go +++ b/services/agents-api/internal/api/runtime_history.go @@ -216,7 +216,14 @@ func runtimeHistoryResponse(value runtimehistory.Response, expectedTenantID, exp Object: "agent.runtime_history", Source: "durable", SessionID: value.SessionID, RequestedRange: v1.RuntimeHistoryRange{Start: value.Requested.Start.Unix(), End: value.Requested.End.Unix()}, ResolutionSeconds: int64(value.Resolution / time.Second), GeneratedAt: value.GeneratedAt.Unix(), Coverage: coverage, - Series: make([]v1.RuntimeHistorySeries, 0, len(value.Series)), + Series: make([]v1.RuntimeHistorySeries, 0, len(value.Series)), + TokenUsage: make([]v1.RuntimeHistoryTokenUsagePoint, 0, len(value.TokenUsage)), + } + for _, point := range value.TokenUsage { + response.TokenUsage = append(response.TokenUsage, v1.RuntimeHistoryTokenUsagePoint{ + Start: point.Start.Unix(), End: point.End.Unix(), SampledAt: point.SampledAt.Unix(), + InputTokens: point.InputTokens, OutputTokens: point.OutputTokens, + }) } for _, series := range value.Series { converted := v1.RuntimeHistorySeries{ diff --git a/services/agents-api/internal/api/runtime_history_test.go b/services/agents-api/internal/api/runtime_history_test.go index 3ff2532eb..c2d25fd49 100644 --- a/services/agents-api/internal/api/runtime_history_test.go +++ b/services/agents-api/internal/api/runtime_history_test.go @@ -36,7 +36,7 @@ func historyCapabilities(mode runtimehistory.CollectionMode) runtimehistory.Capa value := runtimehistory.Capabilities{ CollectionMode: mode, Retention: 7 * 24 * time.Hour, MinimumStep: 30 * time.Second, MaximumRange: 24 * time.Hour, MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, - Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory}, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory, runtimehistory.MetricTokens}, } if mode == runtimehistory.CollectionPeriodic { value.SampleInterval = 30 * time.Second @@ -113,8 +113,9 @@ func TestRuntimeHistoryRouteBindsAuthenticatedSessionAndPreservesCoverage(t *tes service.response = runtimehistory.Response{ Capabilities: service.capabilities, Scope: scope, Requested: runtimehistory.Range{Start: start, End: now, MaxPoints: 60}, Resolution: time.Minute, GeneratedAt: now, RetainedFrom: &start, - Coverage: []runtimehistory.CoveragePoint{coveragePoint}, - Series: []runtimehistory.Series{{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker", Points: []runtimehistory.Point{resourcePoint}}}, + Coverage: []runtimehistory.CoveragePoint{coveragePoint}, + Series: []runtimehistory.Series{{Scope: scope, AllocationID: allocationID, StartedAt: start.Add(-time.Minute), ProviderType: "docker", Points: []runtimehistory.Point{resourcePoint}}}, + TokenUsage: []runtimehistory.TokenUsagePoint{{Start: start, End: start.Add(time.Minute), SampledAt: start.Add(10 * time.Second), InputTokens: 120, OutputTokens: 30}}, } secondIncarnation := service.response.Series[0] secondIncarnation.StartedAt = secondIncarnation.StartedAt.Add(time.Nanosecond) @@ -124,7 +125,7 @@ func TestRuntimeHistoryRouteBindsAuthenticatedSessionAndPreservesCoverage(t *tes t.Fatalf("history returned %d: %s", response.Code, response.Body) } var value v1.RuntimeHistory - if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history" || value.Source != "durable" || value.SessionID != sessionID || value.ResolutionSeconds != 60 || value.Coverage.SampleCount != 1 || value.Coverage.ExpectedSampleCount != 120 || len(value.Coverage.Buckets) != 1 || len(value.Series) != 2 || len(value.Series[0].Points) != 1 || value.Series[0].StartedAt.Seconds != value.Series[1].StartedAt.Seconds || value.Series[0].StartedAt.Nanoseconds == value.Series[1].StartedAt.Nanoseconds { + if json.Unmarshal(response.Body.Bytes(), &value) != nil || value.Object != "agent.runtime_history" || value.Source != "durable" || value.SessionID != sessionID || value.ResolutionSeconds != 60 || value.Coverage.SampleCount != 1 || value.Coverage.ExpectedSampleCount != 120 || len(value.Coverage.Buckets) != 1 || len(value.Series) != 2 || len(value.Series[0].Points) != 1 || value.Series[0].StartedAt.Seconds != value.Series[1].StartedAt.Seconds || value.Series[0].StartedAt.Nanoseconds == value.Series[1].StartedAt.Nanoseconds || len(value.TokenUsage) != 1 || value.TokenUsage[0].InputTokens != 120 || value.TokenUsage[0].OutputTokens != 30 { t.Fatalf("invalid history response: %s", response.Body) } point := value.Series[0].Points[0] diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader.go index 367720497..e80b50eed 100644 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go +++ b/services/agents-api/internal/runtimehistory/clickhousereader/reader.go @@ -57,6 +57,27 @@ GROUP BY ORDER BY resolved_at_unix_nano, allocation_id, started_at_unix_nano, metric_name LIMIT {row_limit:UInt64}` +const tokenUsageQuery = ` +SELECT + resolved_at_unix_nano, + metric_name, + min(value) AS minimum_value, + max(value) AS maximum_value +FROM runtime_history_metrics +WHERE tenant_id = {tenant_id:String} + AND session_id = {session_id:String} + AND environment_id = {environment_id:String} + AND collection_source = 'periodic' + AND resolved_at_unix_nano >= {raw_start:Int64} + AND resolved_at_unix_nano < {end:Int64} + AND metric_name IN ( + 'agents.session.tokens.input', + 'agents.session.tokens.output' + ) +GROUP BY resolved_at_unix_nano, metric_name +ORDER BY resolved_at_unix_nano, metric_name +LIMIT {row_limit:UInt64}` + type Config struct { Address string Database string @@ -173,7 +194,41 @@ func (r *Reader) Query(ctx context.Context, query runtimehistory.Query) (runtime if count >= rowLimit { return runtimehistory.Result{}, errors.New("Runtime history raw result exceeds limit") } - return aggregate(query, generatedAt, raw) + result, err := aggregate(query, generatedAt, raw) + if err != nil || !hasMetric(r.capabilities.Metrics, runtimehistory.MetricTokens) { + return result, err + } + tokenRowLimit := query.MaximumTotalPoints*2 + 1 + tokenRows, err := r.client.Query(queryCtx, tokenUsageQuery, + clickhouse.Named("tenant_id", query.TenantID), + clickhouse.Named("session_id", query.SessionID), + clickhouse.Named("environment_id", query.EnvironmentID), + clickhouse.Named("raw_start", query.Start.UnixNano()), + clickhouse.Named("end", query.End.UnixNano()), + clickhouse.Named("row_limit", uint64(tokenRowLimit)), + ) + if err != nil { + return runtimehistory.Result{}, errors.New("query Runtime history token usage") + } + defer tokenRows.Close() + rawTokens, tokenCount, err := readTokenUsage(tokenRows, query.Start, query.End) + if err != nil { + return runtimehistory.Result{}, err + } + if tokenCount >= tokenRowLimit { + return runtimehistory.Result{}, errors.New("Runtime history token result exceeds limit") + } + result.TokenUsage, err = aggregateTokenUsage(query, rawTokens) + return result, err +} + +func hasMetric(metrics []runtimehistory.Metric, expected runtimehistory.Metric) bool { + for _, metric := range metrics { + if metric == expected { + return true + } + } + return false } func (r *Reader) validateQuery(query runtimehistory.Query) error { @@ -234,6 +289,86 @@ type rawSample struct { metrics map[string]float64 } +type rawTokenUsage struct { + resolvedAt time.Time + inputTokens, outputTokens *uint64 +} + +func readTokenUsage(resultRows rows, start, end time.Time) ([]rawTokenUsage, int, error) { + byResolved := map[int64]*rawTokenUsage{} + rowCount := 0 + for resultRows.Next() { + rowCount++ + var resolvedNano int64 + var metric string + var minimum, maximum float64 + if err := resultRows.Scan(&resolvedNano, &metric, &minimum, &maximum); err != nil { + return nil, rowCount, errors.New("scan Runtime history token row") + } + resolvedAt := time.Unix(0, resolvedNano).UTC() + value, err := safeUint64(minimum) + if resolvedNano < 0 || resolvedAt.Before(start) || !resolvedAt.Before(end) || minimum != maximum || err != nil { + return nil, rowCount, errors.New("invalid Runtime history token row") + } + usage := byResolved[resolvedNano] + if usage == nil { + usage = &rawTokenUsage{resolvedAt: resolvedAt} + byResolved[resolvedNano] = usage + } + switch metric { + case otlpexporter.TokenInputName: + if usage.inputTokens != nil { + return nil, rowCount, errors.New("duplicate Runtime history input token row") + } + usage.inputTokens = &value + case otlpexporter.TokenOutputName: + if usage.outputTokens != nil { + return nil, rowCount, errors.New("duplicate Runtime history output token row") + } + usage.outputTokens = &value + default: + return nil, rowCount, errors.New("invalid Runtime history token metric") + } + } + if err := resultRows.Err(); err != nil { + return nil, rowCount, errors.New("iterate Runtime history token rows") + } + result := make([]rawTokenUsage, 0, len(byResolved)) + for _, usage := range byResolved { + if usage.inputTokens == nil || usage.outputTokens == nil { + return nil, rowCount, errors.New("incomplete Runtime history token usage") + } + result = append(result, *usage) + } + sort.Slice(result, func(i, j int) bool { return result[i].resolvedAt.Before(result[j].resolvedAt) }) + return result, rowCount, nil +} + +func aggregateTokenUsage(query runtimehistory.Query, raw []rawTokenUsage) ([]runtimehistory.TokenUsagePoint, error) { + latest := map[int]*rawTokenUsage{} + for _, usage := range raw { + index := bucketIndex(query, usage.resolvedAt) + if index < 0 { + continue + } + previous, ok := latest[index] + if !ok || usage.resolvedAt.After(previous.resolvedAt) { + value := usage + latest[index] = &value + } + } + result := make([]runtimehistory.TokenUsagePoint, 0, len(latest)) + for _, index := range sortedIndexes(latest) { + usage := latest[index] + start, end := bucketBounds(query, index) + result = append(result, runtimehistory.TokenUsagePoint{ + Start: start, End: end, SampledAt: usage.resolvedAt, + InputTokens: *usage.inputTokens, OutputTokens: *usage.outputTokens, + }) + } + return result, nil +} + func readSamples(resultRows rows, rawStart, end time.Time) ([]*rawSample, int, error) { byKey := map[string]*rawSample{} rowCount := 0 diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go index 69abcdbd0..476a71145 100644 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go +++ b/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go @@ -22,15 +22,23 @@ const ( ) type fakeClient struct { - rows rows - err error - statement string - args []any - closed bool + rows rows + rowsQueue []rows + err error + statement string + statements []string + args []any + closed bool } func (c *fakeClient) Query(_ context.Context, statement string, args ...any) (rows, error) { c.statement, c.args = statement, append([]any(nil), args...) + c.statements = append(c.statements, statement) + if len(c.rowsQueue) > 0 { + value := c.rowsQueue[0] + c.rowsQueue = c.rowsQueue[1:] + return value, c.err + } return c.rows, c.err } @@ -156,6 +164,48 @@ func TestReaderQueriesMandatoryScopeAndAggregatesFencedHistory(t *testing.T) { } } +func TestReaderReturnsSessionTokenUsageIndependentOfRuntimeIncarnation(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + first := start.Add(10 * time.Second).UnixNano() + second := start.Add(40 * time.Second).UnixNano() + client := &fakeClient{rowsQueue: []rows{ + &fakeRows{}, + &fakeRows{values: [][]any{ + {first, otlpexporter.TokenInputName, float64(100), float64(100)}, + {first, otlpexporter.TokenOutputName, float64(20), float64(20)}, + {second, otlpexporter.TokenInputName, float64(160), float64(160)}, + {second, otlpexporter.TokenOutputName, float64(50), float64(50)}, + }}, + }} + capabilities := testCapabilities() + capabilities.Metrics = append(capabilities.Metrics, runtimehistory.MetricTokens) + reader, err := newReader(client, capabilities, 2*time.Second) + if err != nil { + t.Fatal(err) + } + reader.now = func() time.Time { return start.Add(time.Minute) } + result, err := reader.Query(t.Context(), testQuery(start)) + if err != nil { + t.Fatal(err) + } + if len(client.statements) != 2 || !strings.Contains(client.statements[1], "agents.session.tokens.input") { + t.Fatalf("token query was not issued separately: %#v", client.statements) + } + if len(result.TokenUsage) != 2 || result.TokenUsage[0].InputTokens != 100 || result.TokenUsage[1].OutputTokens != 50 { + t.Fatalf("unexpected token usage history: %+v", result.TokenUsage) + } +} + +func TestReaderRejectsIncompleteSessionTokenUsage(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + rows := &fakeRows{values: [][]any{{ + start.Add(10 * time.Second).UnixNano(), otlpexporter.TokenInputName, float64(100), float64(100), + }}} + if _, _, err := readTokenUsage(rows, start, start.Add(time.Minute)); err == nil { + t.Fatal("incomplete token usage was accepted") + } +} + func runtimehistoryResponseValidation(query runtimehistory.Query, result runtimehistory.Result) error { response := runtimehistory.Response{ Capabilities: testCapabilities(), Scope: query.Scope, diff --git a/services/agents-api/internal/runtimehistory/service.go b/services/agents-api/internal/runtimehistory/service.go index 40859ce0f..59043201c 100644 --- a/services/agents-api/internal/runtimehistory/service.go +++ b/services/agents-api/internal/runtimehistory/service.go @@ -83,6 +83,7 @@ func (s *Service) QuerySession(ctx context.Context, tenantID, sessionID string, RetainedFrom: cloneTime(result.RetainedFrom), Coverage: append([]CoveragePoint(nil), result.Coverage...), Series: cloneSeries(result.Series), + TokenUsage: append([]TokenUsagePoint(nil), result.TokenUsage...), }, nil } diff --git a/services/agents-api/internal/runtimehistory/service_test.go b/services/agents-api/internal/runtimehistory/service_test.go index 4e100dbb7..63ba36b8e 100644 --- a/services/agents-api/internal/runtimehistory/service_test.go +++ b/services/agents-api/internal/runtimehistory/service_test.go @@ -47,7 +47,7 @@ func capabilities() Capabilities { MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, - Metrics: []Metric{MetricCPU, MetricMemory}, + Metrics: []Metric{MetricCPU, MetricMemory, MetricTokens}, } } @@ -59,7 +59,9 @@ func TestServiceAuthorizesAndBoundsBackendQuery(t *testing.T) { memory, limit := uint64(1024), uint64(2048) scope := Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID} reader := &fakeReader{capabilities: capabilities()} - reader.result = Result{GeneratedAt: now, RetainedFrom: &start, Series: []Series{{ + reader.result = Result{GeneratedAt: now, RetainedFrom: &start, TokenUsage: []TokenUsagePoint{{ + Start: start, End: start.Add(30 * time.Second), SampledAt: start.Add(20 * time.Second), InputTokens: 120, OutputTokens: 30, + }}, Series: []Series{{ Scope: scope, AllocationID: allocationID, StartedAt: startedAt, ProviderType: "docker", Points: []Point{{ Start: start, End: start.Add(30 * time.Second), FirstObservedAt: start.Add(time.Second), LastObservedAt: start.Add(20 * time.Second), @@ -83,6 +85,10 @@ func TestServiceAuthorizesAndBoundsBackendQuery(t *testing.T) { if response.Series[0].Points[0].CPUUtilizationRatio == nil || *response.Series[0].Points[0].CPUUtilizationRatio != .25 { t.Fatal("response aliases backend-owned metric memory") } + reader.result.TokenUsage[0].InputTokens = 999 + if response.TokenUsage[0].InputTokens != 120 { + t.Fatal("response aliases backend-owned token usage memory") + } } func TestServiceNeverQueriesBeforeOwnershipResolution(t *testing.T) { @@ -244,6 +250,16 @@ func TestServiceRejectsMalformedBackendResults(t *testing.T) { point.FirstObservedAt, point.LastObservedAt = point.Start, point.Start.Add(time.Second) result.GeneratedAt = point.Start }, + "token sample outside bucket": func(result *Result) { + result.TokenUsage = []TokenUsagePoint{{ + Start: start, End: start.Add(time.Minute), SampledAt: start.Add(time.Minute), InputTokens: 1, OutputTokens: 1, + }} + }, + "unsafe token count": func(result *Result) { + result.TokenUsage = []TokenUsagePoint{{ + Start: start, End: start.Add(time.Minute), SampledAt: start, InputTokens: maxSafeInteger + 1, + }} + }, } { t.Run(name, func(t *testing.T) { result := Result{GeneratedAt: now, Series: []Series{base}} diff --git a/services/agents-api/internal/runtimehistory/types.go b/services/agents-api/internal/runtimehistory/types.go index 55b83caa5..1f5be4616 100644 --- a/services/agents-api/internal/runtimehistory/types.go +++ b/services/agents-api/internal/runtimehistory/types.go @@ -22,6 +22,7 @@ type Metric string const ( MetricCPU Metric = "cpu" MetricMemory Metric = "memory" + MetricTokens Metric = "tokens" ) var providerTypePattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) @@ -95,12 +96,12 @@ func (c Capabilities) Validate() error { if c.MaximumSeries < 1 || c.MaximumSeries > 1_000 || c.MaximumTotalPoints < c.MaximumPoints || c.MaximumTotalPoints > 100_000 { return errors.New("invalid Runtime history result limits") } - if len(c.Metrics) == 0 || len(c.Metrics) > 2 { + if len(c.Metrics) == 0 || len(c.Metrics) > 3 { return errors.New("invalid Runtime history metrics") } seen := map[Metric]bool{} for _, metric := range c.Metrics { - if metric != MetricCPU && metric != MetricMemory || seen[metric] { + if metric != MetricCPU && metric != MetricMemory && metric != MetricTokens || seen[metric] { return errors.New("invalid Runtime history metrics") } seen[metric] = true @@ -143,6 +144,15 @@ type Result struct { RetainedFrom *time.Time Coverage []CoveragePoint Series []Series + TokenUsage []TokenUsagePoint +} + +// TokenUsagePoint is the final cumulative canonical Session Usage snapshot in +// one bucket. Throughput is derived from ordered adjacent points by clients. +type TokenUsagePoint struct { + Start, End time.Time + SampledAt time.Time + InputTokens, OutputTokens uint64 } type CoveragePoint struct { @@ -183,6 +193,7 @@ type Response struct { RetainedFrom *time.Time Coverage []CoveragePoint Series []Series + TokenUsage []TokenUsagePoint } func (r Response) Validate(now time.Time) error { @@ -200,7 +211,7 @@ func (r Response) Validate(now time.Time) error { Scope: r.Scope, Start: r.Requested.Start, End: r.Requested.End, Step: r.Resolution, Retention: r.Retention, MaxPoints: r.Requested.MaxPoints, MaximumSeries: r.MaximumSeries, MaximumTotalPoints: r.MaximumTotalPoints, } - if err := validateResult(query, Result{GeneratedAt: r.GeneratedAt, RetainedFrom: r.RetainedFrom, Coverage: r.Coverage, Series: r.Series}, now); err != nil { + if err := validateResult(query, Result{GeneratedAt: r.GeneratedAt, RetainedFrom: r.RetainedFrom, Coverage: r.Coverage, Series: r.Series, TokenUsage: r.TokenUsage}, now); err != nil { return ErrInvalidResult } return nil @@ -280,6 +291,23 @@ func validateResult(query Query, result Result, now time.Time) error { } } } + if len(result.TokenUsage) > query.MaxPoints { + return errors.New("Runtime history result exceeds token point limit") + } + totalPoints += len(result.TokenUsage) + if totalPoints > query.MaximumTotalPoints { + return errors.New("Runtime history result exceeds point limit") + } + for index, point := range result.TokenUsage { + if !validPublicBoundary(point.Start) || !validPublicBoundary(point.End) || point.Start.Before(query.Start) || !point.End.After(point.Start) || point.End.After(query.End) || point.End.Sub(point.Start) > query.Step || + !validPublicTime(point.SampledAt) || point.SampledAt.Before(point.Start) || !point.SampledAt.Before(point.End) || point.SampledAt.After(result.GeneratedAt) || point.SampledAt.Before(retainedStart) || + point.InputTokens > maxSafeInteger || point.OutputTokens > maxSafeInteger { + return errors.New("invalid Runtime history token usage point") + } + if index > 0 && result.TokenUsage[index-1].End.After(point.Start) { + return errors.New("Runtime history token usage points overlap or are out of order") + } + } return nil } diff --git a/services/agents-api/internal/runtimeobs/exporter.go b/services/agents-api/internal/runtimeobs/exporter.go index b3aa54267..b224a1207 100644 --- a/services/agents-api/internal/runtimeobs/exporter.go +++ b/services/agents-api/internal/runtimeobs/exporter.go @@ -31,6 +31,7 @@ type ExportRecord struct { ResolvedAt time.Time SourceDuration time.Duration Sample *Sample + TokenUsage *TokenUsage } // Exporter persists or forwards sanitized Runtime observation records. Core @@ -166,9 +167,18 @@ func exportRecord(observation Observation, collectionSource CollectionSource) Ex ResolvedAt: observation.ResolvedAt, SourceDuration: observation.SourceDuration, Sample: cloneSample(observation.Sample), + TokenUsage: cloneTokenUsage(observation.Target.TokenUsage), } } +func cloneTokenUsage(usage *TokenUsage) *TokenUsage { + if usage == nil { + return nil + } + cloned := *usage + return &cloned +} + func cloneSample(sample *Sample) *Sample { if sample == nil { return nil diff --git a/services/agents-api/internal/runtimeobs/identity.go b/services/agents-api/internal/runtimeobs/identity.go index f0ccdd14b..5cb149040 100644 --- a/services/agents-api/internal/runtimeobs/identity.go +++ b/services/agents-api/internal/runtimeobs/identity.go @@ -28,4 +28,14 @@ type Target struct { TenantID, SessionID, EnvironmentID string Mode Mode Instance Instance + // TokenUsage is the latest canonical Core Session Usage snapshot. It is + // Session-scoped and deliberately independent of the provider Runtime. + TokenUsage *TokenUsage +} + +// TokenUsage contains cumulative measured Session counters. Missing native +// measurements remain nil at the Target level and are never treated as zero. +type TokenUsage struct { + InputTokens uint64 + OutputTokens uint64 } diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go index 1e173c10e..3bcdcea55 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter.go @@ -29,6 +29,8 @@ const ( MemoryLimitName = "agents.runtime.memory.limit" SampleName = "agents.runtime.sample" SampleDurationName = "agents.runtime.sample.duration" + TokenInputName = "agents.session.tokens.input" + TokenOutputName = "agents.session.tokens.output" ) var safeLabelPattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,63}$`) @@ -109,6 +111,12 @@ func recordMetrics(record runtimeobs.ExportRecord) ([]metricdata.Metrics, error) if record.SourceDuration > 0 { result = append(result, durationMetric(record, attributes)) } + if record.TokenUsage != nil { + result = append(result, + integerGaugeMetric(TokenInputName, "Cumulative measured Session input tokens", "{token}", int64(record.TokenUsage.InputTokens), record.ResolvedAt, attributes), + integerGaugeMetric(TokenOutputName, "Cumulative measured Session output tokens", "{token}", int64(record.TokenUsage.OutputTokens), record.ResolvedAt, attributes), + ) + } if record.Sample == nil { return result, nil } @@ -176,6 +184,10 @@ func validateRecord(record runtimeobs.ExportRecord) error { if record.SourceDuration > 0 && record.ResolvedAt.Add(-record.SourceDuration).Unix() < 0 { return errors.New("invalid Runtime history source start time") } + const maxSafeInteger = uint64(1<<53 - 1) + if record.TokenUsage != nil && (record.TokenUsage.InputTokens > maxSafeInteger || record.TokenUsage.OutputTokens > maxSafeInteger) { + return errors.New("invalid Runtime history token usage") + } if record.Sample == nil { return nil } @@ -280,5 +292,11 @@ func gaugeMetric(name, description, unit string, value float64, observedAt time. }} } +func integerGaugeMetric(name, description, unit string, value int64, observedAt time.Time, attributes attribute.Set) metricdata.Metrics { + return metricdata.Metrics{Name: name, Description: description, Unit: unit, Data: metricdata.Gauge[int64]{ + DataPoints: []metricdata.DataPoint[int64]{{Attributes: attributes, Time: observedAt, Value: value}}, + }} +} + var _ sdkmetric.Exporter = (*otlpmetrichttp.Exporter)(nil) var _ runtimeobs.Exporter = (*Exporter)(nil) diff --git a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go index 1f59f010f..0fbfeb95b 100644 --- a/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go +++ b/services/agents-api/internal/runtimeobs/otlpexporter/exporter_test.go @@ -118,6 +118,29 @@ func TestExporterPreservesUnavailableWithoutInventingResourceValues(t *testing.T } } +func TestExporterEmitsCanonicalSessionTokenGaugesWithoutProviderValues(t *testing.T) { + client := &captureClient{} + exporter := newWithClient(client) + record := runtimeobs.ExportRecord{ + TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", + Mode: runtimeobs.ModeManaged, Status: runtimeobs.StatusUnavailable, Reason: "sample_timeout", + CollectionSource: runtimeobs.CollectionSourcePeriodic, + ResolvedAt: time.Date(2026, 9, 23, 3, 0, 0, 0, time.UTC), + TokenUsage: &runtimeobs.TokenUsage{InputTokens: 120, OutputTokens: 30}, + } + if err := exporter.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + metrics := client.metrics.ScopeMetrics[0].Metrics + if len(metrics) != 3 || metrics[0].Name != SampleName || metrics[1].Name != TokenInputName || metrics[2].Name != TokenOutputName { + t.Fatalf("unexpected token metrics: %#v", metrics) + } + input, ok := metrics[1].Data.(metricdata.Gauge[int64]) + if !ok || len(input.DataPoints) != 1 || input.DataPoints[0].Value != 120 { + t.Fatalf("unexpected input token gauge: %#v", metrics[1].Data) + } +} + func TestExporterPreservesObservedCoverageWithoutUnfencedResourceValues(t *testing.T) { client := &captureClient{} exporter := newWithClient(client) diff --git a/services/agents-api/internal/runtimeobs/storeresolver/resolver.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go index 79a988a49..d956532ae 100644 --- a/services/agents-api/internal/runtimeobs/storeresolver/resolver.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver.go @@ -1,10 +1,12 @@ package storeresolver import ( + "bytes" "context" "encoding/json" "errors" "fmt" + "io" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" @@ -39,6 +41,11 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (run return runtimeobs.Target{}, errors.New("invalid stored Runtime environment configuration") } target := runtimeobs.Target{TenantID: session.TenantID, SessionID: session.ID, Mode: runtimeobs.Mode(configuration.Environment.Type)} + usage, err := decodeTokenUsage(session.Usage) + if err != nil { + return runtimeobs.Target{}, err + } + target.TokenUsage = usage switch target.Mode { case runtimeobs.ModeNone: if session.Environment != nil { @@ -83,6 +90,54 @@ func (r *Resolver) Resolve(ctx context.Context, tenantID, sessionID string) (run } } +type storedTokenUsage struct { + InputTokens *int64 `json:"input_tokens"` + InputTokensDetails *storedInputTokenDetails `json:"input_tokens_details"` + OutputTokens *int64 `json:"output_tokens"` + OutputTokensDetails *storedOutputTokenDetails `json:"output_tokens_details"` + TotalTokens *int64 `json:"total_tokens"` +} + +type storedInputTokenDetails struct { + CachedTokens *int64 `json:"cached_tokens"` +} + +type storedOutputTokenDetails struct { + ReasoningTokens *int64 `json:"reasoning_tokens"` +} + +func decodeTokenUsage(raw json.RawMessage) (*runtimeobs.TokenUsage, error) { + trimmed := bytes.TrimSpace(raw) + if len(trimmed) == 0 || bytes.Equal(trimmed, []byte("null")) { + return nil, nil + } + var usage storedTokenUsage + decoder := json.NewDecoder(bytes.NewReader(trimmed)) + decoder.DisallowUnknownFields() + if err := decoder.Decode(&usage); err != nil { + return nil, errors.New("invalid stored Session token usage") + } + if err := decoder.Decode(&struct{}{}); !errors.Is(err, io.EOF) { + return nil, errors.New("invalid stored Session token usage") + } + if usage.InputTokens == nil || usage.OutputTokens == nil || usage.TotalTokens == nil || + usage.InputTokensDetails == nil || usage.InputTokensDetails.CachedTokens == nil || + usage.OutputTokensDetails == nil || usage.OutputTokensDetails.ReasoningTokens == nil { + return nil, errors.New("invalid stored Session token usage") + } + const maxSafeInteger = int64(1<<53 - 1) + values := []int64{*usage.InputTokens, *usage.OutputTokens, *usage.TotalTokens, *usage.InputTokensDetails.CachedTokens, *usage.OutputTokensDetails.ReasoningTokens} + for _, value := range values { + if value < 0 || value > maxSafeInteger { + return nil, errors.New("invalid stored Session token usage") + } + } + if *usage.InputTokens+*usage.OutputTokens != *usage.TotalTokens { + return nil, errors.New("invalid stored Session token usage") + } + return &runtimeobs.TokenUsage{InputTokens: uint64(*usage.InputTokens), OutputTokens: uint64(*usage.OutputTokens)}, nil +} + func (r *Resolver) ListRuntimeObservationSessions(ctx context.Context, after string, limit int) (runtimeobs.SessionPage, error) { page, err := r.store.ListRuntimeObservationSessions(ctx, after, limit) if err != nil { diff --git a/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go index 9b86e304a..855749093 100644 --- a/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go +++ b/services/agents-api/internal/runtimeobs/storeresolver/resolver_test.go @@ -44,7 +44,7 @@ func TestResolverListsOnlyProviderNeutralSessionIdentity(t *testing.T) { func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { r, err := NewResolver(resolverStore{ - session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, + session: store.Session{ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"openai_hosted"}}`), Usage: []byte(`{"input_tokens":120,"input_tokens_details":{"cached_tokens":20},"output_tokens":30,"output_tokens_details":{"reasoning_tokens":10},"total_tokens":150}`), Environment: &store.Environment{ID: "environment", TenantID: "tenant", SessionID: "session"}}, allocation: store.RuntimeAllocation{ ID: "allocation", TenantID: "tenant", SessionID: "session", EnvironmentID: "environment", ProviderKey: "provider", DeviceID: "device", ComputePhase: "running", ComputeState: []byte(`{"current":{"name":"sandbox"}}`), @@ -66,6 +66,42 @@ func TestResolverBindsManagedSessionEnvironmentAndAllocation(t *testing.T) { if target.Instance.ComputePhase != "running" { t.Fatalf("compute phase was not retained: %s", target.Instance.ComputePhase) } + if target.TokenUsage == nil || target.TokenUsage.InputTokens != 120 || target.TokenUsage.OutputTokens != 30 { + t.Fatalf("canonical Session usage was not retained: %+v", target.TokenUsage) + } +} + +func TestResolverRejectsInvalidCanonicalSessionUsage(t *testing.T) { + for _, usage := range []string{ + `{"input_tokens":2,"input_tokens_details":{"cached_tokens":0},"output_tokens":3,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":4}`, + `{}`, + `{"total_tokens":0}`, + `{"input_tokens":0,"input_tokens_details":{"cached_tokens":-1},"output_tokens":0,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":0}`, + `{"input_tokens":0,"input_tokens_details":{"cached_tokens":0},"output_tokens":0,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":0,"unknown":0}`, + } { + resolver, err := NewResolver(resolverStore{session: store.Session{ + ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"none"}}`), Usage: []byte(usage), + }}) + if err != nil { + t.Fatal(err) + } + if _, err := resolver.Resolve(t.Context(), "tenant", "session"); err == nil { + t.Fatalf("invalid Session token usage was accepted: %s", usage) + } + } +} + +func TestResolverKeepsNullCanonicalSessionUsageAbsent(t *testing.T) { + resolver, err := NewResolver(resolverStore{session: store.Session{ + ID: "session", TenantID: "tenant", Configuration: []byte(`{"environment":{"type":"none"}}`), Usage: []byte(" \n null \t"), + }}) + if err != nil { + t.Fatal(err) + } + target, err := resolver.Resolve(t.Context(), "tenant", "session") + if err != nil || target.TokenUsage != nil { + t.Fatalf("null Session usage was not kept absent: %+v %v", target.TokenUsage, err) + } } func TestResolverKeepsUnsupportedModesDistinct(t *testing.T) { diff --git a/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql b/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql new file mode 100644 index 000000000..f2e2d5f9c --- /dev/null +++ b/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql @@ -0,0 +1,31 @@ +-- Session-scoped cumulative token counters. Apply after 001_runtime_history.sql. +-- A separate materialized view keeps usage independent of provider Runtime +-- incarnation fields while reusing the bounded runtime_history_metrics table. + +CREATE MATERIALIZED VIEW IF NOT EXISTS runtime_history_session_usage_gauge_mv +TO runtime_history_metrics AS +SELECT + fromUnixTimestamp64Nano(toInt64(Attributes['agents.runtime.resolved_at_unix_nano'])) AS timestamp, + Attributes['agents.tenant.id'] AS tenant_id, + Attributes['agents.session.id'] AS session_id, + Attributes['agents.environment.id'] AS environment_id, + Attributes['agents.runtime.allocation.id'] AS allocation_id, + toInt64OrNull(Attributes['agents.runtime.compute.started_at_unix_nano']) AS started_at_unix_nano, + Attributes['agents.runtime.provider.type'] AS provider_type, + Attributes['agents.runtime.status'] AS status, + Attributes['agents.runtime.collection.source'] AS collection_source, + toInt64(Attributes['agents.runtime.resolved_at_unix_nano']) AS resolved_at_unix_nano, + toInt64OrNull(Attributes['agents.runtime.observed_at_unix_nano']) AS observed_at_unix_nano, + MetricName AS metric_name, + toFloat64(Value) AS value +FROM metrics_gauge +WHERE MetricName IN +( + 'agents.session.tokens.input', + 'agents.session.tokens.output' +) +AND Attributes['agents.tenant.id'] != '' +AND Attributes['agents.session.id'] != '' +AND Attributes['agents.environment.id'] != '' +AND Attributes['agents.runtime.collection.source'] IN ('on_read', 'periodic') +AND toInt64OrNull(Attributes['agents.runtime.resolved_at_unix_nano']) IS NOT NULL; diff --git a/services/agents-api/runtime-history/clickhouse/README.md b/services/agents-api/runtime-history/clickhouse/README.md index c8919bfa8..6eabe4ec5 100644 --- a/services/agents-api/runtime-history/clickhouse/README.md +++ b/services/agents-api/runtime-history/clickhouse/README.md @@ -30,8 +30,11 @@ or lifecycle decisions. exporter creates the generic metric tables. Keep the OTLP receiver private or authenticate it at the network/proxy boundary. 2. After `metrics_gauge` and `metrics_sum` exist, apply - [`001_runtime_history.sql`](001_runtime_history.sql) to the same database. - The materialized views retain only the five values needed by the history API; + [`001_runtime_history.sql`](001_runtime_history.sql), then + [`002_session_token_usage.sql`](002_session_token_usage.sql), to the same + database. + The materialized views retain only the five Runtime values and two canonical + Session token counters needed by the history API; provider receipts, native container identities, paths, and raw errors never enter the projection. Grant the Collector identity the minimum source-column reads required when From 0dafa92378a725b509ce047efb4bdde89a6b916b Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 17:13:22 +0800 Subject: [PATCH 36/51] Fix microsandbox metrics incarnation timing precision --- .../sandbox/microsandbox/resources.go | 4 +- .../sandbox/microsandbox/resources_test.go | 42 ++++++ .../internal/sandbox/microsandbox/types.go | 3 +- .../tools/microsandbox-provider/README.md | 28 +++- .../tools/microsandbox-provider/backend.go | 61 --------- .../tools/microsandbox-provider/main.go | 3 +- .../tools/microsandbox-provider/metrics.go | 106 +++++++++++++++ .../microsandbox-provider/metrics_test.go | 126 ++++++++++-------- 8 files changed, 250 insertions(+), 123 deletions(-) create mode 100644 services/agents-api/tools/microsandbox-provider/metrics.go diff --git a/services/agents-api/internal/sandbox/microsandbox/resources.go b/services/agents-api/internal/sandbox/microsandbox/resources.go index 4408dddb5..02821c98c 100644 --- a/services/agents-api/internal/sandbox/microsandbox/resources.go +++ b/services/agents-api/internal/sandbox/microsandbox/resources.go @@ -14,7 +14,7 @@ var _ runtimeobs.Source = (*Provider)(nil) func (*Provider) ObservationProviderType() string { return "microsandbox" } -// Observe reads one point-in-time SDK metrics snapshot through the existing +// Observe reads one point-in-time native metrics snapshot through the existing // one-shot helper. The persisted compute receipt selects the exact generation; // browser input and provider display names never select a sandbox. func (p *Provider) Observe(ctx context.Context, target runtimeobs.Target) (runtimeobs.Sample, error) { @@ -59,6 +59,8 @@ func sampleFromMetrics(config Config, metrics Metrics) (runtimeobs.Sample, error if metrics.ObservedAt.IsZero() || metrics.Uptime < 0 || metrics.MemoryLimitBytes == 0 || config.CPUs == 0 { return runtimeobs.Sample{}, ErrUnconfirmed } + // Both values come from one native registry sample at millisecond precision. + // Helper wall time and the SDK's rounded uptime cannot establish this fence. startedAt := metrics.ObservedAt.Add(-metrics.Uptime) if startedAt.Unix() < 0 || startedAt.After(metrics.ObservedAt) { return runtimeobs.Sample{}, ErrUnconfirmed diff --git a/services/agents-api/internal/sandbox/microsandbox/resources_test.go b/services/agents-api/internal/sandbox/microsandbox/resources_test.go index 02bb3d2a8..819e251d1 100644 --- a/services/agents-api/internal/sandbox/microsandbox/resources_test.go +++ b/services/agents-api/internal/sandbox/microsandbox/resources_test.go @@ -127,3 +127,45 @@ func TestSampleFromMetricsRejectsInvalidTimingAndLimit(t *testing.T) { } } } + +func TestObserveKeepsIncarnationAcrossPollsAndCoreRestart(t *testing.T) { + config, reference := testConfig(), testRef() + compute := Compute{Name: Name(config, reference, 0), ID: "local:first"} + startedAt := time.Date(2026, 9, 22, 12, 0, 0, 122000000, time.UTC) + var previous runtimeobs.Sample + for index, uptime := range []time.Duration{300001 * time.Millisecond, 301877 * time.Millisecond, time.Millisecond} { + currentStart := startedAt + cpu := uint64(2_500_000_000 + index*1_000_000_000) + if index == 2 { + compute = Compute{Name: Name(config, reference, 1), ID: "local:restored", Generation: 1, RestoredFrom: func() *SnapshotIdentity { s := testSnapshot(); return &s }()} + currentStart = startedAt.Add(5 * time.Minute) + cpu = 1_000_000 + } + // Each poll uses a fresh Core provider, with no process-local time cache. + provider, err := NewWithCaller(config, callerFunc(func(_ context.Context, request Request) (Response, error) { + if request.Compute.ID != compute.ID || request.Compute.Generation != compute.Generation { + t.Fatalf("wrong compute: %+v", request.Compute) + } + return Response{Version: ProtocolVersion, Metrics: &Metrics{ObservedAt: currentStart.Add(uptime), Uptime: uptime, VCPUTimeNs: cpu, MemoryLimitBytes: 8192}}, nil + })) + if err != nil { + t.Fatal(err) + } + sample, err := provider.Observe(deadline(t), observationTarget(t, compute)) + if err != nil { + t.Fatal(err) + } + if sample.StartedAt == nil || !sample.StartedAt.Equal(currentStart) { + t.Fatalf("wrong start: %+v", sample) + } + if index == 1 { + if !sample.StartedAt.Equal(*previous.StartedAt) || *sample.CPUUsageSecondsTotal-*previous.CPUUsageSecondsTotal != 1 { + t.Fatal("repeated polls lost the CPU delta") + } + } + if index == 2 && sample.StartedAt.Equal(*previous.StartedAt) { + t.Fatal("restored compute reused the source incarnation") + } + previous = sample + } +} diff --git a/services/agents-api/internal/sandbox/microsandbox/types.go b/services/agents-api/internal/sandbox/microsandbox/types.go index 414aef67a..671590d32 100644 --- a/services/agents-api/internal/sandbox/microsandbox/types.go +++ b/services/agents-api/internal/sandbox/microsandbox/types.go @@ -71,7 +71,8 @@ type Response struct { } // Metrics is the bounded provider-helper projection used by Core observability. -// It intentionally excludes instantaneous CPU percent and provider-native names. +// ObservedAt and Uptime preserve the same native registry sample at millisecond +// precision. It excludes instantaneous CPU percent and provider-native names. type Metrics struct { ObservedAt time.Time Uptime time.Duration diff --git a/services/agents-api/tools/microsandbox-provider/README.md b/services/agents-api/tools/microsandbox-provider/README.md index 7d967df68..efe8ba243 100644 --- a/services/agents-api/tools/microsandbox-provider/README.md +++ b/services/agents-api/tools/microsandbox-provider/README.md @@ -128,13 +128,27 @@ successful stdin completion and an explicit guest exit before returning a result Timeouts, output overflow and missing receipts return ErrCommandUnconfirmed. Closing an SDK exec handle alone does not prove the guest process exited. -The read-only metrics operation verifies the same exact allocation and compute -incarnation before calling the pinned SDK's point-in-time `SandboxHandle.Metrics`. -It returns only observation time, uptime, cumulative vCPU time and guest memory -usage/limit to Core. It does not connect to the guest, renew activity, resume paused -compute or mutate lifecycle state. SDK metrics-disabled and no-current-sample errors -are reduced to one safe unavailable code; raw diagnostics never cross the helper -boundary. +The read-only metrics operation holds the allocation lock and verifies the exact +compute ID through the SDK before and after `msb metrics NAME --format json`. +The CLI is the same checksum/version-pinned runtime binary. Its registry report +preserves the native sample timestamp and fractional-second uptime; parsing both +at native millisecond precision reconstructs the real run start consistently +across polls and Core restarts, while new runs have a new start. The Go SDK v0.7.2 +projection drops that timestamp and truncates uptime to whole seconds, so it +cannot supply this fence. Sandbox creation time is not a run start time. + +The helper returns only the native observation time, exact uptime, cumulative +vCPU time and guest memory usage/limit. It rejects stale/exited reports and +malformed or missing fields. CLI output is bounded; command failures expose only +an unavailable code, never native diagnostics. Observation does not connect to +the guest, renew activity, resume paused compute or mutate lifecycle state. + +The pinned source evidence is tag `v0.7.2`, commit +`1c59b8dbf0ad47dda2f807c0214b529aceb81c74`: `crates/metrics/lib/registry.rs` +reads `sampled_at_unix_ms` and `started_at_unix_ms` coherently and subtracts them +for uptime; `crates/cli/lib/commands/metrics.rs` serializes the timestamp and +`uptime.as_secs_f64()`. The private native fields and identifiers do not cross +the helper boundary. ## Acceptance boundary diff --git a/services/agents-api/tools/microsandbox-provider/backend.go b/services/agents-api/tools/microsandbox-provider/backend.go index 5168b1e78..412d92570 100644 --- a/services/agents-api/tools/microsandbox-provider/backend.go +++ b/services/agents-api/tools/microsandbox-provider/backend.go @@ -5,7 +5,6 @@ package main import ( "context" "errors" - "time" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" wire "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox/microsandbox" @@ -16,11 +15,6 @@ const bootstrapLabel = "io.parsar.bootstrap" type backend struct{ q wire.Request } -type liveMetricsSource interface { - Metrics(context.Context) (*sdk.Metrics, error) - Detach(context.Context) error -} - func (b backend) run(ctx context.Context) (wire.Response, error) { switch b.q.Operation { case "create": @@ -70,61 +64,6 @@ func (b backend) run(ctx context.Context) (wire.Response, error) { return wire.Response{}, sandbox.ErrInvalid } -func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, error) { - metricsCtx, cancel := context.WithDeadline(ctx, b.q.Deadline) - defer cancel() - h, state, err := b.inspect(metricsCtx, c) - if err != nil { - return nil, err - } - if state.Status != string(sdk.SandboxStatusRunning) && state.Status != "draining" { - return nil, sandbox.ErrNotFound - } - // v0.7.2's name-based handle metrics omit cumulative vCPU time. Connect to - // the already-running, identity-qualified instance so CPU usage remains a - // monotonic counter that Core can safely derive rates from. Connect never - // starts stopped compute, and Detach leaves the runtime lifecycle unchanged. - metrics, err := observeConnectedMetrics(metricsCtx, func(ctx context.Context) (liveMetricsSource, error) { - return h.Connect(ctx) - }) - if err != nil { - return nil, err - } - projected := projectMetrics(metrics, time.Now().UTC()) - if projected == nil { - return nil, wire.ErrUnconfirmed - } - return projected, nil -} - -func observeConnectedMetrics(ctx context.Context, connect func(context.Context) (liveMetricsSource, error)) (*sdk.Metrics, error) { - live, err := connect(ctx) - if err != nil { - return nil, err - } - metrics, metricsErr := live.Metrics(ctx) - detachCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - detachErr := live.Detach(detachCtx) - cancel() - if metricsErr != nil { - return nil, metricsErr - } - if detachErr != nil { - return nil, detachErr - } - return metrics, nil -} - -func projectMetrics(metrics *sdk.Metrics, observedAt time.Time) *wire.Metrics { - if metrics == nil { - return nil - } - return &wire.Metrics{ - ObservedAt: observedAt, Uptime: metrics.Uptime, - VCPUTimeNs: metrics.VCPUTimeNs, MemoryBytes: metrics.MemoryBytes, - MemoryLimitBytes: metrics.MemoryLimitBytes, - } -} func (b backend) inspect(ctx context.Context, c wire.Compute) (*sdk.SandboxHandle, wire.State, error) { state := wire.State{Compute: c} h, e := sdk.GetSandbox(ctx, c.Name) diff --git a/services/agents-api/tools/microsandbox-provider/main.go b/services/agents-api/tools/microsandbox-provider/main.go index d1276b411..a67f21ceb 100644 --- a/services/agents-api/tools/microsandbox-provider/main.go +++ b/services/agents-api/tools/microsandbox-provider/main.go @@ -19,6 +19,7 @@ import ( "syscall" "time" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" wire "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox/microsandbox" sdk "github.com/superradcompany/microsandbox/sdk/go" @@ -75,7 +76,7 @@ func code(err error) string { return "not_found" case errors.Is(err, sandbox.ErrCommandUnconfirmed): return "command_unconfirmed" - case sdk.IsKind(err, sdk.ErrMetricsDisabled), sdk.IsKind(err, sdk.ErrMetricsUnavailable): + case errors.Is(err, runtimeobs.ErrUnavailable), sdk.IsKind(err, sdk.ErrMetricsDisabled), sdk.IsKind(err, sdk.ErrMetricsUnavailable): return "metrics_unavailable" default: return "unconfirmed" diff --git a/services/agents-api/tools/microsandbox-provider/metrics.go b/services/agents-api/tools/microsandbox-provider/metrics.go new file mode 100644 index 000000000..907b2a109 --- /dev/null +++ b/services/agents-api/tools/microsandbox-provider/metrics.go @@ -0,0 +1,106 @@ +//go:build linux + +package main + +import ( + "bytes" + "context" + "encoding/json" + "io" + "os/exec" + "strconv" + "strings" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox" + wire "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox/microsandbox" +) + +func (b backend) metrics(ctx context.Context, c wire.Compute) (*wire.Metrics, error) { + ctx, cancel := context.WithDeadline(ctx, b.q.Deadline) + defer cancel() + _, state, err := b.inspect(ctx, c) + if err != nil { + return nil, err + } + if state.Status != "running" && state.Status != "draining" { + return nil, sandbox.ErrNotFound + } + // The allocation lock and exact SDK identity checks fence the name-based + // CLI read. The pinned CLI preserves the registry sample timestamp and + // millisecond uptime that the Go SDK's Metrics projection discards. + metrics, err := readMetrics(ctx, b.q.Config.RuntimePath, c.Name) + if err != nil { + return nil, err + } + _, current, err := b.inspect(ctx, state.Compute) + if err != nil { + return nil, err + } + if current.Status != "running" && current.Status != "draining" { + return nil, sandbox.ErrNotFound + } + return metrics, nil +} + +func readMetrics(ctx context.Context, runtimePath, name string) (*wire.Metrics, error) { + command := exec.CommandContext(ctx, runtimePath, "metrics", name, "--format", "json") + command.WaitDelay = time.Second + output := &metricsOutput{} + command.Stdout = output + command.Stderr = io.Discard + if err := command.Run(); err != nil { + if ctx.Err() != nil { + return nil, ctx.Err() + } + return nil, runtimeobs.ErrUnavailable + } + return projectMetrics(output.Bytes(), name) +} + +type metricsOutput struct{ bytes.Buffer } + +func (b *metricsOutput) Write(value []byte) (int, error) { + if len(value) > 64*1024-b.Len() { + return 0, io.ErrShortWrite + } + return b.Buffer.Write(value) +} + +func projectMetrics(data []byte, name string) (*wire.Metrics, error) { + var report struct { + Name string `json:"name"` + State string `json:"state"` + Timestamp time.Time `json:"timestamp"` + UptimeSeconds json.Number `json:"uptime_secs"` + VCPUTimeNs *uint64 `json:"vcpu_time_ns"` + MemoryBytes *uint64 `json:"memory_bytes"` + MemoryLimitBytes *uint64 `json:"memory_limit_bytes"` + } + if json.Unmarshal(data, &report) != nil || report.Name != name || report.Timestamp.IsZero() || + !report.Timestamp.Equal(report.Timestamp.Truncate(time.Millisecond)) || + report.VCPUTimeNs == nil || report.MemoryBytes == nil || report.MemoryLimitBytes == nil || *report.MemoryLimitBytes == 0 { + return nil, wire.ErrUnconfirmed + } + if report.State != "running" { + return nil, runtimeobs.ErrUnavailable + } + // Registry timestamps are integer milliseconds. Parse the CLI decimal + // exactly, without float-to-nanosecond rounding changing the start fence. + seconds, fraction, _ := strings.Cut(string(report.UptimeSeconds), ".") + if seconds == "" || len(fraction) > 3 || strings.ContainsAny(seconds+fraction, "-+eE/") { + return nil, wire.ErrUnconfirmed + } + milliseconds, err := strconv.ParseInt(seconds+fraction+strings.Repeat("0", 3-len(fraction)), 10, 64) + if err != nil || milliseconds < 0 { + return nil, wire.ErrUnconfirmed + } + if milliseconds > int64((1<<63-1)/time.Millisecond) || report.Timestamp.UnixMilli() < milliseconds { + return nil, wire.ErrUnconfirmed + } + return &wire.Metrics{ + ObservedAt: report.Timestamp, Uptime: time.Duration(milliseconds) * time.Millisecond, + VCPUTimeNs: *report.VCPUTimeNs, MemoryBytes: *report.MemoryBytes, MemoryLimitBytes: *report.MemoryLimitBytes, + }, nil +} diff --git a/services/agents-api/tools/microsandbox-provider/metrics_test.go b/services/agents-api/tools/microsandbox-provider/metrics_test.go index 95dc25644..313185937 100644 --- a/services/agents-api/tools/microsandbox-provider/metrics_test.go +++ b/services/agents-api/tools/microsandbox-provider/metrics_test.go @@ -5,71 +5,93 @@ package main import ( "context" "errors" + "fmt" + "os" + "path/filepath" + "strings" "testing" "time" - sdk "github.com/superradcompany/microsandbox/sdk/go" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + wire "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/sandbox/microsandbox" ) -type fakeLiveMetrics struct { - metrics *sdk.Metrics - metricsErr error - detachErr error - metricsCalled bool - detachCalled bool - detachBounded bool +func nativeMetrics(timestamp, uptime string, cpu uint64) []byte { + return []byte(fmt.Sprintf(`{"name":"owned","state":"running","timestamp":%q,"uptime_secs":%s,"vcpu_time_ns":%d,"memory_bytes":4096,"memory_limit_bytes":8192,"cpu_percent":72.5}`, timestamp, uptime, cpu)) } -func (f *fakeLiveMetrics) Metrics(context.Context) (*sdk.Metrics, error) { - f.metricsCalled = true - return f.metrics, f.metricsErr -} - -func (f *fakeLiveMetrics) Detach(ctx context.Context) error { - f.detachCalled = true - _, f.detachBounded = ctx.Deadline() - return f.detachErr -} - -func TestObserveConnectedMetricsSamplesLiveInstanceAndDetaches(t *testing.T) { - want := &sdk.Metrics{VCPUTimeNs: 123} - live := &fakeLiveMetrics{metrics: want} - connected := false - got, err := observeConnectedMetrics(context.Background(), func(context.Context) (liveMetricsSource, error) { - connected = true - return live, nil - }) - if err != nil || got != want || !connected || !live.metricsCalled || !live.detachCalled || !live.detachBounded { - t.Fatalf("unexpected observation: metrics=%+v err=%v connected=%t live=%+v", got, err, connected, live) +func TestProjectMetricsRetainsNativeStartAcrossPollsAndReplacement(t *testing.T) { + var start time.Time + for index, input := range []struct { + timestamp, uptime string + cpu uint64 + }{ + {"2026-09-22T12:00:00.123Z", "300.001", 2_500_000_000}, + {"2026-09-22T12:00:01.999Z", "301.877", 3_500_000_000}, + {"2026-09-22T12:00:04.124Z", "0.001", 1_000_000}, + } { + sample, err := projectMetrics(nativeMetrics(input.timestamp, input.uptime, input.cpu), "owned") + if err != nil { + t.Fatal(err) + } + if sample.VCPUTimeNs != input.cpu || sample.MemoryBytes != 4096 || sample.MemoryLimitBytes != 8192 { + t.Fatalf("bad sample: %+v", sample) + } + actualStart := sample.ObservedAt.Add(-sample.Uptime) + if index == 0 { + start = actualStart + } + if (index < 2) != actualStart.Equal(start) { + t.Fatalf("poll %d start=%s first=%s", index, actualStart, start) + } } } -func TestObserveConnectedMetricsPropagatesDetachFailure(t *testing.T) { - want := errors.New("detach failed") - live := &fakeLiveMetrics{metrics: &sdk.Metrics{VCPUTimeNs: 123}, detachErr: want} - got, err := observeConnectedMetrics(context.Background(), func(context.Context) (liveMetricsSource, error) { - return live, nil - }) - if got != nil || !errors.Is(err, want) || !live.detachCalled || !live.detachBounded { - t.Fatalf("detach failure not propagated: metrics=%+v err=%v live=%+v", got, err, live) +func TestProjectMetricsRejectsUnqualifiedReports(t *testing.T) { + valid := string(nativeMetrics("2026-09-22T12:00:00.123Z", "300.001", 123)) + for _, input := range []string{ + `{}`, valid + `{}`, strings.Replace(valid, `"owned"`, `"foreign"`, 1), + strings.Replace(valid, `"timestamp":"2026-09-22T12:00:00.123Z"`, `"timestamp":"2026-09-22T12:00:00.123456Z"`, 1), + strings.Replace(valid, `"vcpu_time_ns":123,`, ``, 1), + strings.Replace(valid, `"memory_limit_bytes":8192`, `"memory_limit_bytes":0`, 1), + strings.Replace(valid, `"uptime_secs":300.001,`, ``, 1), + } { + if _, err := projectMetrics([]byte(input), "owned"); !errors.Is(err, wire.ErrUnconfirmed) { + t.Fatalf("accepted malformed report: %s, %v", input, err) + } } -} - -func TestProjectMetricsKeepsOnlyCumulativeCommonFields(t *testing.T) { - observed := time.Date(2026, 9, 22, 12, 0, 0, 0, time.UTC) - projected := projectMetrics(&sdk.Metrics{ - CPUPercent: 72.5, VCPUTimeNs: 2_500_000_000, - MemoryBytes: 4096, MemoryLimitBytes: 8192, - DiskReadBytes: 100, DiskWriteBytes: 200, NetRxBytes: 300, NetTxBytes: 400, - Uptime: 5 * time.Minute, - }, observed) - if projected.ObservedAt != observed || projected.Uptime != 5*time.Minute || projected.VCPUTimeNs != 2_500_000_000 || projected.MemoryBytes != 4096 || projected.MemoryLimitBytes != 8192 { - t.Fatalf("bad projection: %+v", projected) + for _, uptime := range []string{"-1", "0.0001", "1e100000000", "999999999999999999999999999", "99999999999", "null"} { + if _, err := projectMetrics(nativeMetrics("2026-09-22T12:00:00.123Z", uptime, 123), "owned"); !errors.Is(err, wire.ErrUnconfirmed) { + t.Fatalf("accepted uptime %s: %v", uptime, err) + } + } + for _, state := range []string{"stalled", "exited", ""} { + if _, err := projectMetrics([]byte(strings.Replace(valid, `"running"`, fmt.Sprintf("%q", state), 1)), "owned"); !errors.Is(err, runtimeobs.ErrUnavailable) { + t.Fatalf("accepted state %q: %v", state, err) + } } } -func TestProjectMetricsRejectsMissingSDKSample(t *testing.T) { - if projected := projectMetrics(nil, time.Now()); projected != nil { - t.Fatalf("missing SDK sample projected: %+v", projected) +func TestReadMetricsUsesBoundedReadOnlyNativeCommand(t *testing.T) { + binary := filepath.Join(t.TempDir(), "msb") + body := "#!/bin/sh\n[ \"$*\" = 'metrics owned --format json' ] || exit 2\nprintf '%s' '" + string(nativeMetrics("2026-09-22T12:00:00.123Z", "300.001", 123)) + "'\n" + if err := os.WriteFile(binary, []byte(body), 0700); err != nil { + t.Fatal(err) + } + ctx, cancel := context.WithTimeout(context.Background(), time.Second) + defer cancel() + if sample, err := readMetrics(ctx, binary, "owned"); err != nil || sample.VCPUTimeNs != 123 { + t.Fatalf("sample=%+v error=%v", sample, err) + } + if _, err := readMetrics(ctx, binary, "foreign"); !errors.Is(err, runtimeobs.ErrUnavailable) { + t.Fatalf("command failure leaked: %v", err) + } + cancel() + if _, err := readMetrics(ctx, binary, "owned"); !errors.Is(err, context.Canceled) { + t.Fatalf("deadline ignored: %v", err) + } + output := &metricsOutput{} + if _, err := output.Write(make([]byte, 64*1024+1)); err == nil || output.Len() != 0 { + t.Fatal("unbounded metrics output") } } From c1e669c55778868890131b184a8a753ecf19e872 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 17:15:16 +0800 Subject: [PATCH 37/51] Recover native millisecond precision from CLI decimal rounding --- .../tools/microsandbox-provider/metrics.go | 24 +++++++++++++++---- .../microsandbox-provider/metrics_test.go | 9 +++++++ 2 files changed, 28 insertions(+), 5 deletions(-) diff --git a/services/agents-api/tools/microsandbox-provider/metrics.go b/services/agents-api/tools/microsandbox-provider/metrics.go index 907b2a109..f0795959f 100644 --- a/services/agents-api/tools/microsandbox-provider/metrics.go +++ b/services/agents-api/tools/microsandbox-provider/metrics.go @@ -7,8 +7,8 @@ import ( "context" "encoding/json" "io" + "math/big" "os/exec" - "strconv" "strings" "time" @@ -88,14 +88,28 @@ func projectMetrics(data []byte, name string) (*wire.Metrics, error) { } // Registry timestamps are integer milliseconds. Parse the CLI decimal // exactly, without float-to-nanosecond rounding changing the start fence. - seconds, fraction, _ := strings.Cut(string(report.UptimeSeconds), ".") - if seconds == "" || len(fraction) > 3 || strings.ContainsAny(seconds+fraction, "-+eE/") { + decimal := string(report.UptimeSeconds) + if len(decimal) == 0 || len(decimal) > 32 || strings.ContainsAny(decimal, "-+eE/") { return nil, wire.ErrUnconfirmed } - milliseconds, err := strconv.ParseInt(seconds+fraction+strings.Repeat("0", 3-len(fraction)), 10, 64) - if err != nil || milliseconds < 0 { + uptime, ok := new(big.Rat).SetString(decimal) + if !ok { return nil, wire.ErrUnconfirmed } + uptime.Mul(uptime, big.NewRat(1000, 1)) + // as_secs_f64 can serialize 1.118 seconds as 1.1179999999999999. + // Recover the nearest native millisecond exactly. Within Go's duration + // range, floating serialization error stays below one microsecond. + rounded := new(big.Rat).Add(uptime, big.NewRat(1, 2)) + millisecondsValue := new(big.Int).Quo(rounded.Num(), rounded.Denom()) + if !millisecondsValue.IsInt64() { + return nil, wire.ErrUnconfirmed + } + deviation := new(big.Rat).Sub(uptime, new(big.Rat).SetInt(millisecondsValue)) + if deviation.Abs(deviation).Cmp(big.NewRat(1, 1000)) > 0 { + return nil, wire.ErrUnconfirmed + } + milliseconds := millisecondsValue.Int64() if milliseconds > int64((1<<63-1)/time.Millisecond) || report.Timestamp.UnixMilli() < milliseconds { return nil, wire.ErrUnconfirmed } diff --git a/services/agents-api/tools/microsandbox-provider/metrics_test.go b/services/agents-api/tools/microsandbox-provider/metrics_test.go index 313185937..8b361f08e 100644 --- a/services/agents-api/tools/microsandbox-provider/metrics_test.go +++ b/services/agents-api/tools/microsandbox-provider/metrics_test.go @@ -47,6 +47,15 @@ func TestProjectMetricsRetainsNativeStartAcrossPollsAndReplacement(t *testing.T) } } +func TestProjectMetricsRecoversNativeFloatingSerializationAtMillisecondPrecision(t *testing.T) { + for _, uptime := range []string{"1.118", "1.1179999999999999", "1.1180000000000001"} { + sample, err := projectMetrics(nativeMetrics("2026-09-22T12:00:00.123Z", uptime, 123), "owned") + if err != nil || sample.Uptime != 1118*time.Millisecond { + t.Fatalf("uptime %s: sample=%+v error=%v", uptime, sample, err) + } + } +} + func TestProjectMetricsRejectsUnqualifiedReports(t *testing.T) { valid := string(nativeMetrics("2026-09-22T12:00:00.123Z", "300.001", 123)) for _, input := range []string{ From e8efa8498990f967b9f918f4f651a74238b6cf70 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 17:15:45 +0800 Subject: [PATCH 38/51] Document native runtime metrics incarnation contract --- CONTRIBUTING.md | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a6c72d8d2..43e3e9ee7 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -299,7 +299,10 @@ may inform operators, but automatic suspension requires durable Core-owned activ state and must not use a monitoring backend as lifecycle authority. Managed Docker observes one non-streaming Inspect/Stats sample. Managed microsandbox observes the exact persisted compute generation through the existing one-shot helper -and pinned SDK Metrics call. Preserve cumulative CPU seconds, memory usage/limit and +and pinned native CLI metrics report, with SDK identity checks before and after +observation. Derive compute start from the same native sample timestamp and precise +uptime; never subtract rounded uptime from a new wall-clock timestamp. Preserve +cumulative CPU seconds, memory usage/limit and compute uptime semantics across both. Do not use microsandbox's instantaneous CPU percent, wake suspended compute, or expose provider-native identifiers to fill a common field. From 4c0bf007392d22f28e2c43b7d46aab5c850fd3d7 Mon Sep 17 00:00:00 2001 From: sam Date: Wed, 23 Sep 2026 19:37:11 +0800 Subject: [PATCH 39/51] Allow Session token metrics through Collector --- services/agents-api/runtime-history/clickhouse/README.md | 4 ++++ .../runtime-history/clickhouse/otel-collector.example.yaml | 1 + 2 files changed, 5 insertions(+) diff --git a/services/agents-api/runtime-history/clickhouse/README.md b/services/agents-api/runtime-history/clickhouse/README.md index 6eabe4ec5..106f44ccc 100644 --- a/services/agents-api/runtime-history/clickhouse/README.md +++ b/services/agents-api/runtime-history/clickhouse/README.md @@ -37,6 +37,10 @@ or lifecycle decisions. Session token counters needed by the history API; provider receipts, native container identities, paths, and raw errors never enter the projection. + Keep both the `agents.runtime.*` family and the exact + `agents.session.tokens.input` / `agents.session.tokens.output` gauges in the + Collector allow-list, as shown in `otel-collector.example.yaml`; otherwise + the Collector discards Session usage before ClickHouse receives it. Grant the Collector identity the minimum source-column reads required when those views run: diff --git a/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml b/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml index f21915892..0d2262a88 100644 --- a/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml +++ b/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml @@ -13,6 +13,7 @@ processors: match_type: regexp metric_names: - '^agents\.runtime\..*$' + - '^agents\.session\.tokens\.(input|output)$' batch/runtime_history: timeout: 5s send_batch_size: 5000 From d07ce3f2a91a01f493d077f32ba2f7114db35657 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:40:49 +0800 Subject: [PATCH 40/51] Preserve sparse runtime history gaps and remove backend-specific UI copy --- apps/web/e2e/agents-lifecycle.spec.ts | 22 +++--- .../features/dashboard/DashboardView.test.tsx | 2 +- .../src/features/dashboard/DashboardView.tsx | 2 +- .../dashboard/RuntimeObservabilityContent.tsx | 2 +- .../dashboard/runtime-history.test.ts | 70 +++++++++++++++++++ .../src/features/dashboard/runtime-history.ts | 4 ++ .../agents-client/src/runtime-history.test.ts | 2 +- 7 files changed, 90 insertions(+), 14 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index d1af2bc37..bd72d98b7 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2803,7 +2803,7 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await attachElementScreenshot(dashboard.locator(".dashboard-runtime-panel"), testInfo, "runtime-visual-dashboard-narrow"); }); -test("restores ClickHouse Runtime history after a Dashboard reload", async ({ page }) => { +test("restores retained Runtime history after a Dashboard reload", async ({ page }) => { const sessionId = "11111111-1111-4111-8111-111111111111"; const environmentId = "22222222-2222-4222-8222-222222222222"; const allocationId = "33333333-3333-4333-8333-333333333333"; @@ -2879,7 +2879,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa return pointStart === gapStart ? gap : point(pointStart, .25 + index / 1_000, 536_870_912 + index * 1_048_576); - }); + }).filter((point) => point.start !== end - 90); await route.fulfill({ status: 200, contentType: "application/json", @@ -2895,12 +2895,12 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa environment_id: environmentId, allocation_id: allocationId, started_at: { seconds: startedAt, nanoseconds: 123_456_789 }, provider_type: "docker", points, }], - token_usage: points.map((_, index) => ({ - start: start + index * 30, - end: start + (index + 1) * 30, - sampled_at: start + index * 30 + 20, - input_tokens: 10_000 + index * 300, - output_tokens: 2_000 + index * 60, + token_usage: points.map((point) => ({ + start: point.start, + end: point.end, + sampled_at: point.start + 20, + input_tokens: 10_000 + (point.start - start) * 10, + output_tokens: 2_000 + (point.start - start) * 2, })), }), }); @@ -2913,7 +2913,7 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await expect(dashboard.getByLabel("Runtime durable-history charts")).toBeVisible(); await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); await expect(dashboard).toContainText("120 buckets"); - await expect(dashboard).toContainText("120/120 observations"); + await expect(dashboard).toContainText("119/120 observations"); await expect(dashboard.getByText("Token throughput durable trend available")).toBeAttached(); const durableCpuChart = dashboard.getByLabel("CPU usage: 120 retained buckets"); await expect(dashboard.getByRole("region", { name: "CPU usage durable history chart" })).toBeVisible(); @@ -2921,6 +2921,8 @@ test("restores ClickHouse Runtime history after a Dashboard reload", async ({ pa await durableCpuChart.focus(); await durableCpuChart.press("ArrowLeft"); await expect(durableCpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Unavailable"); + await durableCpuChart.press("ArrowLeft"); + await expect(durableCpuCard.locator(".dashboard-runtime-trend-tooltip")).toContainText("Unavailable"); const durableMemoryCard = dashboard.getByRole("region", { name: "Memory usage durable history chart" }); const durableMemorySpan = await durableMemoryCard.locator("canvas").evaluate((canvas: HTMLCanvasElement) => { const context = canvas.getContext("2d"); @@ -3036,7 +3038,7 @@ test("publishes Dashboard counts only after every top-level Agent and Session pa expect(agentAfters).toEqual([null, "agent_b"]); // Session collection loads once for the page and once per Runtime snapshot. // Returning to Dashboard refreshes Runtime immediately instead of waiting 30 seconds. - expect(sessionAfters).toHaveLength(6); + await expect.poll(() => sessionAfters.length).toBe(6); expect(sessionAfters.filter((after) => after === null)).toHaveLength(3); expect(sessionAfters.filter((after) => after === "session_snapshot")).toHaveLength(3); }); diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index bce7423f0..887b61802 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -248,7 +248,7 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain("512 MiB / 2.00 GiB"); expect(html).toContain('aria-label="Runtime durable-history charts"'); expect(html).toContain("Resource trends"); - expect(html).toContain("ClickHouse-backed retained samples · explicit history source"); + expect(html).toContain("Retained samples · durable history"); expect(html).not.toContain('aria-label="Runtime trend source"'); expect(html).toContain("History · loading"); expect(html).toContain("0 buckets"); diff --git a/apps/web/src/features/dashboard/DashboardView.tsx b/apps/web/src/features/dashboard/DashboardView.tsx index 499311ee0..de2637444 100644 --- a/apps/web/src/features/dashboard/DashboardView.tsx +++ b/apps/web/src/features/dashboard/DashboardView.tsx @@ -384,7 +384,7 @@ export function DashboardView({

Runtime monitoring

-

Current provider status · ClickHouse-backed metrics history when configured

+

Current provider status · retained metrics history

{runtimeModel ? ( diff --git a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx index 25ab16d98..b11424475 100644 --- a/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx +++ b/apps/web/src/features/dashboard/RuntimeObservabilityContent.tsx @@ -383,7 +383,7 @@ export function RuntimeObservabilityContent({

Resource trends

-

{source === "durable" ? "ClickHouse-backed retained samples · explicit history source" : "Browser-local samples · reset on reload"}

+

{source === "durable" ? "Retained samples · durable history" : "Browser-local samples · reset on reload"}

{ expect(samples.every((sample) => sample.memoryUsageBytes === null && sample.memoryLimitBytes === null)).toBe(true); }); + it("keeps omitted buckets between distant observations as gaps", () => { + const source = history(); + const secondPoint = { + ...source.series[0]!.points[1]!, + start: 400, end: 430, first_observed_at: 410, last_observed_at: 410, + }; + const sparse = history({ + requested_range: { start: 100, end: 430 }, + generated_at: 431, + coverage: { + ...source.coverage, + last_sample_at: 410, + expected_sample_count: 11, + buckets: [source.coverage.buckets[0]!, { + ...source.coverage.buckets[1]!, + start: 400, end: 430, first_observed_at: 410, last_observed_at: 410, + }], + }, + series: [{ ...source.series[0]!, points: [source.series[0]!.points[0]!, secondPoint] }], + token_usage: [ + source.token_usage[0]!, + { ...source.token_usage[1]!, start: 400, end: 430, sampled_at: 410 }, + ], + }); + const samples = runtimeDurableTrendSamples([session], [sparse]); + expect(samples.map((sample) => sample.sampledAt)).toEqual( + Array.from({ length: 11 }, (_, index) => (130 + index * 30) * 1_000), + ); + expect(samples[0]?.targets[0]?.cpuRatio).toBe(.25); + expect(samples[10]?.targets[0]?.cpuRatio).toBe(.5); + expect(samples[10]?.inputTokensPerMinute).toBeNull(); + expect(samples[10]?.outputTokensPerMinute).toBeNull(); + for (const sample of samples.slice(1, -1)) { + expect(sample).toMatchObject({ + targets: [], memoryUsageBytes: null, memoryLimitBytes: null, + inputTokensPerMinute: null, outputTokensPerMinute: null, + }); + } + }); + + it("includes leading and trailing gaps with a shortened final bucket", () => { + const samples = runtimeDurableTrendSamples([session], [history({ + requested_range: { start: 70, end: 205 }, + generated_at: 206, + })]); + expect(samples.map((sample) => sample.sampledAt)).toEqual([100_000, 130_000, 160_000, 190_000, 205_000]); + expect(samples.map((sample) => sample.memoryUsageBytes)).toEqual([null, 512, 768, null, null]); + expect(samples.map((sample) => sample.targets.length)).toEqual([0, 1, 1, 0, 0]); + }); + + it("represents an entirely missing range without fabricating zero measurements", () => { + const samples = runtimeDurableTrendSamples([session], [history({ + requested_range: { start: 100, end: 175 }, + generated_at: 176, + coverage: { + retained_start: 100, first_sample_at: null, last_sample_at: null, + sample_count: 0, expected_sample_count: 3, buckets: [], + }, + series: [], + token_usage: [], + })]); + expect(samples.map((sample) => sample.sampledAt)).toEqual([130_000, 160_000, 175_000]); + for (const sample of samples) { + expect(sample).toMatchObject({ + targets: [], memoryUsageBytes: null, memoryLimitBytes: null, + inputTokensPerMinute: null, outputTokensPerMinute: null, + }); + } + }); + it("keeps aggregate token throughput absent when any queried Session lacks usage", () => { const second = { ...session, id: "44444444-4444-4444-8444-444444444444" } as AgentSession; const secondHistory = history({ session_id: second.id, token_usage: [] }); diff --git a/apps/web/src/features/dashboard/runtime-history.ts b/apps/web/src/features/dashboard/runtime-history.ts index 57b28ce38..201b154a3 100644 --- a/apps/web/src/features/dashboard/runtime-history.ts +++ b/apps/web/src/features/dashboard/runtime-history.ts @@ -101,6 +101,10 @@ export function runtimeDurableTrendSamples( }; for (const history of histories) { + const { start, end } = history.requested_range; + for (let bucketStart = start; bucketStart < end; bucketStart += history.resolution_seconds) { + bucket(Math.min(bucketStart + history.resolution_seconds, end) * 1_000); + } for (const coverage of history.coverage.buckets) bucket(coverage.end * 1_000); for (const usage of history.token_usage) { bucket(usage.end * 1_000).tokens.set(history.session_id, { diff --git a/packages/agents-client/src/runtime-history.test.ts b/packages/agents-client/src/runtime-history.test.ts index 476f593bd..2b7f5f652 100644 --- a/packages/agents-client/src/runtime-history.test.ts +++ b/packages/agents-client/src/runtime-history.test.ts @@ -138,7 +138,7 @@ describe("Runtime history client", () => { }); for (const [name, body] of [ - ["unknown field", capabilities({ backend: "clickhouse" })], + ["unknown field", capabilities({ backend: "postgres" })], ["inconsistent availability", capabilities({ available: false })], ["duplicate metrics", capabilities({ metrics: ["cpu", "cpu"] })], ["invalid retention", capabilities({ retention_seconds: 60, maximum_range_seconds: 120 })], From 4c58e12207410d7b94724790a2278771c01561eb Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:43:01 +0800 Subject: [PATCH 41/51] test: add opt-in real-model runtime history acceptance --- .../agents-api/tests/runtime_history_live.py | 175 ++++++++++++++++++ 1 file changed, 175 insertions(+) create mode 100644 services/agents-api/tests/runtime_history_live.py diff --git a/services/agents-api/tests/runtime_history_live.py b/services/agents-api/tests/runtime_history_live.py new file mode 100644 index 000000000..ba06f92fc --- /dev/null +++ b/services/agents-api/tests/runtime_history_live.py @@ -0,0 +1,175 @@ +#!/usr/bin/env python3 +"""Opt-in real-model Runtime history acceptance against an isolated deployment. + +Run `collect`, restart the owned Core externally, then run `verify-restart`. +The caller owns deployment, private model configuration and final Session cleanup. +Use a dedicated report directory per provider. This never restarts any service. +Raw native/SQL sample comparisons are separate operator evidence: successive live +observations cannot be expected to have identical counter values. +""" + +import argparse +import importlib.util +import json +import os +from pathlib import Path +import time +import uuid + + +ROOT = Path(__file__).resolve().parents[3] +SPEC = importlib.util.spec_from_file_location("installer_acceptance", ROOT / "deploy/install/acceptance.py") +base = importlib.util.module_from_spec(SPEC) +SPEC.loader.exec_module(base) + + +class SafeAPI(base.API): + def __init__(self, origin, token, evidence, secrets): + super().__init__(origin, token, evidence) + self.secrets = secrets + + def request(self, *args, **kwargs): + value = super().request(*args, **kwargs) + encoded = json.dumps(value) + base.require(not any(secret in encoded for secret in self.secrets), "public_response_exposed_secret") + base.require(not any(marker in encoded for marker in ( + "postgres://", "postgresql://", "runtime_history_samples", "token_sha256", + "unix:///", "host.microsandbox.internal", "/home/parsar-acceptance/")), + "public_response_exposed_backend_configuration") + return value + + +def stable_history(value): + return {key: item for key, item in value.items() if key != "generated_at"} + + +def check_isolation(api, foreign, record, query): + path = base.session_path(record) + for suffix in ("/runtime-observation", "/runtime-history"): + try: + foreign.get(path + suffix, "foreign" + suffix, query if suffix.endswith("history") else None) + except base.Failure: + base.require(foreign.evidence[-1].get("http_status") == 404, "foreign_resource_not_hidden") + else: + raise base.Failure("foreign_resource_visible") + page = foreign.get("/runtime-observations", "foreign.observations") + base.require(not any(row["session_id"] == record["session_id"] for row in page["data"]), + "foreign_observation_list_exposed_session") + record["checks"].append("foreign_history_and_current_observation_hidden") + + +def collect(api, foreign, record, config, args, save): + base.require("session_id" not in record, "collection_already_attempted") + capabilities = api.get("/runtime-history/capabilities", "history.capabilities") + base.require(capabilities.get("available") is True and capabilities.get("collection_mode") == "periodic", + "builtin_periodic_history_unavailable") + record["capabilities"] = capabilities + record["start"] = int(time.time()) - 60 + nonce = uuid.uuid4().hex + prompt = ("Use your native shell tool to run exactly once: python3 -c 'import time; " + "end=time.monotonic()+45; x=0\nwhile time.monotonic()= 2, "insufficient_live_observations") + base.require(all(sample["provider_type"] == args.provider for sample in observed), "wrong_runtime_provider") + base.require(all(sample["cpu"]["usage_seconds_total"] is not None and sample["memory"]["usage_bytes"] is not None + for sample in observed), "missing_native_counters") + base.require(any(sample["cpu"]["usage_seconds_total"] > observed[0]["cpu"]["usage_seconds_total"] + for sample in observed[1:]), "cpu_counter_did_not_advance_during_execution") + record["observations"] = observations + record["turn"] = completed + record["checks"].append("provider_current_cpu_and_memory_observed_during_execution") + time.sleep(min(capabilities["sample_interval_seconds"] + 2, 60)) + query = {"start": record["start"], "end": int(time.time()), "max_points": 120} + history = api.get(base.session_path(record) + "/runtime-history", "history.before_restart", query) + base.require(history.get("source") == "durable" and history["coverage"]["sample_count"] >= 2, + "periodic_history_not_persisted") + base.require(bool(history["series"]), "native_resource_history_missing") + base.require(bool(history["token_usage"]), "canonical_usage_history_missing") + record.update(query=query, history=history) + check_isolation(api, foreign, record, query) + record["checks"].append("periodic_cpu_memory_and_usage_history_read") + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("stage", choices=("collect", "verify-restart")) + parser.add_argument("--base-url", required=True) + parser.add_argument("--caller-key-file", type=Path, required=True) + parser.add_argument("--foreign-key-file", type=Path, required=True) + parser.add_argument("--model-config-file", type=Path, required=True) + parser.add_argument("--report-dir", type=Path, required=True) + parser.add_argument("--provider", choices=("docker", "microsandbox"), required=True) + parser.add_argument("--harness", choices=base.HARNESSES, default="codex") + parser.add_argument("--timeout", type=int, default=300) + args = parser.parse_args() + os.umask(0o077) + token, models, secrets = base.settings(args) + foreign_token = base.private_file(args.foreign_key_file).strip() + secrets.append(foreign_token) + origin = base.validate_origin(args.base_url) + base.require(args.report_dir.is_absolute(), "absolute_report_directory_required") + args.report_dir.mkdir(mode=0o700, parents=True, exist_ok=True) + base.require(args.report_dir.stat().st_mode & 0o077 == 0, "report_directory_must_be_private") + path = args.report_dir / "history.json" + record = json.loads(base.private_file(path)) if path.exists() else { + "provider": args.provider, "harness": args.harness, "checks": [], "stages": {}} + base.require(record["provider"] == args.provider and record["harness"] == args.harness, + "report_configuration_mismatch") + base.require(args.stage not in record["stages"], "stage_already_attempted") + stage = {"passed": False, "requests": []} + record["stages"][args.stage] = stage + + def save(): + encoded = json.dumps(record, indent=2) + "\n" + base.require(not any(secret in encoded for secret in secrets), "secret_in_report") + path.write_text(encoded) + path.chmod(0o600) + + api = SafeAPI(origin, token, stage["requests"], secrets) + foreign = SafeAPI(origin, foreign_token, stage["requests"], secrets) + try: + if args.stage == "collect": + collect(api, foreign, record, models[args.harness], args, save) + else: + base.require(record["stages"].get("collect", {}).get("passed"), "collection_did_not_pass") + history = api.get(base.session_path(record) + "/runtime-history", "history.after_restart", record["query"]) + base.require(stable_history(history) == stable_history(record["history"]), "history_changed_across_restart") + record["after_restart_history"] = history + check_isolation(api, foreign, record, record["query"]) + record["checks"].append("fixed_range_history_survives_operator_core_restart_and_new_reader") + stage["passed"] = True + except Exception as error: + stage["failure"] = str(error) if isinstance(error, base.Failure) else "unexpected_local_or_response_error" + stage["cleanup"] = base.settle_failed_work(api, record) + save() + print(args.provider + " " + args.stage + (" passed" if stage["passed"] else " FAILED: " + stage["failure"])) + return 0 if stage["passed"] else 1 + + +if __name__ == "__main__": + try: + raise SystemExit(main()) + except Exception: + raise SystemExit("Acceptance configuration failed; private values withheld.") from None From 66e73d8e1ba93d335a371c958410c22e74882eb5 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:44:08 +0800 Subject: [PATCH 42/51] Use Core database history by default with independent optional export --- .env.example | 4 +- CONTRIBUTING.md | 12 + contracts/agents-api/runtime-history-api.md | 17 +- .../runtime-observability-design.md | 148 ++++-------- contracts/agents-api/runtime-observability.md | 75 +++--- docs/web/README.md | 14 +- docs/web/README.zh-CN.md | 11 +- services/agents-api/cmd/server/main.go | 18 +- .../agents-api/cmd/server/runtime_history.go | 223 ++++++++---------- .../cmd/server/runtime_history_test.go | 103 +++----- .../internal/runtimeobs/exporter.go | 9 +- .../runtimeobs/exporter_independence_test.go | 33 +++ .../agents-api/internal/runtimeobs/service.go | 17 +- services/agents-api/runtime-history/README.md | 56 +++++ .../clickhouse/001_runtime_history.sql | 90 ------- .../clickhouse/002_session_token_usage.sql | 31 --- .../runtime-history/clickhouse/README.md | 94 -------- .../clickhouse/otel-collector.example.yaml | 44 ---- 18 files changed, 349 insertions(+), 650 deletions(-) create mode 100644 services/agents-api/internal/runtimeobs/exporter_independence_test.go create mode 100644 services/agents-api/runtime-history/README.md delete mode 100644 services/agents-api/runtime-history/clickhouse/001_runtime_history.sql delete mode 100644 services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql delete mode 100644 services/agents-api/runtime-history/clickhouse/README.md delete mode 100644 services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml diff --git a/.env.example b/.env.example index cbb3e5d05..482c6add0 100644 --- a/.env.example +++ b/.env.example @@ -7,8 +7,8 @@ AGENTS_API_PROXY_TARGET=http://127.0.0.1:8091 AGENTS_API_PROXY_TOKEN_FILE=~/.parsar/agents-api/web-token # Optional server-only OTLP/HTTP Runtime history exporter configuration. # When unset, current observations and the Dashboard require no telemetry backend. -# Optional OTLP export, periodic sampling, and ClickHouse Durable history config: -# services/agents-api/runtime-history/clickhouse/README.md +# Runtime history uses the Core PostgreSQL database with 30-second periodic sampling. +# Optional sampling/OTLP export overrides: services/agents-api/runtime-history/README.md # AGENTS_API_RUNTIME_HISTORY_FILE=/absolute/path/runtime-history.json # Set sample_interval_seconds (5..300) in that server-only file to enable the # execution-owner singleton sampler. Omit it for on-read export only. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 43e3e9ee7..128a2f06f 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -307,6 +307,18 @@ compute uptime semantics across both. Do not use microsandbox's instantaneous CP percent, wake suspended compute, or expose provider-native identifiers to fill a common field. +Runtime history uses the existing Core PostgreSQL database: one sanitized row per +periodic observation, seven-day retention and bounded reads. It is best-effort +operational evidence, not execution or Usage authority. The execution owner samples +by default every 30 seconds. Existing canonical Session Usage supplies token +snapshots; never aggregate provider counters as model tokens. Preserve missing data +and reset CPU derivation across compute incarnations or counter regressions. +The bounded asynchronous database writer and optional OTLP exporter have independent +queues; external telemetry outages must not stall local history or execution. +Retention cleanup also runs without active Runtimes. The browser queries only Core, +never storage or a Collector, and stays a lightweight administrator console. +No additional metrics database or Collector is required for retained charts. + In V1, our daemon fills the user-side executor role. Users deploy daemon, the selected harness, local tools and workspace together. Do not require Codex `exec-server`, a service-side harness, registry/Noise transport or remote tool diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md index 2a73576c4..b394998c9 100644 --- a/contracts/agents-api/runtime-history-api.md +++ b/contracts/agents-api/runtime-history-api.md @@ -1,14 +1,13 @@ # Runtime history API -Status: public contract, strict TypeScript client, optional ClickHouse reference -Reader, and capability-gated Core Web History ranges implemented. No Reader is -configured by default. The capability route therefore advertises -`available=false` until an operator supplies both a validated Reader and qualified -periodic collection. +Status: public contract, strict TypeScript client, PostgreSQL history and Core Web +History ranges implemented. Core uses its existing database; the execution owner +samples every 30 seconds by default. API-only processes without an execution worker +advertise on-read collection rather than claiming periodic coverage. -This is an optional Agents Core extension. It is read-only and backend-neutral. -The browser never receives a ClickHouse endpoint, OTLP credential, provider-native -identity, or tenant selector. +This is an Agents Core extension. It is read-only and backend-neutral. The browser +never receives a storage endpoint, OTLP credential, provider-native identity or +tenant selector. ## Capability discovery @@ -165,7 +164,7 @@ or malformed data reject the entire response with a 502 client projection error. ## Explicit boundaries - The routes never sample a live provider, provision compute, or mutate lifecycle. -- The contract does not choose ClickHouse, Prometheus, Mimir, or another backend. +- The public contract does not expose a storage backend. - History availability does not imply current Runtime readiness. - Current observations and Durable history have separate freshness and retention semantics and must remain separately labelled in Web. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 39ba49f79..0a383e235 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -4,8 +4,8 @@ Status: provider abstraction with Docker and microsandbox sampling, the current-snapshot API/client contract, and Core Web Live and capability-gated Durable Dashboard sources are implemented. Phase 4 includes the bounded sanitized exporter seam, optional OTLP/HTTP transport, execution-owner singleton background -sampling, ClickHouse projection/Reader, public history API, and 1h/6h/24h Web -ranges. History remains disabled by default. Other provider sources are not implemented. +sampling, PostgreSQL history, public history API, and 1h/6h/24h Web ranges. +History uses the existing Core database by default. Other provider sources are not implemented. Microsandbox idle suspension is a separate durable lifecycle feature; it does not consume this telemetry as authority. @@ -243,57 +243,23 @@ Provider type, Runtime mode, and coarse status are safe low-cardinality labels. High-cardinality identities require tenant-scoped access and retention policies; they are not global Prometheus labels by default. -### 10.1 Qualified public implementation reference - -The first Phase 4 qualification uses E2B Runtime commit -`ccf2a64ee40472645209b92525a5459d413bce76` as implementation evidence, not as -an API contract to copy. Its sandbox observer samples every five seconds, exports -provider metrics through OTLP, and attaches sandbox and team identity. The -OpenTelemetry Collector sends ordinary operational metrics to Mimir but routes -the high-cardinality `e2b.*` sandbox series to ClickHouse. Its authenticated API -derives team identity from the caller, queries with both `team_id` and -`sandbox_id`, validates the requested time range, calculates a bounded step, and -retains the specialized sandbox table for seven days. Relevant public files are: - -- [`packages/orchestrator/pkg/metrics/sandboxes.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/orchestrator/pkg/metrics/sandboxes.go) - for bounded collection and identity attributes; -- [`packages/local-dev/otel-collector.yaml`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/local-dev/otel-collector.yaml) - for OTLP fan-out to Mimir and ClickHouse; -- [`packages/clickhouse/migrations/20250717135224_sandbox_metrics.sql`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/clickhouse/migrations/20250717135224_sandbox_metrics.sql) - for the high-cardinality history schema and retention; and -- [`packages/api/internal/clusters/resources_local.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/api/internal/clusters/resources_local.go) - plus [`packages/clickhouse/pkg/sandbox.go`](https://github.com/e2b-dev/runtime/blob/ccf2a64ee40472645209b92525a5459d413bce76/packages/clickhouse/pkg/sandbox.go) - for tenant-scoped, downsampled reads. - -Dify commit `a068c47ea993ccc0f943131c274b7830b16de9f4` independently demonstrates an -optional OTLP exporter that becomes a no-op when disabled, but it does not -provide a Runtime-incarnation history query boundary. It supports the exporter -choice, not the history adapter design. - -For Core, the qualified topology is therefore: - -1. a bounded, best-effort provider-neutral handoff after validated current - observations; -2. an optional operator-managed OTLP Collector; -3. a high-cardinality history store behind a separate server-side adapter; and -4. authenticated Core history routes that resolve tenant and Session ownership - before issuing a backend query. - -Mimir remains suitable for low-cardinality service health. The initial Runtime -history qualification does not treat a shared Prometheus label filter as a -tenant security boundary and does not let Web query Mimir, ClickHouse, or the -Collector directly. ClickHouse is the first reference backend because its query -shape can require tenant, Session, allocation, and incarnation predicates, but -the public API and `runtimehistory` interface must remain backend-neutral. The -backend, Collector, and exporter are disabled by default and are not required for -Session execution or current observations. +### 10.1 Lightweight deployment + +Core reuses its PostgreSQL database for bounded recent Runtime history. The +provider-neutral observation service hands each sanitized periodic result to a +bounded asynchronous writer. One typed row contains the observation and optional +canonical Session Usage snapshot. The public API resolves caller ownership before +issuing bounded queries; Web never queries storage directly. + +External OTLP export remains optional. Each destination has an independent queue, +so a Collector outage cannot delay local persistence. Neither history nor export +is execution or lifecycle authority. No additional metrics service is deployed. ### 10.2 OTLP transport configuration and instruments -Core enables Runtime history export only when -`AGENTS_API_RUNTIME_HISTORY_FILE` points to a server-only JSON file. With the -variable unset, no exporter is created and no Collector or history store is -required. A minimal configuration is: +Core enables external export only when `AGENTS_API_RUNTIME_HISTORY_FILE` contains +an OTLP endpoint. With the variable unset, local history and 30-second sampling +remain enabled. An optional server-only configuration is: ```json { @@ -311,11 +277,10 @@ committed. Plain HTTP requires the explicit combination of an `http` endpoint and `"insecure": true`; HTTPS rejects that flag. Endpoint userinfo, query strings, fragments, invalid headers, reserved transport headers, queues above 4096 records, and timeouts above 30 seconds fail startup without echoing config -contents. `sample_interval_seconds` is optional; values from 5 through 300 enable -the deployment sampler, while omission retains on-read export only. Periodic -sampling requires the execution Worker because its database lease is the -deployment singleton boundary. Export is best effort through the bounded queue -documented above. +contents. `sample_interval_seconds` accepts 5 through 300; omission uses 30. +Periodic sampling requires the execution Worker and its existing database lease. +API-only processes without that worker advertise on-read collection. External +export is optional and best effort; PostgreSQL writes use a separate bounded queue. The OTLP request uses standard protobuf metrics and these instruments: @@ -359,19 +324,17 @@ A history API must expose actual sample coverage. Core Web must not advertise a durable range until the operator backend, query adapter, and a qualified periodic collection cadence are all configured. -The reference ClickHouse Reader uses the same server-only file under an optional -`clickhouse` object. Its native address, database, reader username, password, TLS -mode, and bounded dial/query timeouts are never capability fields. The fixed -reference policy is seven-day retention, a 24-hour maximum range, at most 1,000 -buckets per series, 64 series, and 10,000 returned points. Operator schema and -Collector examples live under `services/agents-api/runtime-history/clickhouse`. -Exactly one of `secure` or `insecure` must be set for the native connection; -plaintext transport is never inferred from an omitted TLS flag. +The PostgreSQL Reader shares Core database access and SQLC ownership. It limits +queries to seven-day retention, a 24-hour range, 1,000 buckets per series, 64 series +and 10,000 output points. Raw input has a separate bounded budget, so dense sampling +does not consume the output budget before downsampling. A bounded periodic cleanup +removes expired rows even when no Runtime is active. Expired rows are excluded from +reads immediately; physical removal is incremental. ### 10.3 Backend-neutral history query boundary `services/agents-api/internal/runtimehistory` defines the server-side query -contract independently from ClickHouse, OTLP, and the public HTTP shape. Its +contract independently from SQL, OTLP, and the public HTTP shape. Its service resolves the authenticated tenant and Session to durable Core identity before calling a Reader. Reader queries always carry tenant, Session, and Environment scope plus a bounded start, exclusive end, server-selected step, @@ -396,11 +359,10 @@ is not sufficient to advertise a Durable Dashboard source. Backend identity, URLs, credentials, and tenant data are never capability fields. This internal boundary, the Session-scoped public extension, capability discovery, -strict client, and production ClickHouse reference Reader are implemented. The -Reader queries only the specialized projection, always includes tenant, Session, -Environment, bounded time, and `collection_source = 'periodic'` predicates, and -aggregates resource points by allocation, resetting CPU derivation after a -cumulative-counter regression. Durable +strict client, and PostgreSQL Reader are implemented. The Reader includes tenant, +Session, Environment and bounded-time predicates. Only periodic records are stored. +It aggregates resource points by allocation, resetting CPU derivation across compute +incarnations, missing counters and cumulative-counter regressions. Durable Web ranges remain gated on real retention, isolation, restart, and incarnation acceptance. @@ -490,8 +452,8 @@ history sweep, the Core resolver reads the existing canonical cumulative Session Usage snapshot from the execution store alongside Runtime identity. The exporter emits Session-scoped input/output token gauges with the same Session and sampling time, independently of Docker, microsandbox, Kubernetes, or another provider. -ClickHouse retains those cumulative points separately from allocation/incarnation -series. Web derives throughput from adjacent nondecreasing points. Missing or +PostgreSQL retains those cumulative snapshots alongside the sample; query results +keep Session token points separate from allocation series. Web derives throughput from adjacent nondecreasing points. Missing or incomplete native usage and counter regressions remain gaps, never zero. The current snapshot API still does not duplicate Usage fields: Web joins its @@ -516,9 +478,9 @@ authorized backend, never by exposing product credentials to Core Web. ## 14. Data model impact -Current snapshot and Dashboard work require no migration. Existing -`runtime_allocations`, `environments`, Sessions, Turns, and Usage are sufficient. -No time-series table is proposed. +Current snapshots reuse existing control-plane resources. Retained history adds +one bounded observation table to the Core database. It does not change canonical +Sessions, Turns or Usage and cannot authorize execution or lifecycle changes. Automatic idle shutdown is a separate feature. It requires durable fields such as `activity_revision`, `idle_since`, and `shutdown_requested_at` with fenced state @@ -559,34 +521,16 @@ Implemented for the browser-local current-snapshot live window. - Build an explicitly ephemeral live window from complete Web snapshots. - Expose 15-minute and one-hour views with an explicit browser-local source label. -### Phase 4: optional history - -- Implemented: bounded asynchronous handoff of sanitized, validated current - observation results. It is disabled by default, drops on queue saturation, and - cannot fail the current-observation request path. -- Implemented: optional server-only OTLP/HTTP protobuf transport for the six - documented Runtime instruments, including allocation identity and provider - start metadata. Configuration is strict and secrets never reach Web. -- Implemented: optional execution-owner singleton sampling across all nondeleted - managed Sessions. Keyset scans, provider concurrency, source deadlines, and - non-overlapping sweeps are bounded; collection source is exported explicitly. -- Implemented: backend-neutral `runtimehistory` types and service validation. - Tenant/Session/Environment scope precedes every Reader query; allocation, - coverage, nullability, ordering, range and total-point invariants are enforced. -- Implemented: safe public capability discovery, bounded Session-scoped history - query routes, generated OpenAPI schemas, and strict `packages/agents-client` - projection. Unconfigured or on-read-only deployments cannot advertise Durable. -- Implemented: optional ClickHouse Reader, seven-day schema/TTL projection, - Collector example, strict server-only configuration, and mandatory - tenant/Session/Environment/periodic-source query predicates. No backend is a - Core execution dependency. -- Qualified: real OTLP Collector-to-ClickHouse acceptance covers tenant isolation, - periodic-only public reads, Reader restart persistence, continuous allocation - series, and CPU baseline reset after cumulative-counter regression. -- Implemented: Core Web discovers capabilities, reloads bounded Session histories - with bounded concurrency, and exposes explicit Live versus History sources with - 1h, 6h, and 24h Durable ranges, including canonical Session token throughput. -- Not implemented: exporter queue/drop/error coverage telemetry. +### Phase 4: retained history + +- Bounded asynchronous writes of sanitized periodic observations to Core PostgreSQL. +- Optional OTLP/HTTP export with server-only credentials and a separate queue. +- Execution-owner sampling with bounded pages, concurrency and source deadlines. +- Tenant-scoped history queries with input/output limits, explicit gaps and CPU fences. +- Seven-day retention and bounded cleanup independent of active Runtime count. +- Core Web history restoration after refresh, including canonical token snapshots. +- Real Docker/microsandbox and PostgreSQL acceptance remains mandatory for delivery. +- Exporter queue/drop/error coverage counters remain outside this batch. ### Phase 5: additional sources diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index 76018fb8d..305424b7c 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -78,51 +78,30 @@ an activity revision and timestamps such as `idle_since` and `shutdown_requested_at`. Metrics, an in-memory cache, or a monitoring backend must not become the lifecycle authority. -## Optional history export boundary - -An optional server-only OTLP/HTTP exporter can forward validated observation -results to an operator Collector when `AGENTS_API_RUNTIME_HISTORY_FILE` is set. -It is disabled by default and does not change the current API result, execution -ownership, or lifecycle authority. The file may contain endpoint authorization -headers and is never returned to Web. Provider receipts, native identifiers, raw -errors, paths, and credentials are excluded from metric attributes. CPU, -capacity, and memory points require the provider-qualified compute `started_at` -fence; unfenced observations export coverage and read duration only. Core -resolved/observed nanosecond attributes provide a backend join key even when -generic OTLP storage lowers the event timestamp precision. - -The server-only file may also enable a bounded periodic cadence. Only the Core -service holding the execution database lease runs that deployment-wide sampler; -it scans nondeleted managed Sessions with bounded pages and concurrency, gives -each provider read an independent deadline, monitors lease ownership throughout -the sweep, rechecks ownership before export, and never mutates Runtime state. -Exports distinguish `periodic` samples from `on_read` samples. - -The Collector, high-cardinality history backend, server-side history query -adapter, retention policy, and durable Web ranges remain separate optional -capabilities in the [full design](runtime-observability-design.md). A ClickHouse -Reader plus reference projection/Collector configuration is available under -`services/agents-api/runtime-history/clickhouse`; it is disabled by default. When -configured with qualified periodic sampling, Web advertises explicit 1h, 6h, and -24h History ranges and reconstructs them after reload. - -The internal `runtimehistory` boundary is backend-neutral and validates Core -scope, incarnation fences, bucket coverage, nullability, time bounds and total -point limits. The public capability and Session history routes plus strict client -projection are implemented, but they do not make history available by themselves: -an operator must configure the Reader and qualified periodic collection. - -## First-phase boundary - -The current-snapshot implementation adds no migration, metrics backend, token -duplication, Kubernetes/E2B source, or telemetry-driven lifecycle action. Optional -Durable history separately samples canonical Session Usage into its configured -metrics backend; provider sources still do not own token accounting. The -internal source interface admits Docker and microsandbox without changing Session -attribution or the existing sandbox lifecycle interface. - -The current API and browser-local Web live window are documented in the -[full design](runtime-observability-design.md) and the -[current-snapshot extension](runtime-observability-api.md), and the optional -[history extension](runtime-history-api.md). Additional providers, self-hosted -telemetry, exporter health counters, and idle-policy authority remain later phases. +## Retained history and optional export + +Periodic Runtime observations are persisted asynchronously in the existing Core +PostgreSQL database. The execution owner samples every 30 seconds by default, +using bounded pages, concurrency and source deadlines. Collection never wakes or +mutates compute. The worker lease is checked during the sweep and before each +handoff. Only periodic samples populate durable history; current API reads cannot +inflate cadence coverage. Retention is seven days; public queries span at most +24 hours and have explicit input and output limits. + +`AGENTS_API_RUNTIME_HISTORY_FILE` optionally changes sampling and adds OTLP/HTTP +export. The local database and external exporter have independent bounded queues. +No Collector is required for the Dashboard. Failures and queue saturation remain +missing observations rather than fabricated zeroes or failed executions. Transport +credentials stay server-only; native identifiers, receipts, raw errors, paths and +credentials are excluded from observations. + +The internal `runtimehistory` boundary validates Core scope, bucket coverage, +nullability, time bounds and point limits. One chart series represents an allocation; +CPU deltas reset across compute incarnations or counter regressions. Canonical +Session Usage supplies independently sampled token counters. History queries survive +Core restart and browser reload without replaying execution. + +See the [design](runtime-observability-design.md), +[current API](runtime-observability-api.md), [history API](runtime-history-api.md) +and [configuration](../../services/agents-api/runtime-history/README.md). +Additional provider telemetry and idle-policy authority remain separate work. diff --git a/docs/web/README.md b/docs/web/README.md index a3fc3319f..ee7eaa86c 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -33,14 +33,12 @@ outside the browser. Dashboard is the starting point. It summarizes the current Agent and Session results and loads complete tenant-scoped Runtime observation snapshots without inventing missing values. The searchable, filterable, sortable, paginated semantic table -remains available on demand. When the operator enables qualified periodic sampling -plus the ClickHouse Reader, Web automatically discovers the capability and uses the -retained 1-hour, 6-hour, and 24-hour History ranges as the primary chart source. -A reload reconstructs History from Core instead of briefly publishing a new -browser-local window; Web never receives ClickHouse credentials or queries it -directly. An unconfigured deployment falls back to a bounded browser-local Live -window. Token throughput stays Live-only because Runtime telemetry does not -duplicate canonical Session Usage. +remains available on demand. Core stores periodic observations in its existing +PostgreSQL database, with no separate monitoring stack. Web discovers periodic +history and offers 1-hour, 6-hour and 24-hour ranges; reload restores data through +Core. API-only deployments without a sampler retain the browser-local Live view. +Token throughput uses snapshots of canonical Session Usage, preserving missing data. +Web never connects to a database or receives telemetry credentials. When a provider reports only cumulative CPU time, Web derives interval utilization only across adjacent samples from the same verified Runtime incarnation; restarts and counter regressions create diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index 2d33df4d8..6ead8d1a9 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -30,13 +30,10 @@ Core 部署提供完整的产品界面,同时让凭据和执行能力始终留 Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,并加载完整的租户级 Runtime 观测快照,不会把缺失值伪装成 0。支持搜索、状态/模式筛选、排序、分页的语义表格按需 展开;页面也展示需要关注的 Session,并可直接进入创建 Agent 或启动 Session 的流程。 -运维方配置周期采样和 ClickHouse Reader 后,页面优先使用可跨刷新的 1 小时、6 小时和 -24 小时持久 History,不会先发布一个新的浏览器本地窗口再切换数据源。Web 只通过 Core -读取,不接收 ClickHouse 凭据,也不会直接连接 ClickHouse。未配置 History 的部署仍回退 -到有界的浏览器本地 live window。若 provider -只报告累计 CPU 时间,Web 仅在 -相邻样本属于同一已验证 Runtime incarnation 时计算区间利用率;重启或计数回退会形成 -数据缺口,不会制造峰值。 +Core 默认将周期采样存入现有 PostgreSQL,无需额外监控服务。页面通过 Core 查询 +1 小时、6 小时和 24 小时历史,刷新后仍可读取。未启用执行 worker 的部署保留 +浏览器本地 Live 视图。Token 速率来自 Session 的实际用量快照,缺失数据不按零计算。 +Web 不直接访问数据库,也不接收监控凭据。 ![Runtime 监控实时趋势](images/runtime-dashboard.png) diff --git a/services/agents-api/cmd/server/main.go b/services/agents-api/cmd/server/main.go index 91a27cc92..09cd09617 100644 --- a/services/agents-api/cmd/server/main.go +++ b/services/agents-api/cmd/server/main.go @@ -111,18 +111,11 @@ func run() error { if err != nil { return err } - history, err := runtimeHistory(ctx) + history, err := runtimeHistory(ctx, executionStore, os.Getenv("AGENTS_API_DAEMON_WS_URL") != "") if err != nil { return err } - if history.Reader != nil { - defer closeRuntimeHistoryReader(history.Reader) - } - observationOptions := []runtimeobs.ServiceOption{} - if history.Option != nil { - observationOptions = append(observationOptions, history.Option) - } - observationService, err := runtimeobs.NewService(observationResolver, observationSources, observationOptions...) + observationService, err := runtimeobs.NewService(observationResolver, observationSources, history.Options...) if err != nil { if history.Exporter != nil { closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) @@ -139,6 +132,13 @@ func run() error { closeRuntimeHistory(closeCtx, history.Exporter) } }() + cleanupCtx, cancelCleanup := context.WithCancel(ctx) + cleanupDone := make(chan struct{}) + go func() { + defer close(cleanupDone) + runHistoryCleanup(cleanupCtx, history.Prune) + }() + defer func() { cancelCleanup(); <-cleanupDone }() var workerDone chan error var worker *execution.Worker options := []api.Option{api.WithSubagents(executionStore), api.WithSkills(executionStore), api.WithSourceFiles(executionStore), api.WithSessionArtifacts(executionStore), api.WithRuntimeObservations(observationService)} diff --git a/services/agents-api/cmd/server/runtime_history.go b/services/agents-api/cmd/server/runtime_history.go index d75f2bde8..131c5261e 100644 --- a/services/agents-api/cmd/server/runtime_history.go +++ b/services/agents-api/cmd/server/runtime_history.go @@ -6,58 +6,44 @@ import ( "encoding/json" "errors" "io" - "net" "net/url" "os" "strings" "time" - "golang.org/x/net/http/httpguts" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory/clickhousereader" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory/postgresreader" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" + "golang.org/x/net/http/httpguts" ) const ( - defaultRuntimeHistoryQueueCapacity = 256 - defaultRuntimeHistoryTimeoutSeconds = 2 - maxRuntimeHistoryQueueCapacity = 4096 - maxRuntimeHistoryTimeoutSeconds = 30 - minRuntimeHistorySampleIntervalSeconds = 5 - maxRuntimeHistorySampleIntervalSeconds = 300 - defaultRuntimeHistoryDialTimeoutSeconds = 5 - defaultRuntimeHistoryQueryTimeoutSeconds = 10 + defaultRuntimeHistoryQueueCapacity = 256 + defaultRuntimeHistoryTimeoutSeconds = 2 + maxRuntimeHistoryQueueCapacity = 4096 + maxRuntimeHistoryTimeoutSeconds = 30 + minRuntimeHistorySampleIntervalSeconds = 5 + maxRuntimeHistorySampleIntervalSeconds = 300 ) type runtimeHistoryConfig struct { - Transport string `json:"transport"` - Endpoint string `json:"endpoint"` - Insecure bool `json:"insecure"` - Headers map[string]string `json:"headers,omitempty"` - QueueCapacity int `json:"queue_capacity,omitempty"` - TimeoutSeconds int `json:"timeout_seconds,omitempty"` - SampleIntervalSeconds int `json:"sample_interval_seconds,omitempty"` - ClickHouse *runtimeHistoryClickHouseConfig `json:"clickhouse,omitempty"` -} - -type runtimeHistoryClickHouseConfig struct { - Address string `json:"address"` - Database string `json:"database"` - Username string `json:"username"` - Password string `json:"password,omitempty"` - Secure bool `json:"secure,omitempty"` - Insecure bool `json:"insecure,omitempty"` - DialTimeoutSeconds int `json:"dial_timeout_seconds,omitempty"` - QueryTimeoutSeconds int `json:"query_timeout_seconds,omitempty"` + Transport string `json:"transport,omitempty"` + Endpoint string `json:"endpoint,omitempty"` + Insecure bool `json:"insecure,omitempty"` + Headers map[string]string `json:"headers,omitempty"` + QueueCapacity int `json:"queue_capacity,omitempty"` + TimeoutSeconds int `json:"timeout_seconds,omitempty"` + SampleIntervalSeconds int `json:"sample_interval_seconds,omitempty"` } type runtimeHistorySetup struct { - Option runtimeobs.ServiceOption + Options []runtimeobs.ServiceOption Exporter runtimeHistoryExporter - Reader runtimeHistoryReader + Reader runtimehistory.Reader SampleInterval time.Duration + Prune func(context.Context) error } type runtimeHistoryExporter interface { @@ -65,32 +51,57 @@ type runtimeHistoryExporter interface { Close(context.Context) error } -type runtimeHistoryReader interface { - runtimehistory.Reader - Close() error -} - -var openRuntimeHistoryReader = func(ctx context.Context, config clickhousereader.Config) (runtimeHistoryReader, error) { - return clickhousereader.Open(ctx, config) -} - -func runtimeHistory(ctx context.Context) (runtimeHistorySetup, error) { - file := os.Getenv("AGENTS_API_RUNTIME_HISTORY_FILE") - if file == "" { - return runtimeHistorySetup{}, nil +func runtimeHistory(ctx context.Context, coreStore *store.Store, executionEnabled bool) (runtimeHistorySetup, error) { + config, err := loadRuntimeHistoryConfig() + if err != nil { + return runtimeHistorySetup{}, err + } + interval := time.Duration(config.SampleIntervalSeconds) * time.Second + mode := runtimehistory.CollectionPeriodic + if !executionEnabled { + interval = 0 + mode = runtimehistory.CollectionOnRead } - raw, err := os.ReadFile(file) + capabilities := runtimehistory.Capabilities{ + CollectionMode: mode, SampleInterval: interval, Retention: 7 * 24 * time.Hour, + MinimumStep: max(30*time.Second, interval), MaximumRange: 24 * time.Hour, + MaximumPoints: 1000, MaximumSeries: 64, MaximumTotalPoints: 10000, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory, runtimehistory.MetricTokens}, + } + timeout := time.Duration(config.TimeoutSeconds) * time.Second + backend, err := postgresreader.New(coreStore, postgresreader.Config{Capabilities: capabilities, QueryTimeout: timeout}) if err != nil { - return runtimeHistorySetup{}, errors.New("cannot read AGENTS_API_RUNTIME_HISTORY_FILE") + return runtimeHistorySetup{}, err } + options := runtimeobs.ExportOptions{QueueCapacity: config.QueueCapacity, Timeout: timeout} + setup := runtimeHistorySetup{Options: []runtimeobs.ServiceOption{runtimeobs.WithExporter(backend, options)}, Reader: backend, SampleInterval: interval, Prune: backend.Prune} + if config.Endpoint != "" { + exporter, err := otlpexporter.New(ctx, otlpexporter.Config{Endpoint: config.Endpoint, Headers: config.Headers, Insecure: config.Insecure, RequestTimeout: timeout}) + if err != nil { + return runtimeHistorySetup{}, err + } + setup.Exporter = exporter + // Independent queues keep an external Collector outage from delaying local history. + setup.Options = append(setup.Options, runtimeobs.WithExporter(exporter, options)) + } + return setup, nil +} + +func loadRuntimeHistoryConfig() (runtimeHistoryConfig, error) { var config runtimeHistoryConfig - decoder := json.NewDecoder(bytes.NewReader(raw)) - decoder.DisallowUnknownFields() - if decoder.Decode(&config) != nil || decoder.Decode(new(any)) != io.EOF { - return runtimeHistorySetup{}, errors.New("invalid Runtime history configuration") + if file := os.Getenv("AGENTS_API_RUNTIME_HISTORY_FILE"); file != "" { + raw, err := os.ReadFile(file) + if err != nil { + return config, errors.New("cannot read AGENTS_API_RUNTIME_HISTORY_FILE") + } + decoder := json.NewDecoder(bytes.NewReader(raw)) + decoder.DisallowUnknownFields() + if decoder.Decode(&config) != nil || decoder.Decode(new(any)) != io.EOF { + return config, errors.New("invalid Runtime history configuration") + } } if err := validateRuntimeHistoryConfig(config); err != nil { - return runtimeHistorySetup{}, err + return config, err } if config.QueueCapacity == 0 { config.QueueCapacity = defaultRuntimeHistoryQueueCapacity @@ -98,70 +109,49 @@ func runtimeHistory(ctx context.Context) (runtimeHistorySetup, error) { if config.TimeoutSeconds == 0 { config.TimeoutSeconds = defaultRuntimeHistoryTimeoutSeconds } - timeout := time.Duration(config.TimeoutSeconds) * time.Second - exporter, err := otlpexporter.New(ctx, otlpexporter.Config{ - Endpoint: config.Endpoint, Headers: config.Headers, Insecure: config.Insecure, RequestTimeout: timeout, - }) - if err != nil { - return runtimeHistorySetup{}, err + if config.SampleIntervalSeconds == 0 { + config.SampleIntervalSeconds = 30 } - var reader runtimeHistoryReader - if config.ClickHouse != nil { - clickhouse := *config.ClickHouse - if clickhouse.DialTimeoutSeconds == 0 { - clickhouse.DialTimeoutSeconds = defaultRuntimeHistoryDialTimeoutSeconds - } - if clickhouse.QueryTimeoutSeconds == 0 { - clickhouse.QueryTimeoutSeconds = defaultRuntimeHistoryQueryTimeoutSeconds - } - mode := runtimehistory.CollectionOnRead - interval := time.Duration(config.SampleIntervalSeconds) * time.Second - if interval > 0 { - mode = runtimehistory.CollectionPeriodic - } - minimumStep := 30 * time.Second - if interval > minimumStep { - minimumStep = interval - } - reader, err = openRuntimeHistoryReader(ctx, clickhousereader.Config{ - Address: clickhouse.Address, Database: clickhouse.Database, Username: clickhouse.Username, Password: clickhouse.Password, Secure: clickhouse.Secure, Insecure: clickhouse.Insecure, - DialTimeout: time.Duration(clickhouse.DialTimeoutSeconds) * time.Second, QueryTimeout: time.Duration(clickhouse.QueryTimeoutSeconds) * time.Second, - Capabilities: runtimehistory.Capabilities{ - CollectionMode: mode, SampleInterval: interval, Retention: 7 * 24 * time.Hour, MinimumStep: minimumStep, MaximumRange: 24 * time.Hour, - MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, - Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory, runtimehistory.MetricTokens}, - }, - }) - if err != nil { - closeCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - defer cancel() - closeRuntimeHistory(closeCtx, exporter) - return runtimeHistorySetup{}, err - } - } - return runtimeHistorySetup{ - Option: runtimeobs.WithExporter(exporter, runtimeobs.ExportOptions{QueueCapacity: config.QueueCapacity, Timeout: timeout}), - Exporter: exporter, - Reader: reader, - SampleInterval: time.Duration(config.SampleIntervalSeconds) * time.Second, - }, nil + return config, nil } func closeRuntimeHistory(ctx context.Context, exporter runtimeHistoryExporter) { - defer func() { - _ = recover() - }() + defer func() { _ = recover() }() _ = exporter.Close(ctx) } -func closeRuntimeHistoryReader(reader runtimeHistoryReader) { - defer func() { - _ = recover() - }() - _ = reader.Close() +// Retention also runs without active Runtimes. Each bounded pass has its own deadline. +func runHistoryCleanup(ctx context.Context, prune func(context.Context) error) { + ticker := time.NewTicker(time.Minute) + defer ticker.Stop() + for { + pruneCtx, cancel := context.WithTimeout(ctx, 2*time.Second) + _ = prune(pruneCtx) + cancel() + select { + case <-ctx.Done(): + return + case <-ticker.C: + } + } } func validateRuntimeHistoryConfig(config runtimeHistoryConfig) error { + if config.QueueCapacity < 0 || config.QueueCapacity > maxRuntimeHistoryQueueCapacity { + return errors.New("Runtime history queue_capacity is out of range") + } + if config.TimeoutSeconds < 0 || config.TimeoutSeconds > maxRuntimeHistoryTimeoutSeconds { + return errors.New("Runtime history timeout_seconds is out of range") + } + if config.SampleIntervalSeconds != 0 && (config.SampleIntervalSeconds < minRuntimeHistorySampleIntervalSeconds || config.SampleIntervalSeconds > maxRuntimeHistorySampleIntervalSeconds) { + return errors.New("Runtime history sample_interval_seconds is out of range") + } + if config.Endpoint == "" { + if config.Transport != "" || config.Insecure || len(config.Headers) != 0 { + return errors.New("Runtime history export options require an endpoint") + } + return nil + } if config.Transport != "otlp_http" { return errors.New("Runtime history transport must be otlp_http") } @@ -181,25 +171,6 @@ func validateRuntimeHistoryConfig(config runtimeHistoryConfig) error { default: return errors.New("Runtime history endpoint scheme must be https or explicit insecure http") } - if config.QueueCapacity < 0 || config.QueueCapacity > maxRuntimeHistoryQueueCapacity { - return errors.New("Runtime history queue_capacity is out of range") - } - if config.TimeoutSeconds < 0 || config.TimeoutSeconds > maxRuntimeHistoryTimeoutSeconds { - return errors.New("Runtime history timeout_seconds is out of range") - } - if config.SampleIntervalSeconds != 0 && (config.SampleIntervalSeconds < minRuntimeHistorySampleIntervalSeconds || config.SampleIntervalSeconds > maxRuntimeHistorySampleIntervalSeconds) { - return errors.New("Runtime history sample_interval_seconds is out of range") - } - if config.ClickHouse != nil { - clickhouse := config.ClickHouse - host, _, addressErr := net.SplitHostPort(clickhouse.Address) - if addressErr != nil || host == "" || clickhouse.Database == "" || clickhouse.Username == "" || clickhouse.Secure == clickhouse.Insecure || len(clickhouse.Database) > 128 || len(clickhouse.Username) > 128 || strings.ContainsAny(clickhouse.Database+clickhouse.Username, "\x00\r\n") { - return errors.New("Runtime history ClickHouse configuration is invalid") - } - if clickhouse.DialTimeoutSeconds < 0 || clickhouse.DialTimeoutSeconds > maxRuntimeHistoryTimeoutSeconds || clickhouse.QueryTimeoutSeconds < 0 || clickhouse.QueryTimeoutSeconds > maxRuntimeHistoryTimeoutSeconds { - return errors.New("Runtime history ClickHouse timeout is out of range") - } - } for key, value := range config.Headers { lower := strings.ToLower(key) if !httpguts.ValidHeaderFieldName(key) || !httpguts.ValidHeaderFieldValue(value) || lower == "host" || lower == "content-length" || lower == "content-type" || lower == "content-encoding" { diff --git a/services/agents-api/cmd/server/runtime_history_test.go b/services/agents-api/cmd/server/runtime_history_test.go index c72fe57a8..385a656ce 100644 --- a/services/agents-api/cmd/server/runtime_history_test.go +++ b/services/agents-api/cmd/server/runtime_history_test.go @@ -2,15 +2,13 @@ package main import ( "context" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" "os" "path/filepath" "strings" "testing" "time" - - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory/clickhousereader" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" ) type panicHistoryExporter struct{} @@ -18,72 +16,42 @@ type panicHistoryExporter struct{} func (panicHistoryExporter) Export(context.Context, runtimeobs.ExportRecord) error { return nil } func (panicHistoryExporter) Close(context.Context) error { panic("close") } -type fakeRuntimeHistoryReader struct { - capabilities runtimehistory.Capabilities - closed bool -} - -func (r *fakeRuntimeHistoryReader) Capabilities() runtimehistory.Capabilities { return r.capabilities } -func (r *fakeRuntimeHistoryReader) Query(context.Context, runtimehistory.Query) (runtimehistory.Result, error) { - return runtimehistory.Result{}, nil -} -func (r *fakeRuntimeHistoryReader) Close() error { r.closed = true; return nil } - -func TestRuntimeHistoryIsDisabledByDefault(t *testing.T) { +func TestRuntimeHistoryUsesCoreDatabaseByDefault(t *testing.T) { t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", "") - setup, err := runtimeHistory(t.Context()) - if err != nil || setup.Option != nil || setup.Exporter != nil || setup.Reader != nil || setup.SampleInterval != 0 { - t.Fatalf("disabled history created dependencies: option=%v exporter=%v reader=%v interval=%v err=%v", setup.Option != nil, setup.Exporter != nil, setup.Reader != nil, setup.SampleInterval, err) + for _, enabled := range []bool{true, false} { + setup, err := runtimeHistory(t.Context(), store.New(nil), enabled) + if err != nil { + t.Fatal(err) + } + if len(setup.Options) != 1 || setup.Exporter != nil || setup.Reader == nil || setup.Prune == nil { + t.Fatal("default history requires extra deployment") + } + capabilities := setup.Reader.Capabilities() + if capabilities.Durable() != enabled || capabilities.Retention != 7*24*time.Hour { + t.Fatalf("incorrect default capabilities: %+v", capabilities) + } + if enabled && setup.SampleInterval != 30*time.Second { + t.Fatal("default cadence missing") + } + if !enabled && setup.SampleInterval != 0 { + t.Fatal("sampler requires an execution owner") + } } } -func TestRuntimeHistoryWiresStrictClickHouseReaderWithoutExposingBackend(t *testing.T) { - file := filepath.Join(t.TempDir(), "runtime-history.json") - config := `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","sample_interval_seconds":30,"clickhouse":{"address":"clickhouse.example.test:9440","database":"runtime_history","username":"reader","password":"must-not-leak","secure":true,"dial_timeout_seconds":4,"query_timeout_seconds":6}}` - if err := os.WriteFile(file, []byte(config), 0600); err != nil { +func TestRuntimeHistoryOptionalExportAndSamplingConfiguration(t *testing.T) { + file := filepath.Join(t.TempDir(), "history.json") + if err := os.WriteFile(file, []byte(`{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Authorization":"Bearer private"},"sample_interval_seconds":60}`), 0600); err != nil { t.Fatal(err) } t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) - original := openRuntimeHistoryReader - defer func() { openRuntimeHistoryReader = original }() - var captured clickhousereader.Config - reader := &fakeRuntimeHistoryReader{} - openRuntimeHistoryReader = func(_ context.Context, value clickhousereader.Config) (runtimeHistoryReader, error) { - captured = value - reader.capabilities = value.Capabilities - return reader, nil - } - setup, err := runtimeHistory(t.Context()) - if err != nil || setup.Reader != reader { - t.Fatalf("ClickHouse Reader was not wired: reader=%v err=%v", setup.Reader != nil, err) - } - if captured.Address != "clickhouse.example.test:9440" || captured.Database != "runtime_history" || captured.Username != "reader" || !captured.Secure || captured.DialTimeout != 4*time.Second || captured.QueryTimeout != 6*time.Second || !captured.Capabilities.Durable() { - t.Fatalf("unexpected ClickHouse Reader configuration: %+v", captured) - } - closeRuntimeHistoryReader(setup.Reader) - if !reader.closed { - t.Fatal("ClickHouse Reader was not closed") - } - closeCtx, cancel := context.WithTimeout(t.Context(), time.Second) - defer cancel() - closeRuntimeHistory(closeCtx, setup.Exporter) -} - -func TestRuntimeHistoryLoadsStrictServerOnlyConfig(t *testing.T) { - file := filepath.Join(t.TempDir(), "runtime-history.json") - config := `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","headers":{"Authorization":"Bearer test-only"},"queue_capacity":12,"timeout_seconds":3,"sample_interval_seconds":30}` - if err := os.WriteFile(file, []byte(config), 0600); err != nil { + setup, err := runtimeHistory(t.Context(), store.New(nil), true) + if err != nil { t.Fatal(err) } - t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) - setup, err := runtimeHistory(t.Context()) - if err != nil || setup.Option == nil || setup.Exporter == nil || setup.SampleInterval != 30*time.Second { - t.Fatalf("valid history config was rejected: option=%v exporter=%v interval=%v err=%v", setup.Option != nil, setup.Exporter != nil, setup.SampleInterval, err) - } - closeCtx, cancel := context.WithTimeout(t.Context(), time.Second) - defer cancel() - if err := setup.Exporter.Close(closeCtx); err != nil { - t.Fatal(err) + defer closeRuntimeHistory(t.Context(), setup.Exporter) + if len(setup.Options) != 2 || setup.SampleInterval != time.Minute || setup.Reader.Capabilities().MinimumStep != time.Minute { + t.Fatal("export and local history are not independent") } } @@ -101,8 +69,8 @@ func TestRuntimeHistoryConfigFailsClosedWithoutLeakingSecrets(t *testing.T) { {name: "oversized timeout", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","timeout_seconds":31}`}, {name: "too frequent sampling", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","sample_interval_seconds":4}`}, {name: "oversized sampling interval", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","sample_interval_seconds":301}`}, - {name: "implicit insecure ClickHouse", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","clickhouse":{"address":"clickhouse.example.test:9000","database":"runtime_history","username":"reader","password":"must-not-leak"}}`}, - {name: "conflicting ClickHouse transport", config: `{"transport":"otlp_http","endpoint":"https://collector.example.test/v1/metrics","clickhouse":{"address":"clickhouse.example.test:9440","database":"runtime_history","username":"reader","secure":true,"insecure":true}}`}, + {name: "removed backend", config: `{"clickhouse":{"password":"must-not-leak"}}`}, + {name: "transport without endpoint", config: `{"transport":"otlp_http"}`}, } for _, test := range tests { t.Run(test.name, func(t *testing.T) { @@ -111,9 +79,9 @@ func TestRuntimeHistoryConfigFailsClosedWithoutLeakingSecrets(t *testing.T) { t.Fatal(err) } t.Setenv("AGENTS_API_RUNTIME_HISTORY_FILE", file) - setup, err := runtimeHistory(t.Context()) - if err == nil || setup.Option != nil || setup.Exporter != nil { - t.Fatalf("unsafe history config was accepted: option=%v exporter=%v err=%v", setup.Option != nil, setup.Exporter != nil, err) + _, err := loadRuntimeHistoryConfig() + if err == nil { + t.Fatal("unsafe history configuration accepted") } if strings.Contains(err.Error(), "must-not-leak") { t.Fatalf("history error leaked config content: %v", err) @@ -126,9 +94,6 @@ func TestRuntimeHistoryAllowsExplicitLocalHTTPCollector(t *testing.T) { config := runtimeHistoryConfig{ Transport: "otlp_http", Endpoint: "http://127.0.0.1:4318/v1/metrics", Insecure: true, Headers: map[string]string{"X-Scope-OrgID": "operator-history"}, - ClickHouse: &runtimeHistoryClickHouseConfig{ - Address: "127.0.0.1:9000", Database: "runtime_history", Username: "reader", Insecure: true, - }, } if err := validateRuntimeHistoryConfig(config); err != nil { t.Fatal(err) diff --git a/services/agents-api/internal/runtimeobs/exporter.go b/services/agents-api/internal/runtimeobs/exporter.go index b224a1207..33ce5de63 100644 --- a/services/agents-api/internal/runtimeobs/exporter.go +++ b/services/agents-api/internal/runtimeobs/exporter.go @@ -48,11 +48,15 @@ type ExportOptions struct { type ServiceOption func(*serviceOptions) error -type serviceOptions struct { +type exporterConfig struct { exporter Exporter exportOptions ExportOptions } +type serviceOptions struct { + exporters []exporterConfig +} + // WithExporter enables best-effort history export. Queue saturation drops the // newest handoff; it never changes the Runtime observation result. func WithExporter(exporter Exporter, options ExportOptions) ServiceOption { @@ -66,8 +70,7 @@ func WithExporter(exporter Exporter, options ExportOptions) ServiceOption { if options.Timeout < 0 { return errors.New("Runtime observation export timeout cannot be negative") } - config.exporter = exporter - config.exportOptions = options + config.exporters = append(config.exporters, exporterConfig{exporter: exporter, exportOptions: options}) return nil } } diff --git a/services/agents-api/internal/runtimeobs/exporter_independence_test.go b/services/agents-api/internal/runtimeobs/exporter_independence_test.go new file mode 100644 index 000000000..add9f2276 --- /dev/null +++ b/services/agents-api/internal/runtimeobs/exporter_independence_test.go @@ -0,0 +1,33 @@ +package runtimeobs + +import ( + "testing" + "time" +) + +func TestExporterOutageDoesNotDelayOtherDestinations(t *testing.T) { + now := time.Now() + target := Target{EnvironmentID: "environment", Mode: ModeManaged, Instance: Instance{AllocationID: "allocation", ProviderKey: "provider", AllocationState: "running"}} + blocked := &gatedExporter{started: make(chan struct{}, 1), release: make(chan struct{})} + records := make(chan ExportRecord, 1) + service, err := NewService(fixedResolver{target: target}, map[string]Source{"provider": &fixedSource{sample: Sample{ObservedAt: now}}}, + WithExporter(blocked, ExportOptions{QueueCapacity: 1, Timeout: time.Second}), + WithExporter(channelExporter{records: records}, ExportOptions{QueueCapacity: 1, Timeout: time.Second})) + if err != nil { + t.Fatal(err) + } + defer func() { close(blocked.release); _ = service.Close(t.Context()) }() + if _, err := service.ObserveSession(t.Context(), "tenant", "session"); err != nil { + t.Fatal(err) + } + select { + case <-blocked.started: + case <-time.After(time.Second): + t.Fatal("blocked exporter never started") + } + select { + case <-records: + case <-time.After(time.Second): + t.Fatal("local history waited for external exporter") + } +} diff --git a/services/agents-api/internal/runtimeobs/service.go b/services/agents-api/internal/runtimeobs/service.go index b414fbc5f..4f86104b1 100644 --- a/services/agents-api/internal/runtimeobs/service.go +++ b/services/agents-api/internal/runtimeobs/service.go @@ -31,7 +31,7 @@ type Service struct { resolver TargetResolver sources map[string]Source now func() time.Time - exports *exportDispatcher + exports []*exportDispatcher } func NewService(resolver TargetResolver, sources map[string]Source, options ...ServiceOption) (*Service, error) { @@ -55,8 +55,8 @@ func NewService(resolver TargetResolver, sources map[string]Source, options ...S copySources[key] = source } service := &Service{resolver: resolver, sources: copySources, now: time.Now} - if config.exporter != nil { - service.exports = newExportDispatcher(config.exporter, config.exportOptions) + for _, export := range config.exporters { + service.exports = append(service.exports, newExportDispatcher(export.exporter, export.exportOptions)) } return service, nil } @@ -64,18 +64,19 @@ func NewService(resolver TargetResolver, sources map[string]Source, options ...S // Close drains pending history handoffs within ctx. Current-observation callers // may keep using a Service without an exporter; Close is then a no-op. func (s *Service) Close(ctx context.Context) error { - if s.exports == nil { - return nil + var result error + for _, exporter := range s.exports { + result = errors.Join(result, exporter.close(ctx)) } - return s.exports.close(ctx) + return result } func (s *Service) finish(ctx context.Context, observation Observation, source CollectionSource, owner OwnershipChecker) (Observation, error) { if err := checkHistoryOwnership(ctx, owner); err != nil { return Observation{}, err } - if s.exports != nil { - s.exports.enqueue(exportRecord(observation, source)) + for _, exporter := range s.exports { + exporter.enqueue(exportRecord(observation, source)) } return observation, nil } diff --git a/services/agents-api/runtime-history/README.md b/services/agents-api/runtime-history/README.md new file mode 100644 index 000000000..6dc77e6ed --- /dev/null +++ b/services/agents-api/runtime-history/README.md @@ -0,0 +1,56 @@ +# Runtime history + +Core stores recent Runtime observations in its existing PostgreSQL database. +No additional database or Collector is required. Apply Core migrations normally; +the server enables 30-second periodic sampling when an execution worker is present. +Only the execution lease owner samples, using the same provider-neutral current +observation sources. Observation never changes Runtime lifecycle. + +History retains seven days and queries at most 24 hours. Writes use a bounded +asynchronous queue; failures or overflow create gaps, never execution failures. +CPU and memory unknowns remain null. CPU counter deltas never bridge compute +restarts or counter resets. Token snapshots come from canonical Session Usage; +the history table does not become billing or execution authority. + +Tenant, Session and Environment scope are mandatory on reads and writes. The +schema stores sanitized measurements and Core identifiers, not credentials, +provider receipts, native process identifiers, paths or raw provider errors. +Nanosecond identity timestamps are lossless integers. Retention cleanup runs in +bounded batches once per minute even when there are no active Runtimes. Queries +exclude expired records before physical cleanup completes. + +## Optional configuration + +Set `AGENTS_API_RUNTIME_HISTORY_FILE` to an absolute server-only JSON file to change +sampling or export to an existing OTLP receiver. Sampling-only example: + +```json +{"sample_interval_seconds": 60} +``` + +Sampling accepts 5–300 seconds; omitted or zero selects 30. Queue capacity defaults +to 256 (maximum 4096); write/export timeout defaults to two seconds (maximum30). +A server without an execution worker advertises on-read collection and does not +claim periodic coverage. Retained queries remain available through Core. + +Optional external export: + +```json +{ + "transport": "otlp_http", + "endpoint": "https://collector.example.com/v1/metrics", + "headers": {"Authorization": "Bearer operator-secret"}, + "sample_interval_seconds": 30, + "queue_capacity": 256, + "timeout_seconds": 2 +} +``` + +HTTP requires explicit `insecure: true`. Transport credentials stay server-only; +the browser uses only Core. External export has a separate bounded queue, so its +outage cannot starve local history. Do not commit this private configuration. + +Apply the standard `make check` gate with a dedicated PostgreSQL database. Real +Runtime acceptance must include provider observations during real model execution, +tenant isolation and retained history across Core restart; injected rows alone do +not qualify provider sampling. diff --git a/services/agents-api/runtime-history/clickhouse/001_runtime_history.sql b/services/agents-api/runtime-history/clickhouse/001_runtime_history.sql deleted file mode 100644 index e698f71db..000000000 --- a/services/agents-api/runtime-history/clickhouse/001_runtime_history.sql +++ /dev/null @@ -1,90 +0,0 @@ --- Parsar Core Runtime history projection for the OpenTelemetry Collector --- ClickHouse exporter generic metrics_gauge and metrics_sum tables. --- Apply this file only after those generic tables exist. - -CREATE TABLE IF NOT EXISTS runtime_history_metrics -( - timestamp DateTime64(9) CODEC(ZSTD(1)), - tenant_id String CODEC(ZSTD(1)), - session_id String CODEC(ZSTD(1)), - environment_id String CODEC(ZSTD(1)), - allocation_id String CODEC(ZSTD(1)), - started_at_unix_nano Nullable(Int64) CODEC(ZSTD(1)), - provider_type LowCardinality(String) CODEC(ZSTD(1)), - status LowCardinality(String) CODEC(ZSTD(1)), - collection_source LowCardinality(String) CODEC(ZSTD(1)), - resolved_at_unix_nano Int64 CODEC(ZSTD(1)), - observed_at_unix_nano Nullable(Int64) CODEC(ZSTD(1)), - metric_name LowCardinality(String) CODEC(ZSTD(1)), - value Float64 CODEC(ZSTD(1)) -) -ENGINE = MergeTree -PARTITION BY toDate(timestamp) -ORDER BY -( - tenant_id, - session_id, - environment_id, - collection_source, - resolved_at_unix_nano, - allocation_id, - ifNull(started_at_unix_nano, -1), - metric_name -) -TTL toDateTime(timestamp) + INTERVAL 7 DAY; - -CREATE MATERIALIZED VIEW IF NOT EXISTS runtime_history_metrics_gauge_mv -TO runtime_history_metrics AS -SELECT - fromUnixTimestamp64Nano(toInt64(Attributes['agents.runtime.resolved_at_unix_nano'])) AS timestamp, - Attributes['agents.tenant.id'] AS tenant_id, - Attributes['agents.session.id'] AS session_id, - Attributes['agents.environment.id'] AS environment_id, - Attributes['agents.runtime.allocation.id'] AS allocation_id, - toInt64OrNull(Attributes['agents.runtime.compute.started_at_unix_nano']) AS started_at_unix_nano, - Attributes['agents.runtime.provider.type'] AS provider_type, - Attributes['agents.runtime.status'] AS status, - Attributes['agents.runtime.collection.source'] AS collection_source, - toInt64(Attributes['agents.runtime.resolved_at_unix_nano']) AS resolved_at_unix_nano, - toInt64OrNull(Attributes['agents.runtime.observed_at_unix_nano']) AS observed_at_unix_nano, - MetricName AS metric_name, - toFloat64(Value) AS value -FROM metrics_gauge -WHERE MetricName IN -( - 'agents.runtime.cpu.capacity', - 'agents.runtime.memory.usage', - 'agents.runtime.memory.limit' -) -AND Attributes['agents.tenant.id'] != '' -AND Attributes['agents.session.id'] != '' -AND Attributes['agents.environment.id'] != '' -AND Attributes['agents.runtime.collection.source'] IN ('on_read', 'periodic') -AND toInt64OrNull(Attributes['agents.runtime.resolved_at_unix_nano']) IS NOT NULL; -CREATE MATERIALIZED VIEW IF NOT EXISTS runtime_history_metrics_sum_mv -TO runtime_history_metrics AS -SELECT - fromUnixTimestamp64Nano(toInt64(Attributes['agents.runtime.resolved_at_unix_nano'])) AS timestamp, - Attributes['agents.tenant.id'] AS tenant_id, - Attributes['agents.session.id'] AS session_id, - Attributes['agents.environment.id'] AS environment_id, - Attributes['agents.runtime.allocation.id'] AS allocation_id, - toInt64OrNull(Attributes['agents.runtime.compute.started_at_unix_nano']) AS started_at_unix_nano, - Attributes['agents.runtime.provider.type'] AS provider_type, - Attributes['agents.runtime.status'] AS status, - Attributes['agents.runtime.collection.source'] AS collection_source, - toInt64(Attributes['agents.runtime.resolved_at_unix_nano']) AS resolved_at_unix_nano, - toInt64OrNull(Attributes['agents.runtime.observed_at_unix_nano']) AS observed_at_unix_nano, - MetricName AS metric_name, - toFloat64(Value) AS value -FROM metrics_sum -WHERE MetricName IN -( - 'agents.runtime.sample', - 'agents.runtime.cpu.usage' -) -AND Attributes['agents.tenant.id'] != '' -AND Attributes['agents.session.id'] != '' -AND Attributes['agents.environment.id'] != '' -AND Attributes['agents.runtime.collection.source'] IN ('on_read', 'periodic') -AND toInt64OrNull(Attributes['agents.runtime.resolved_at_unix_nano']) IS NOT NULL; diff --git a/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql b/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql deleted file mode 100644 index f2e2d5f9c..000000000 --- a/services/agents-api/runtime-history/clickhouse/002_session_token_usage.sql +++ /dev/null @@ -1,31 +0,0 @@ --- Session-scoped cumulative token counters. Apply after 001_runtime_history.sql. --- A separate materialized view keeps usage independent of provider Runtime --- incarnation fields while reusing the bounded runtime_history_metrics table. - -CREATE MATERIALIZED VIEW IF NOT EXISTS runtime_history_session_usage_gauge_mv -TO runtime_history_metrics AS -SELECT - fromUnixTimestamp64Nano(toInt64(Attributes['agents.runtime.resolved_at_unix_nano'])) AS timestamp, - Attributes['agents.tenant.id'] AS tenant_id, - Attributes['agents.session.id'] AS session_id, - Attributes['agents.environment.id'] AS environment_id, - Attributes['agents.runtime.allocation.id'] AS allocation_id, - toInt64OrNull(Attributes['agents.runtime.compute.started_at_unix_nano']) AS started_at_unix_nano, - Attributes['agents.runtime.provider.type'] AS provider_type, - Attributes['agents.runtime.status'] AS status, - Attributes['agents.runtime.collection.source'] AS collection_source, - toInt64(Attributes['agents.runtime.resolved_at_unix_nano']) AS resolved_at_unix_nano, - toInt64OrNull(Attributes['agents.runtime.observed_at_unix_nano']) AS observed_at_unix_nano, - MetricName AS metric_name, - toFloat64(Value) AS value -FROM metrics_gauge -WHERE MetricName IN -( - 'agents.session.tokens.input', - 'agents.session.tokens.output' -) -AND Attributes['agents.tenant.id'] != '' -AND Attributes['agents.session.id'] != '' -AND Attributes['agents.environment.id'] != '' -AND Attributes['agents.runtime.collection.source'] IN ('on_read', 'periodic') -AND toInt64OrNull(Attributes['agents.runtime.resolved_at_unix_nano']) IS NOT NULL; diff --git a/services/agents-api/runtime-history/clickhouse/README.md b/services/agents-api/runtime-history/clickhouse/README.md deleted file mode 100644 index 6eabe4ec5..000000000 --- a/services/agents-api/runtime-history/clickhouse/README.md +++ /dev/null @@ -1,94 +0,0 @@ -# ClickHouse Runtime history reference deployment - -This optional deployment turns periodic, provider-neutral Runtime observations -into the Durable history source exposed by Agents Core. ClickHouse remains an -operator-owned read model. It is not required for execution, current observations, -or lifecycle decisions. - -## Security and ownership - -- The Collector writer may insert into the generic `metrics_gauge` and - `metrics_sum` tables. ClickHouse executes materialized-view queries under the - inserting identity, so that writer also needs `SELECT` on only the three - source columns consumed by each reference view. -- The Agents API reader should receive `SELECT` on - `runtime_history_metrics` only. It does not need access to generic telemetry, - schema mutation, or another tenant selector. -- Every Reader query contains the authenticated tenant plus resolved Session and - Environment identity, an exclusive time bound, and - `collection_source = 'periodic'`. -- Web calls only the public Agents API. Do not expose ClickHouse or Collector - credentials to a browser. -- The reference table has a seven-day TTL. The server advertises that same - retention and permits at most a 24-hour query range, 1,000 buckets per series, - 64 series, and 10,000 total returned points. - -## Install - -1. Run an OpenTelemetry Collector with - [`otel-collector.example.yaml`](otel-collector.example.yaml). Its ClickHouse - exporter creates the generic metric tables. Keep the OTLP receiver private or - authenticate it at the network/proxy boundary. -2. After `metrics_gauge` and `metrics_sum` exist, apply - [`001_runtime_history.sql`](001_runtime_history.sql), then - [`002_session_token_usage.sql`](002_session_token_usage.sql), to the same - database. - The materialized views retain only the five Runtime values and two canonical - Session token counters needed by the history API; - provider receipts, native container identities, paths, and raw errors never - enter the projection. - Grant the Collector identity the minimum source-column reads required when - those views run: - - ```sql - GRANT SELECT(Attributes, MetricName, Value) - ON runtime_history.metrics_gauge TO agents_runtime_writer; - GRANT SELECT(Attributes, MetricName, Value) - ON runtime_history.metrics_sum TO agents_runtime_writer; - ``` - - Keep the existing `INSERT` grants on both generic tables. No broader table - read or projection read is required by the writer. -3. Create a read-only ClickHouse account for Core: - - ```sql - GRANT SELECT ON runtime_history.runtime_history_metrics TO agents_runtime_reader; - ``` - -4. Store the following JSON outside the repository with mode `0600`, then point - `AGENTS_API_RUNTIME_HISTORY_FILE` at its absolute path: - - ```json - { - "transport": "otlp_http", - "endpoint": "https://collector.example.com/v1/metrics", - "headers": {"Authorization": "Bearer operator-managed-secret"}, - "queue_capacity": 256, - "timeout_seconds": 2, - "sample_interval_seconds": 30, - "clickhouse": { - "address": "clickhouse.example.com:9440", - "database": "runtime_history", - "username": "agents_runtime_reader", - "password": "operator-managed-secret", - "secure": true, - "dial_timeout_seconds": 5, - "query_timeout_seconds": 10 - } - } - ``` - -When `clickhouse` is omitted, export and background sampling continue but Durable -history remains unavailable. When `sample_interval_seconds` is omitted, a -configured Reader is advertised as on-read only and Web must not call it Durable. -Startup validates configuration without echoing credentials. The connection is -opened lazily: a ClickHouse outage returns a sanitized `503` for history reads but -does not stop Core execution or current observations. - -Exactly one of `secure` or `insecure` is required. Production TCP connections -should use `secure: true`; plaintext local validation must opt in explicitly with -`insecure: true` so omitting TLS never silently transmits credentials or history. - -The specialized projection intentionally stores both `on_read` and `periodic` -records for operator inspection. Public Durable queries use only `periodic` -records so Dashboard traffic cannot inflate or fabricate cadence coverage. diff --git a/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml b/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml deleted file mode 100644 index f21915892..000000000 --- a/services/agents-api/runtime-history/clickhouse/otel-collector.example.yaml +++ /dev/null @@ -1,44 +0,0 @@ -receivers: - otlp: - protocols: - grpc: - endpoint: 0.0.0.0:4317 - http: - endpoint: 0.0.0.0:4318 - -processors: - filter/runtime_history: - metrics: - include: - match_type: regexp - metric_names: - - '^agents\.runtime\..*$' - batch/runtime_history: - timeout: 5s - send_batch_size: 5000 - -exporters: - clickhouse/runtime_history: - endpoint: ${env:CLICKHOUSE_ENDPOINT} - database: ${env:CLICKHOUSE_DATABASE} - username: ${env:CLICKHOUSE_WRITER_USERNAME} - password: ${env:CLICKHOUSE_WRITER_PASSWORD} - async_insert: true - create_schema: true - metrics_tables: - gauge: - name: metrics_gauge - sum: - name: metrics_sum - -extensions: - health_check: - endpoint: 0.0.0.0:13133 - -service: - extensions: [health_check] - pipelines: - metrics/runtime_history: - receivers: [otlp] - processors: [filter/runtime_history, batch/runtime_history] - exporters: [clickhouse/runtime_history] From 094b0dfb8edf094eace373f32ee800d527b977a6 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:44:31 +0800 Subject: [PATCH 43/51] Persist bounded Runtime history in existing PostgreSQL store --- .../internal/db/queries/runtime_history.sql | 29 + .../agents-api/internal/db/sqlc/models.go | 18 + .../internal/db/sqlc/runtime_history.sql.go | 144 ++++ .../clickhousereader/acceptance_test.go | 165 ---- .../runtimehistory/clickhousereader/reader.go | 726 ------------------ .../clickhousereader/reader_test.go | 440 ----------- .../postgresreader/aggregate.go | 318 ++++++++ .../postgresreader/aggregate_test.go | 94 +++ .../runtimehistory/postgresreader/exporter.go | 97 +++ .../runtimehistory/postgresreader/reader.go | 156 ++++ .../postgresreader/reader_test.go | 195 +++++ .../internal/store/runtime_history.go | 132 ++++ .../migrations/000056_runtime_history.sql | 36 + 13 files changed, 1219 insertions(+), 1331 deletions(-) create mode 100644 services/agents-api/internal/db/queries/runtime_history.sql create mode 100644 services/agents-api/internal/db/sqlc/runtime_history.sql.go delete mode 100644 services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go delete mode 100644 services/agents-api/internal/runtimehistory/clickhousereader/reader.go delete mode 100644 services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go create mode 100644 services/agents-api/internal/runtimehistory/postgresreader/aggregate.go create mode 100644 services/agents-api/internal/runtimehistory/postgresreader/aggregate_test.go create mode 100644 services/agents-api/internal/runtimehistory/postgresreader/exporter.go create mode 100644 services/agents-api/internal/runtimehistory/postgresreader/reader.go create mode 100644 services/agents-api/internal/runtimehistory/postgresreader/reader_test.go create mode 100644 services/agents-api/internal/store/runtime_history.go create mode 100644 services/agents-api/migrations/000056_runtime_history.sql diff --git a/services/agents-api/internal/db/queries/runtime_history.sql b/services/agents-api/internal/db/queries/runtime_history.sql new file mode 100644 index 000000000..d0c019024 --- /dev/null +++ b/services/agents-api/internal/db/queries/runtime_history.sql @@ -0,0 +1,29 @@ +-- name: InsertRuntimeHistorySample :exec +INSERT INTO runtime_history_samples ( + tenant_id, session_id, environment_id, resolved_at_ns, allocation_id, provider_type, + status, observed_at_ns, started_at_ns, cpu_usage_seconds, cpu_capacity_cores, + memory_usage_bytes, memory_limit_bytes, input_tokens, output_tokens +) +SELECT s.tenant_id, s.id, e.id, sqlc.arg(resolved_at_ns), sqlc.narg(allocation_id), sqlc.arg(provider_type), + sqlc.arg(status), sqlc.narg(observed_at_ns), sqlc.narg(started_at_ns), sqlc.narg(cpu_usage_seconds), sqlc.narg(cpu_capacity_cores), + sqlc.narg(memory_usage_bytes), sqlc.narg(memory_limit_bytes), sqlc.narg(input_tokens), sqlc.narg(output_tokens) +FROM sessions s JOIN environments e ON e.session_id = s.id +WHERE s.tenant_id = sqlc.arg(tenant_id) AND s.id = sqlc.arg(session_id) AND e.id = sqlc.arg(environment_id) +ON CONFLICT (tenant_id, session_id, environment_id, resolved_at_ns) DO NOTHING; + +-- name: ListRuntimeHistorySamples :many +SELECT * FROM runtime_history_samples +WHERE tenant_id = sqlc.arg(tenant_id) AND session_id = sqlc.arg(session_id) AND environment_id = sqlc.arg(environment_id) + AND resolved_at_ns >= sqlc.arg(start_ns) AND resolved_at_ns < sqlc.arg(end_ns) +ORDER BY resolved_at_ns +LIMIT sqlc.arg(row_limit); + +-- name: PruneRuntimeHistorySamples :execrows +WITH expired AS ( + SELECT p.tenant_id, p.session_id, p.environment_id, p.resolved_at_ns FROM runtime_history_samples p + WHERE p.resolved_at_ns < sqlc.arg(before_ns) + ORDER BY p.resolved_at_ns LIMIT 256 FOR UPDATE SKIP LOCKED +) +DELETE FROM runtime_history_samples h USING expired e +WHERE h.tenant_id = e.tenant_id AND h.session_id = e.session_id + AND h.environment_id = e.environment_id AND h.resolved_at_ns = e.resolved_at_ns; diff --git a/services/agents-api/internal/db/sqlc/models.go b/services/agents-api/internal/db/sqlc/models.go index 7b5a834e0..96d2f482d 100644 --- a/services/agents-api/internal/db/sqlc/models.go +++ b/services/agents-api/internal/db/sqlc/models.go @@ -177,6 +177,24 @@ type RuntimeDeviceAuthority struct { CredentialHash string `json:"credential_hash"` } +type RuntimeHistorySample struct { + TenantID pgtype.UUID `json:"tenant_id"` + SessionID pgtype.UUID `json:"session_id"` + EnvironmentID pgtype.UUID `json:"environment_id"` + ResolvedAtNs int64 `json:"resolved_at_ns"` + AllocationID pgtype.UUID `json:"allocation_id"` + ProviderType string `json:"provider_type"` + Status string `json:"status"` + ObservedAtNs pgtype.Int8 `json:"observed_at_ns"` + StartedAtNs pgtype.Int8 `json:"started_at_ns"` + CpuUsageSeconds pgtype.Float8 `json:"cpu_usage_seconds"` + CpuCapacityCores pgtype.Float8 `json:"cpu_capacity_cores"` + MemoryUsageBytes pgtype.Int8 `json:"memory_usage_bytes"` + MemoryLimitBytes pgtype.Int8 `json:"memory_limit_bytes"` + InputTokens pgtype.Int8 `json:"input_tokens"` + OutputTokens pgtype.Int8 `json:"output_tokens"` +} + type Session struct { ID pgtype.UUID `json:"id"` TenantID pgtype.UUID `json:"tenant_id"` diff --git a/services/agents-api/internal/db/sqlc/runtime_history.sql.go b/services/agents-api/internal/db/sqlc/runtime_history.sql.go new file mode 100644 index 000000000..1c6a84eac --- /dev/null +++ b/services/agents-api/internal/db/sqlc/runtime_history.sql.go @@ -0,0 +1,144 @@ +// Code generated by sqlc. DO NOT EDIT. +// versions: +// sqlc v1.29.0 +// source: runtime_history.sql + +package sqlc + +import ( + "context" + + "github.com/jackc/pgx/v5/pgtype" +) + +const insertRuntimeHistorySample = `-- name: InsertRuntimeHistorySample :exec +INSERT INTO runtime_history_samples ( + tenant_id, session_id, environment_id, resolved_at_ns, allocation_id, provider_type, + status, observed_at_ns, started_at_ns, cpu_usage_seconds, cpu_capacity_cores, + memory_usage_bytes, memory_limit_bytes, input_tokens, output_tokens +) +SELECT s.tenant_id, s.id, e.id, $1, $2, $3, + $4, $5, $6, $7, $8, + $9, $10, $11, $12 +FROM sessions s JOIN environments e ON e.session_id = s.id +WHERE s.tenant_id = $13 AND s.id = $14 AND e.id = $15 +ON CONFLICT (tenant_id, session_id, environment_id, resolved_at_ns) DO NOTHING +` + +type InsertRuntimeHistorySampleParams struct { + ResolvedAtNs int64 `json:"resolved_at_ns"` + AllocationID pgtype.UUID `json:"allocation_id"` + ProviderType string `json:"provider_type"` + Status string `json:"status"` + ObservedAtNs pgtype.Int8 `json:"observed_at_ns"` + StartedAtNs pgtype.Int8 `json:"started_at_ns"` + CpuUsageSeconds pgtype.Float8 `json:"cpu_usage_seconds"` + CpuCapacityCores pgtype.Float8 `json:"cpu_capacity_cores"` + MemoryUsageBytes pgtype.Int8 `json:"memory_usage_bytes"` + MemoryLimitBytes pgtype.Int8 `json:"memory_limit_bytes"` + InputTokens pgtype.Int8 `json:"input_tokens"` + OutputTokens pgtype.Int8 `json:"output_tokens"` + TenantID pgtype.UUID `json:"tenant_id"` + SessionID pgtype.UUID `json:"session_id"` + EnvironmentID pgtype.UUID `json:"environment_id"` +} + +func (q *Queries) InsertRuntimeHistorySample(ctx context.Context, arg InsertRuntimeHistorySampleParams) error { + _, err := q.db.Exec(ctx, insertRuntimeHistorySample, + arg.ResolvedAtNs, + arg.AllocationID, + arg.ProviderType, + arg.Status, + arg.ObservedAtNs, + arg.StartedAtNs, + arg.CpuUsageSeconds, + arg.CpuCapacityCores, + arg.MemoryUsageBytes, + arg.MemoryLimitBytes, + arg.InputTokens, + arg.OutputTokens, + arg.TenantID, + arg.SessionID, + arg.EnvironmentID, + ) + return err +} + +const listRuntimeHistorySamples = `-- name: ListRuntimeHistorySamples :many +SELECT tenant_id, session_id, environment_id, resolved_at_ns, allocation_id, provider_type, status, observed_at_ns, started_at_ns, cpu_usage_seconds, cpu_capacity_cores, memory_usage_bytes, memory_limit_bytes, input_tokens, output_tokens FROM runtime_history_samples +WHERE tenant_id = $1 AND session_id = $2 AND environment_id = $3 + AND resolved_at_ns >= $4 AND resolved_at_ns < $5 +ORDER BY resolved_at_ns +LIMIT $6 +` + +type ListRuntimeHistorySamplesParams struct { + TenantID pgtype.UUID `json:"tenant_id"` + SessionID pgtype.UUID `json:"session_id"` + EnvironmentID pgtype.UUID `json:"environment_id"` + StartNs int64 `json:"start_ns"` + EndNs int64 `json:"end_ns"` + RowLimit int32 `json:"row_limit"` +} + +func (q *Queries) ListRuntimeHistorySamples(ctx context.Context, arg ListRuntimeHistorySamplesParams) ([]RuntimeHistorySample, error) { + rows, err := q.db.Query(ctx, listRuntimeHistorySamples, + arg.TenantID, + arg.SessionID, + arg.EnvironmentID, + arg.StartNs, + arg.EndNs, + arg.RowLimit, + ) + if err != nil { + return nil, err + } + defer rows.Close() + items := []RuntimeHistorySample{} + for rows.Next() { + var i RuntimeHistorySample + if err := rows.Scan( + &i.TenantID, + &i.SessionID, + &i.EnvironmentID, + &i.ResolvedAtNs, + &i.AllocationID, + &i.ProviderType, + &i.Status, + &i.ObservedAtNs, + &i.StartedAtNs, + &i.CpuUsageSeconds, + &i.CpuCapacityCores, + &i.MemoryUsageBytes, + &i.MemoryLimitBytes, + &i.InputTokens, + &i.OutputTokens, + ); err != nil { + return nil, err + } + items = append(items, i) + } + if err := rows.Err(); err != nil { + return nil, err + } + return items, nil +} + +const pruneRuntimeHistorySamples = `-- name: PruneRuntimeHistorySamples :execrows +WITH expired AS ( + SELECT p.tenant_id, p.session_id, p.environment_id, p.resolved_at_ns FROM runtime_history_samples p + WHERE p.resolved_at_ns < $1 + ORDER BY p.resolved_at_ns LIMIT 256 FOR UPDATE SKIP LOCKED +) +DELETE FROM runtime_history_samples h USING expired e +WHERE h.tenant_id = e.tenant_id AND h.session_id = e.session_id + AND h.environment_id = e.environment_id AND h.resolved_at_ns = e.resolved_at_ns +` + +func (q *Queries) PruneRuntimeHistorySamples(ctx context.Context, beforeNs int64) (int64, error) { + result, err := q.db.Exec(ctx, pruneRuntimeHistorySamples, beforeNs) + if err != nil { + return 0, err + } + return result.RowsAffected(), nil +} diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go b/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go deleted file mode 100644 index b4d38e831..000000000 --- a/services/agents-api/internal/runtimehistory/clickhousereader/acceptance_test.go +++ /dev/null @@ -1,165 +0,0 @@ -package clickhousereader - -import ( - "context" - "math" - "os" - "testing" - "time" - - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" - "github.com/google/uuid" -) - -func TestClickHouseCollectorAcceptance(t *testing.T) { - address := os.Getenv("PARSAR_RUNTIME_HISTORY_ACCEPTANCE_CLICKHOUSE_ADDRESS") - otlpEndpoint := os.Getenv("PARSAR_RUNTIME_HISTORY_ACCEPTANCE_OTLP_ENDPOINT") - if address == "" || otlpEndpoint == "" { - t.Skip("real ClickHouse/Collector acceptance is not configured") - } - database := envOr("PARSAR_RUNTIME_HISTORY_ACCEPTANCE_CLICKHOUSE_DATABASE", "runtime_history") - username := envOr("PARSAR_RUNTIME_HISTORY_ACCEPTANCE_CLICKHOUSE_USERNAME", "default") - password := os.Getenv("PARSAR_RUNTIME_HISTORY_ACCEPTANCE_CLICKHOUSE_PASSWORD") - - exporter, err := otlpexporter.New(t.Context(), otlpexporter.Config{ - Endpoint: otlpEndpoint, Insecure: true, RequestTimeout: 5 * time.Second, - }) - if err != nil { - t.Fatal(err) - } - t.Cleanup(func() { - ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - defer cancel() - _ = exporter.Close(ctx) - }) - - end := time.Now().UTC().Truncate(time.Second) - start := end.Add(-3 * time.Minute) - tenantID := uuid.NewString() - otherTenantID := uuid.NewString() - sessionID := uuid.NewString() - environmentID := uuid.NewString() - allocationID := uuid.NewString() - firstStartedAt := start.Add(5*time.Second + 123*time.Nanosecond) - secondStartedAt := start.Add(95*time.Second + 456*time.Nanosecond) - onRead := acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID, firstStartedAt, start.Add(80*time.Second), 120, 777) - onRead.CollectionSource = runtimeobs.CollectionSourceOnRead - - records := []runtimeobs.ExportRecord{ - acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID, firstStartedAt, start.Add(30*time.Second), 0, 100), - acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID, firstStartedAt, start.Add(60*time.Second), 30, 120), - acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID, secondStartedAt, start.Add(110*time.Second), 0, 200), - acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID, secondStartedAt, start.Add(140*time.Second), 10, 240), - { - TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID, - Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusUnavailable, Reason: "sample_timeout", - CollectionSource: runtimeobs.CollectionSourcePeriodic, ResolvedAt: start.Add(155 * time.Second), - }, - onRead, - acceptanceObservedRecord(otherTenantID, sessionID, environmentID, allocationID, firstStartedAt, start.Add(70*time.Second), 90, 999), - } - for _, record := range records { - if err := exporter.Export(t.Context(), record); err != nil { - t.Fatal(err) - } - } - - capabilities := testCapabilities() - open := func() *Reader { - reader, err := Open(t.Context(), Config{ - Address: address, Database: database, Username: username, Password: password, Insecure: true, - DialTimeout: 5 * time.Second, QueryTimeout: 5 * time.Second, Capabilities: capabilities, - }) - if err != nil { - t.Fatal(err) - } - return reader - } - query := runtimehistory.Query{ - Scope: runtimehistory.Scope{TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID}, - Start: start, End: end, Step: 30 * time.Second, Retention: capabilities.Retention, - MaxPoints: 6, MaximumSeries: capabilities.MaximumSeries, MaximumTotalPoints: capabilities.MaximumTotalPoints, - } - reader := open() - var result runtimehistory.Result - deadline := time.Now().Add(20 * time.Second) - for { - result, err = reader.Query(t.Context(), query) - if err == nil && coverageCount(result.Coverage) == 5 && len(result.Series) == 1 { - break - } - if time.Now().After(deadline) { - t.Fatalf("history did not arrive through Collector: coverage=%d series=%d err=%v", coverageCount(result.Coverage), len(result.Series), err) - } - time.Sleep(250 * time.Millisecond) - } - if err := reader.Close(); err != nil { - t.Fatal(err) - } - assertAcceptanceResult(t, result, firstStartedAt) - - // A fresh Reader proves history is backend-owned, not process memory. - reader = open() - t.Cleanup(func() { _ = reader.Close() }) - reloaded, err := reader.Query(t.Context(), query) - if err != nil { - t.Fatal(err) - } - if coverageCount(reloaded.Coverage) != 5 || len(reloaded.Series) != 1 { - t.Fatalf("history did not survive Reader restart: %+v", reloaded) - } -} - -func acceptanceObservedRecord(tenantID, sessionID, environmentID, allocationID string, startedAt, observedAt time.Time, cpuSeconds float64, memoryBytes uint64) runtimeobs.ExportRecord { - capacity := 2.0 - limit := uint64(2048) - return runtimeobs.ExportRecord{ - TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID, AllocationID: allocationID, - Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusObserved, - CollectionSource: runtimeobs.CollectionSourcePeriodic, ResolvedAt: observedAt.Add(time.Second), - Sample: &runtimeobs.Sample{ - ObservedAt: observedAt, StartedAt: &startedAt, CPUUsageSecondsTotal: &cpuSeconds, CPUCapacityCores: &capacity, - MemoryUsageBytes: &memoryBytes, MemoryLimitBytes: &limit, - }, - } -} - -func coverageCount(values []runtimehistory.CoveragePoint) int { - total := 0 - for _, value := range values { - total += value.ObservationCount - } - return total -} - -func assertAcceptanceResult(t *testing.T, result runtimehistory.Result, firstStartedAt time.Time) { - t.Helper() - if coverageCount(result.Coverage) != 5 { - t.Fatalf("cross-tenant sample leaked into coverage: %+v", result.Coverage) - } - if len(result.Series) != 1 || !result.Series[0].StartedAt.Equal(firstStartedAt) { - t.Fatalf("allocation history was split by compute start estimates: %+v", result.Series) - } - wantRatios := []float64{.5, 1.0 / 6.0} - found := 0 - for _, point := range result.Series[0].Points { - if point.CPUUtilizationRatio != nil { - if found >= len(wantRatios) || math.Abs(*point.CPUUtilizationRatio-wantRatios[found]) > 1e-9 { - t.Fatalf("unexpected CPU ratio after allocation counter reset: %+v", point) - } - found++ - } - } - if found != len(wantRatios) { - t.Fatalf("allocation lost CPU segments across counter reset: %+v", result.Series[0]) - } -} - -func envOr(key, fallback string) string { - if value := os.Getenv(key); value != "" { - return value - } - return fallback -} diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader.go deleted file mode 100644 index eaddb4046..000000000 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader.go +++ /dev/null @@ -1,726 +0,0 @@ -// Package clickhousereader implements the optional Runtime history Reader -// against the operator-owned ClickHouse projection documented with this -// repository. It is a read-only adapter and is never lifecycle authority. -package clickhousereader - -import ( - "context" - "crypto/tls" - "database/sql" - "errors" - "fmt" - "math" - "net" - "sort" - "time" - - "github.com/ClickHouse/clickhouse-go/v2" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" - "github.com/google/uuid" -) - -const historyQuery = ` -SELECT - resolved_at_unix_nano, - observed_at_unix_nano, - allocation_id, - started_at_unix_nano, - provider_type, - status, - metric_name, - min(value) AS minimum_value, - max(value) AS maximum_value -FROM runtime_history_metrics -WHERE tenant_id = {tenant_id:String} - AND session_id = {session_id:String} - AND environment_id = {environment_id:String} - AND collection_source = 'periodic' - AND resolved_at_unix_nano >= {raw_start:Int64} - AND resolved_at_unix_nano < {end:Int64} - AND metric_name IN ( - 'agents.runtime.sample', - 'agents.runtime.cpu.usage', - 'agents.runtime.cpu.capacity', - 'agents.runtime.memory.usage', - 'agents.runtime.memory.limit' - ) -GROUP BY - resolved_at_unix_nano, - observed_at_unix_nano, - allocation_id, - started_at_unix_nano, - provider_type, - status, - metric_name -ORDER BY resolved_at_unix_nano, allocation_id, metric_name -LIMIT {row_limit:UInt64}` - -const tokenUsageQuery = ` -SELECT - resolved_at_unix_nano, - metric_name, - min(value) AS minimum_value, - max(value) AS maximum_value -FROM runtime_history_metrics -WHERE tenant_id = {tenant_id:String} - AND session_id = {session_id:String} - AND environment_id = {environment_id:String} - AND collection_source = 'periodic' - AND resolved_at_unix_nano >= {raw_start:Int64} - AND resolved_at_unix_nano < {end:Int64} - AND metric_name IN ( - 'agents.session.tokens.input', - 'agents.session.tokens.output' - ) -GROUP BY resolved_at_unix_nano, metric_name -ORDER BY resolved_at_unix_nano, metric_name -LIMIT {row_limit:UInt64}` - -type Config struct { - Address string - Database string - Username string - Password string - Secure bool - Insecure bool - DialTimeout time.Duration - QueryTimeout time.Duration - Capabilities runtimehistory.Capabilities -} - -type rows interface { - Next() bool - Scan(...any) error - Err() error - Close() error -} - -type queryClient interface { - Query(context.Context, string, ...any) (rows, error) - Close() error -} - -type sqlClient struct{ db *sql.DB } - -func (c sqlClient) Query(ctx context.Context, statement string, args ...any) (rows, error) { - return c.db.QueryContext(ctx, statement, args...) -} - -func (c sqlClient) Close() error { return c.db.Close() } - -type Reader struct { - client queryClient - capabilities runtimehistory.Capabilities - queryTimeout time.Duration - now func() time.Time -} - -func Open(_ context.Context, config Config) (*Reader, error) { - host, _, err := net.SplitHostPort(config.Address) - if err != nil || host == "" || config.Database == "" || config.Username == "" || config.Secure == config.Insecure || config.DialTimeout <= 0 || config.QueryTimeout <= 0 { - return nil, errors.New("invalid Runtime history ClickHouse configuration") - } - if err := config.Capabilities.Validate(); err != nil { - return nil, errors.New("invalid Runtime history ClickHouse capabilities") - } - options := &clickhouse.Options{ - Addr: []string{config.Address}, - Auth: clickhouse.Auth{Database: config.Database, Username: config.Username, Password: config.Password}, - DialTimeout: config.DialTimeout, - MaxOpenConns: 4, - MaxIdleConns: 2, - ConnMaxLifetime: 10 * time.Minute, - } - if config.Secure { - options.TLS = &tls.Config{MinVersion: tls.VersionTLS12, ServerName: host} - } - db := clickhouse.OpenDB(options) - return newReader(sqlClient{db: db}, config.Capabilities, config.QueryTimeout) -} - -func newReader(client queryClient, capabilities runtimehistory.Capabilities, queryTimeout time.Duration) (*Reader, error) { - if client == nil || queryTimeout <= 0 || capabilities.Validate() != nil { - return nil, errors.New("invalid Runtime history ClickHouse Reader") - } - return &Reader{client: client, capabilities: capabilities, queryTimeout: queryTimeout, now: time.Now}, nil -} - -func (r *Reader) Capabilities() runtimehistory.Capabilities { - value := r.capabilities - value.Metrics = append([]runtimehistory.Metric(nil), value.Metrics...) - return value -} - -func (r *Reader) Close() error { return r.client.Close() } - -func (r *Reader) Query(ctx context.Context, query runtimehistory.Query) (runtimehistory.Result, error) { - if err := r.validateQuery(query); err != nil { - return runtimehistory.Result{}, err - } - generatedAt := r.now().UTC() - lookback := r.capabilities.MinimumStep - if r.capabilities.SampleInterval > 0 { - lookback = 2 * r.capabilities.SampleInterval - } - rawStart := query.Start.Add(-lookback) - retentionStart := generatedAt.Add(-r.capabilities.Retention) - if rawStart.Before(retentionStart) { - rawStart = retentionStart - } - if rawStart.UnixNano() < 0 { - rawStart = time.Unix(0, 0).UTC() - } - rowLimit := query.MaximumTotalPoints*6 + query.MaximumSeries*6 + 1 - queryCtx, cancel := context.WithTimeout(ctx, r.queryTimeout) - defer cancel() - resultRows, err := r.client.Query(queryCtx, historyQuery, - clickhouse.Named("tenant_id", query.TenantID), - clickhouse.Named("session_id", query.SessionID), - clickhouse.Named("environment_id", query.EnvironmentID), - clickhouse.Named("raw_start", rawStart.UnixNano()), - clickhouse.Named("end", query.End.UnixNano()), - clickhouse.Named("row_limit", uint64(rowLimit)), - ) - if err != nil { - return runtimehistory.Result{}, errors.New("query Runtime history") - } - defer resultRows.Close() - raw, count, err := readSamples(resultRows, rawStart, query.End) - if err != nil { - return runtimehistory.Result{}, err - } - if count >= rowLimit { - return runtimehistory.Result{}, errors.New("Runtime history raw result exceeds limit") - } - result, err := aggregate(query, generatedAt, raw) - if err != nil || !hasMetric(r.capabilities.Metrics, runtimehistory.MetricTokens) { - return result, err - } - tokenRowLimit := query.MaximumTotalPoints*2 + 1 - tokenRows, err := r.client.Query(queryCtx, tokenUsageQuery, - clickhouse.Named("tenant_id", query.TenantID), - clickhouse.Named("session_id", query.SessionID), - clickhouse.Named("environment_id", query.EnvironmentID), - clickhouse.Named("raw_start", query.Start.UnixNano()), - clickhouse.Named("end", query.End.UnixNano()), - clickhouse.Named("row_limit", uint64(tokenRowLimit)), - ) - if err != nil { - return runtimehistory.Result{}, errors.New("query Runtime history token usage") - } - defer tokenRows.Close() - rawTokens, tokenCount, err := readTokenUsage(tokenRows, query.Start, query.End) - if err != nil { - return runtimehistory.Result{}, err - } - if tokenCount >= tokenRowLimit { - return runtimehistory.Result{}, errors.New("Runtime history token result exceeds limit") - } - result.TokenUsage, err = aggregateTokenUsage(query, rawTokens) - return result, err -} - -func hasMetric(metrics []runtimehistory.Metric, expected runtimehistory.Metric) bool { - for _, metric := range metrics { - if metric == expected { - return true - } - } - return false -} - -func (r *Reader) validateQuery(query runtimehistory.Query) error { - for _, value := range []string{query.TenantID, query.SessionID, query.EnvironmentID} { - parsed, err := uuid.Parse(value) - if err != nil || parsed == uuid.Nil || parsed.String() != value { - return errors.New("invalid Runtime history ClickHouse scope") - } - } - if query.Start.IsZero() || query.End.IsZero() || !query.End.After(query.Start) || query.Start.Nanosecond() != 0 || query.End.Nanosecond() != 0 || - query.Step < r.capabilities.MinimumStep || query.Step%time.Second != 0 || query.Retention != r.capabilities.Retention || - query.MaxPoints < 2 || query.MaxPoints > r.capabilities.MaximumPoints || query.MaximumSeries != r.capabilities.MaximumSeries || - query.MaximumTotalPoints != r.capabilities.MaximumTotalPoints || query.End.Sub(query.Start) > r.capabilities.MaximumRange { - return errors.New("invalid Runtime history ClickHouse query") - } - buckets := query.End.Sub(query.Start) / query.Step - if query.End.Sub(query.Start)%query.Step != 0 { - buckets++ - } - if buckets > time.Duration(query.MaxPoints) { - return errors.New("Runtime history ClickHouse query exceeds point budget") - } - return nil -} - -type nullableInt64 struct { - value int64 - valid bool -} - -func (n *nullableInt64) Scan(value any) error { - if value == nil { - n.valid = false - return nil - } - switch typed := value.(type) { - case int64: - n.value, n.valid = typed, true - case uint64: - if typed > math.MaxInt64 { - return errors.New("Runtime history timestamp is out of range") - } - n.value, n.valid = int64(typed), true - default: - return fmt.Errorf("invalid Runtime history timestamp type %T", value) - } - return nil -} - -type rawSample struct { - resolvedAt time.Time - observedAt *time.Time - allocation string - startedAt *time.Time - provider string - status runtimeobs.Status - hasSample bool - metrics map[string]float64 -} - -type rawTokenUsage struct { - resolvedAt time.Time - inputTokens, outputTokens *uint64 -} - -func readTokenUsage(resultRows rows, start, end time.Time) ([]rawTokenUsage, int, error) { - byResolved := map[int64]*rawTokenUsage{} - rowCount := 0 - for resultRows.Next() { - rowCount++ - var resolvedNano int64 - var metric string - var minimum, maximum float64 - if err := resultRows.Scan(&resolvedNano, &metric, &minimum, &maximum); err != nil { - return nil, rowCount, errors.New("scan Runtime history token row") - } - resolvedAt := time.Unix(0, resolvedNano).UTC() - value, err := safeUint64(minimum) - if resolvedNano < 0 || resolvedAt.Before(start) || !resolvedAt.Before(end) || minimum != maximum || err != nil { - return nil, rowCount, errors.New("invalid Runtime history token row") - } - usage := byResolved[resolvedNano] - if usage == nil { - usage = &rawTokenUsage{resolvedAt: resolvedAt} - byResolved[resolvedNano] = usage - } - switch metric { - case otlpexporter.TokenInputName: - if usage.inputTokens != nil { - return nil, rowCount, errors.New("duplicate Runtime history input token row") - } - usage.inputTokens = &value - case otlpexporter.TokenOutputName: - if usage.outputTokens != nil { - return nil, rowCount, errors.New("duplicate Runtime history output token row") - } - usage.outputTokens = &value - default: - return nil, rowCount, errors.New("invalid Runtime history token metric") - } - } - if err := resultRows.Err(); err != nil { - return nil, rowCount, errors.New("iterate Runtime history token rows") - } - result := make([]rawTokenUsage, 0, len(byResolved)) - for _, usage := range byResolved { - if usage.inputTokens == nil || usage.outputTokens == nil { - return nil, rowCount, errors.New("incomplete Runtime history token usage") - } - result = append(result, *usage) - } - sort.Slice(result, func(i, j int) bool { return result[i].resolvedAt.Before(result[j].resolvedAt) }) - return result, rowCount, nil -} - -func aggregateTokenUsage(query runtimehistory.Query, raw []rawTokenUsage) ([]runtimehistory.TokenUsagePoint, error) { - latest := map[int]*rawTokenUsage{} - for _, usage := range raw { - index := bucketIndex(query, usage.resolvedAt) - if index < 0 { - continue - } - previous, ok := latest[index] - if !ok || usage.resolvedAt.After(previous.resolvedAt) { - value := usage - latest[index] = &value - } - } - result := make([]runtimehistory.TokenUsagePoint, 0, len(latest)) - for _, index := range sortedIndexes(latest) { - usage := latest[index] - start, end := bucketBounds(query, index) - result = append(result, runtimehistory.TokenUsagePoint{ - Start: start, End: end, SampledAt: usage.resolvedAt, - InputTokens: *usage.inputTokens, OutputTokens: *usage.outputTokens, - }) - } - return result, nil -} - -func readSamples(resultRows rows, rawStart, end time.Time) ([]*rawSample, int, error) { - byKey := map[string]*rawSample{} - rowCount := 0 - for resultRows.Next() { - rowCount++ - var resolvedNano int64 - var observedNano, startedNano nullableInt64 - var allocation, provider, status, metric string - var minimum, maximum float64 - if err := resultRows.Scan(&resolvedNano, &observedNano, &allocation, &startedNano, &provider, &status, &metric, &minimum, &maximum); err != nil { - return nil, rowCount, errors.New("scan Runtime history row") - } - if resolvedNano < 0 || minimum != maximum || math.IsNaN(minimum) || math.IsInf(minimum, 0) || minimum < 0 { - return nil, rowCount, errors.New("invalid Runtime history metric row") - } - resolvedAt := time.Unix(0, resolvedNano).UTC() - if resolvedAt.Before(rawStart) || !resolvedAt.Before(end) { - return nil, rowCount, errors.New("Runtime history row is outside the bounded query") - } - var observedAt *time.Time - if observedNano.valid { - value := time.Unix(0, observedNano.value).UTC() - if observedNano.value < 0 || value.After(resolvedAt) { - return nil, rowCount, errors.New("invalid Runtime history observed timestamp") - } - observedAt = &value - } - var startedAt *time.Time - if startedNano.valid { - value := time.Unix(0, startedNano.value).UTC() - if startedNano.value < 0 || observedAt == nil || value.After(*observedAt) { - return nil, rowCount, errors.New("invalid Runtime history incarnation timestamp") - } - startedAt = &value - } - parsedStatus := runtimeobs.Status(status) - if parsedStatus != runtimeobs.StatusObserved && parsedStatus != runtimeobs.StatusUnavailable && parsedStatus != runtimeobs.StatusUnsupported { - return nil, rowCount, errors.New("invalid Runtime history status") - } - if metric != otlpexporter.SampleName && metric != otlpexporter.CPUUsageName && metric != otlpexporter.CPUCapacityName && metric != otlpexporter.MemoryUsageName && metric != otlpexporter.MemoryLimitName { - return nil, rowCount, errors.New("invalid Runtime history metric") - } - if err := validateProjectedRow(parsedStatus, metric, minimum, observedAt, startedAt); err != nil { - return nil, rowCount, err - } - // Allocation identity is the durable Dashboard series boundary. Provider - // start-time estimates are sample metadata and may jitter between reads. - key := fmt.Sprintf("%d\x00%s", resolvedNano, allocation) - sample := byKey[key] - if sample == nil { - sample = &rawSample{resolvedAt: resolvedAt, observedAt: observedAt, allocation: allocation, startedAt: startedAt, provider: provider, status: parsedStatus, metrics: map[string]float64{}} - byKey[key] = sample - } else if sample.provider != provider || sample.status != parsedStatus || !sameOptionalTime(sample.observedAt, observedAt) { - return nil, rowCount, errors.New("conflicting Runtime history sample identity") - } else if startedAt != nil && (sample.startedAt == nil || startedAt.Before(*sample.startedAt)) { - // Retain the earliest estimate for compatible display metadata without - // allowing it to split one allocation into multiple series. - sample.startedAt = startedAt - } - if previous, ok := sample.metrics[metric]; ok && previous != minimum { - return nil, rowCount, errors.New("conflicting Runtime history metric value") - } - sample.metrics[metric] = minimum - if metric == otlpexporter.SampleName { - sample.hasSample = true - } - } - if err := resultRows.Err(); err != nil { - return nil, rowCount, errors.New("iterate Runtime history rows") - } - result := make([]*rawSample, 0, len(byKey)) - for _, sample := range byKey { - if !sample.hasSample { - return nil, rowCount, errors.New("Runtime history resource row has no sample marker") - } - result = append(result, sample) - } - sort.Slice(result, func(i, j int) bool { - if result[i].resolvedAt.Equal(result[j].resolvedAt) { - return result[i].allocation < result[j].allocation - } - return result[i].resolvedAt.Before(result[j].resolvedAt) - }) - return result, rowCount, nil -} - -func validateProjectedRow(status runtimeobs.Status, metric string, value float64, observedAt, startedAt *time.Time) error { - if metric == otlpexporter.SampleName && value != 1 { - return errors.New("invalid Runtime history sample marker") - } - if status == runtimeobs.StatusObserved { - if observedAt == nil { - return errors.New("observed Runtime history row has no observation timestamp") - } - if metric != otlpexporter.SampleName && startedAt == nil { - return errors.New("Runtime history resource metric has no incarnation fence") - } - return nil - } - if observedAt != nil || startedAt != nil || metric != otlpexporter.SampleName { - return errors.New("unobserved Runtime history row carries sample data") - } - return nil -} - -func sameOptionalTime(left, right *time.Time) bool { - return left == nil && right == nil || left != nil && right != nil && left.Equal(*right) -} - -type coverageAggregate struct { - first, last time.Time - observations int - observed, unavailable int -} - -type pointAggregate struct { - coverageAggregate - cpuContributors, memoryContributors int - cpuUsageDelta, cpuCapacitySeconds float64 - cpuCapacity *float64 - memoryUsage, memoryLimit *uint64 -} - -type seriesAggregate struct { - allocation string - startedAt time.Time - provider string - samples []*rawSample -} - -func aggregate(query runtimehistory.Query, generatedAt time.Time, raw []*rawSample) (runtimehistory.Result, error) { - coverage := map[int]*coverageAggregate{} - series := map[string]*seriesAggregate{} - for _, sample := range raw { - if !sample.hasSample { - continue - } - if !sample.resolvedAt.Before(query.Start) { - index := bucketIndex(query, sample.resolvedAt) - if index >= 0 { - value := coverage[index] - if value == nil { - value = &coverageAggregate{} - coverage[index] = value - } - addCoverage(value, sample) - } - } - if sample.allocation == "" || sample.startedAt == nil || sample.provider == "" { - continue - } - parsedAllocation, allocationErr := uuid.Parse(sample.allocation) - if allocationErr != nil || parsedAllocation == uuid.Nil || parsedAllocation.String() != sample.allocation { - return runtimehistory.Result{}, errors.New("invalid Runtime history allocation") - } - key := sample.allocation - value := series[key] - if value == nil { - value = &seriesAggregate{allocation: sample.allocation, startedAt: *sample.startedAt, provider: sample.provider} - series[key] = value - } else if value.provider != sample.provider { - return runtimehistory.Result{}, errors.New("conflicting Runtime history provider") - } else if sample.startedAt.Before(value.startedAt) { - value.startedAt = *sample.startedAt - } - value.samples = append(value.samples, sample) - } - result := runtimehistory.Result{GeneratedAt: generatedAt, Coverage: make([]runtimehistory.CoveragePoint, 0, len(coverage)), Series: make([]runtimehistory.Series, 0, len(series))} - for _, index := range sortedIndexes(coverage) { - value := coverage[index] - start, end := bucketBounds(query, index) - result.Coverage = append(result.Coverage, runtimehistory.CoveragePoint{ - Start: start, End: end, FirstObservedAt: value.first, LastObservedAt: value.last, - ObservationCount: value.observations, ObservedCount: value.observed, UnavailableCount: value.unavailable, - }) - } - keys := make([]string, 0, len(series)) - for key := range series { - keys = append(keys, key) - } - sort.Strings(keys) - totalPoints := len(result.Coverage) - for _, key := range keys { - value := series[key] - points, err := aggregateSeries(query, value.samples) - if err != nil { - return runtimehistory.Result{}, err - } - if len(points) == 0 { - continue - } - if len(result.Series) >= query.MaximumSeries { - return runtimehistory.Result{}, errors.New("Runtime history series limit exceeded") - } - totalPoints += len(points) - if len(points) > query.MaxPoints || totalPoints > query.MaximumTotalPoints { - return runtimehistory.Result{}, errors.New("Runtime history point limit exceeded") - } - result.Series = append(result.Series, runtimehistory.Series{ - Scope: query.Scope, AllocationID: value.allocation, StartedAt: value.startedAt, ProviderType: value.provider, Points: points, - }) - } - return result, nil -} - -func aggregateSeries(query runtimehistory.Query, samples []*rawSample) ([]runtimehistory.Point, error) { - samples = append([]*rawSample(nil), samples...) - sort.SliceStable(samples, func(i, j int) bool { - if samples[i].observedAt == nil { - return samples[j].observedAt != nil - } - if samples[j].observedAt == nil { - return false - } - if samples[i].observedAt.Equal(*samples[j].observedAt) { - return samples[i].resolvedAt.Before(samples[j].resolvedAt) - } - return samples[i].observedAt.Before(*samples[j].observedAt) - }) - points := map[int]*pointAggregate{} - var previousUsage, previousCapacity *float64 - var previousAt time.Time - for _, sample := range samples { - usage, hasUsage := sample.metrics[otlpexporter.CPUUsageName] - capacity, hasCapacity := sample.metrics[otlpexporter.CPUCapacityName] - if hasCapacity && capacity <= 0 { - return nil, errors.New("invalid Runtime history CPU capacity") - } - if sample.hasSample && !sample.resolvedAt.Before(query.Start) { - index := bucketIndex(query, sample.resolvedAt) - if index >= 0 { - point := points[index] - if point == nil { - point = &pointAggregate{} - points[index] = point - } - addCoverage(&point.coverageAggregate, sample) - cpuContributed := false - if hasCapacity { - value := capacity - point.cpuCapacity = &value - cpuContributed = true - } - if hasUsage && hasCapacity && sample.observedAt != nil && previousUsage != nil && previousCapacity != nil && !previousAt.IsZero() && sample.observedAt.After(previousAt) && usage >= *previousUsage { - point.cpuUsageDelta += usage - *previousUsage - point.cpuCapacitySeconds += sample.observedAt.Sub(previousAt).Seconds() * capacity - cpuContributed = true - } - if cpuContributed { - point.cpuContributors++ - } - memoryContributed := false - if rawValue, ok := sample.metrics[otlpexporter.MemoryUsageName]; ok { - value, err := safeUint64(rawValue) - if err != nil { - return nil, err - } - point.memoryUsage = &value - memoryContributed = true - } - if rawValue, ok := sample.metrics[otlpexporter.MemoryLimitName]; ok { - value, err := safeUint64(rawValue) - if err != nil || value == 0 { - return nil, errors.New("invalid Runtime history memory limit") - } - point.memoryLimit = &value - memoryContributed = true - } - if memoryContributed { - point.memoryContributors++ - } - } - } - if hasUsage { - value := usage - previousUsage = &value - if sample.observedAt != nil { - previousAt = *sample.observedAt - } else { - previousAt = time.Time{} - } - if hasCapacity { - value := capacity - previousCapacity = &value - } else { - previousCapacity = nil - } - } - } - result := make([]runtimehistory.Point, 0, len(points)) - for _, index := range sortedIndexes(points) { - value := points[index] - start, end := bucketBounds(query, index) - point := runtimehistory.Point{ - Start: start, End: end, FirstObservedAt: value.first, LastObservedAt: value.last, - ObservationCount: value.observations, ObservedCount: value.observed, UnavailableCount: value.unavailable, - CPUContributorCount: value.cpuContributors, MemoryContributorCount: value.memoryContributors, - CPUCapacityCores: value.cpuCapacity, MemoryUsageBytes: value.memoryUsage, MemoryLimitBytes: value.memoryLimit, - } - if value.cpuCapacitySeconds > 0 { - ratio := value.cpuUsageDelta / value.cpuCapacitySeconds - point.CPUUtilizationRatio = &ratio - } - result = append(result, point) - } - return result, nil -} - -func safeUint64(value float64) (uint64, error) { - const maxSafeInteger = uint64(1<<53 - 1) - if value < 0 || value > float64(maxSafeInteger) || math.Trunc(value) != value { - return 0, errors.New("invalid Runtime history byte value") - } - return uint64(value), nil -} - -func addCoverage(value *coverageAggregate, sample *rawSample) { - value.observations++ - if sample.status == runtimeobs.StatusObserved { - value.observed++ - } else { - value.unavailable++ - } - if value.first.IsZero() || sample.resolvedAt.Before(value.first) { - value.first = sample.resolvedAt - } - if value.last.IsZero() || sample.resolvedAt.After(value.last) { - value.last = sample.resolvedAt - } -} - -func bucketIndex(query runtimehistory.Query, value time.Time) int { - if value.Before(query.Start) || !value.Before(query.End) { - return -1 - } - return int(value.Sub(query.Start) / query.Step) -} - -func bucketBounds(query runtimehistory.Query, index int) (time.Time, time.Time) { - start := query.Start.Add(time.Duration(index) * query.Step) - end := start.Add(query.Step) - if end.After(query.End) { - end = query.End - } - return start, end -} - -func sortedIndexes[T any](values map[int]*T) []int { - result := make([]int, 0, len(values)) - for index := range values { - result = append(result, index) - } - sort.Ints(result) - return result -} diff --git a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go b/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go deleted file mode 100644 index e63add572..000000000 --- a/services/agents-api/internal/runtimehistory/clickhousereader/reader_test.go +++ /dev/null @@ -1,440 +0,0 @@ -package clickhousereader - -import ( - "context" - "database/sql" - "errors" - "strings" - "testing" - "time" - - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" - "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs/otlpexporter" - "github.com/google/uuid" -) - -const ( - testTenant = "11111111-1111-4111-8111-111111111111" - testSession = "22222222-2222-4222-8222-222222222222" - testEnvironment = "33333333-3333-4333-8333-333333333333" - testAllocation = "44444444-4444-4444-8444-444444444444" -) - -type fakeClient struct { - rows rows - rowsQueue []rows - err error - statement string - statements []string - args []any - closed bool -} - -func (c *fakeClient) Query(_ context.Context, statement string, args ...any) (rows, error) { - c.statement, c.args = statement, append([]any(nil), args...) - c.statements = append(c.statements, statement) - if len(c.rowsQueue) > 0 { - value := c.rowsQueue[0] - c.rowsQueue = c.rowsQueue[1:] - return value, c.err - } - return c.rows, c.err -} - -func (c *fakeClient) Close() error { - c.closed = true - return nil -} - -type fakeRows struct { - values [][]any - index int - err error -} - -func (r *fakeRows) Next() bool { return r.index < len(r.values) } - -func (r *fakeRows) Scan(destinations ...any) error { - if r.index >= len(r.values) || len(destinations) != len(r.values[r.index]) { - return errors.New("invalid fake scan") - } - for index, value := range r.values[r.index] { - switch destination := destinations[index].(type) { - case *int64: - *destination = value.(int64) - case *float64: - *destination = value.(float64) - case *string: - *destination = value.(string) - case sql.Scanner: - if err := destination.Scan(value); err != nil { - return err - } - default: - return errors.New("unsupported fake scan destination") - } - } - r.index++ - return nil -} - -func (r *fakeRows) Err() error { return r.err } -func (r *fakeRows) Close() error { return nil } - -func testCapabilities() runtimehistory.Capabilities { - return runtimehistory.Capabilities{ - CollectionMode: runtimehistory.CollectionPeriodic, SampleInterval: 30 * time.Second, - Retention: 7 * 24 * time.Hour, MinimumStep: 30 * time.Second, MaximumRange: 24 * time.Hour, - MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, - Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory}, - } -} - -func testQuery(start time.Time) runtimehistory.Query { - return runtimehistory.Query{ - Scope: runtimehistory.Scope{TenantID: testTenant, SessionID: testSession, EnvironmentID: testEnvironment}, - Start: start, End: start.Add(time.Minute), Step: 30 * time.Second, Retention: 7 * 24 * time.Hour, - MaxPoints: 2, MaximumSeries: 64, MaximumTotalPoints: 10_000, - } -} - -func metricRow(resolved, observed int64, allocation string, started *int64, provider, status, metric string, value float64) []any { - var observedValue any = observed - if observed < 0 { - observedValue = nil - } - var startedValue any - if started != nil { - startedValue = *started - } - return []any{resolved, observedValue, allocation, startedValue, provider, status, metric, value, value} -} - -func TestReaderQueriesMandatoryScopeAndAggregatesFencedHistory(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - started := start.Add(-time.Minute).UnixNano() - baseline := start.Add(-30 * time.Second).UnixNano() - observed := start.Add(10 * time.Second).UnixNano() - unavailable := start.Add(40 * time.Second).UnixNano() - client := &fakeClient{rows: &fakeRows{values: [][]any{ - metricRow(baseline, baseline, testAllocation, &started, "docker", "observed", otlpexporter.SampleName, 1), - metricRow(baseline, baseline, testAllocation, &started, "docker", "observed", otlpexporter.CPUUsageName, 1), - metricRow(baseline, baseline, testAllocation, &started, "docker", "observed", otlpexporter.CPUCapacityName, 2), - metricRow(observed, observed, testAllocation, &started, "docker", "observed", otlpexporter.SampleName, 1), - metricRow(observed, observed, testAllocation, &started, "docker", "observed", otlpexporter.CPUUsageName, 3), - metricRow(observed, observed, testAllocation, &started, "docker", "observed", otlpexporter.CPUCapacityName, 2), - metricRow(observed, observed, testAllocation, &started, "docker", "observed", otlpexporter.MemoryUsageName, 1024), - metricRow(observed, observed, testAllocation, &started, "docker", "observed", otlpexporter.MemoryLimitName, 2048), - metricRow(unavailable, -1, "", nil, "docker", "unavailable", otlpexporter.SampleName, 1), - }}} - reader, err := newReader(client, testCapabilities(), 2*time.Second) - if err != nil { - t.Fatal(err) - } - reader.now = func() time.Time { return start.Add(time.Minute) } - result, err := reader.Query(t.Context(), testQuery(start)) - if err != nil { - t.Fatal(err) - } - for _, required := range []string{ - "tenant_id = {tenant_id:String}", "session_id = {session_id:String}", "environment_id = {environment_id:String}", - "collection_source = 'periodic'", "resolved_at_unix_nano >= {raw_start:Int64}", "resolved_at_unix_nano < {end:Int64}", - } { - if !strings.Contains(client.statement, required) { - t.Fatalf("history query lacks mandatory predicate %q", required) - } - } - if len(client.args) != 6 { - t.Fatalf("unexpected bound argument count: %d", len(client.args)) - } - if len(result.Coverage) != 2 || result.Coverage[0].ObservedCount != 1 || result.Coverage[1].UnavailableCount != 1 { - t.Fatalf("unexpected coverage: %+v", result.Coverage) - } - if len(result.Series) != 1 || result.Series[0].AllocationID != testAllocation || !result.Series[0].StartedAt.Equal(time.Unix(0, started)) || len(result.Series[0].Points) != 1 { - t.Fatalf("unexpected fenced series: %+v", result.Series) - } - point := result.Series[0].Points[0] - if point.CPUUtilizationRatio == nil || *point.CPUUtilizationRatio != .025 || point.CPUCapacityCores == nil || *point.CPUCapacityCores != 2 || - point.MemoryUsageBytes == nil || *point.MemoryUsageBytes != 1024 || point.MemoryLimitBytes == nil || *point.MemoryLimitBytes != 2048 { - t.Fatalf("unexpected metric aggregation: %+v", point) - } - if err := runtimehistoryResponseValidation(testQuery(start), result); err != nil { - t.Fatal(err) - } -} - -func TestReaderReturnsSessionTokenUsageIndependentOfRuntimeIncarnation(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - first := start.Add(10 * time.Second).UnixNano() - second := start.Add(40 * time.Second).UnixNano() - client := &fakeClient{rowsQueue: []rows{ - &fakeRows{}, - &fakeRows{values: [][]any{ - {first, otlpexporter.TokenInputName, float64(100), float64(100)}, - {first, otlpexporter.TokenOutputName, float64(20), float64(20)}, - {second, otlpexporter.TokenInputName, float64(160), float64(160)}, - {second, otlpexporter.TokenOutputName, float64(50), float64(50)}, - }}, - }} - capabilities := testCapabilities() - capabilities.Metrics = append(capabilities.Metrics, runtimehistory.MetricTokens) - reader, err := newReader(client, capabilities, 2*time.Second) - if err != nil { - t.Fatal(err) - } - reader.now = func() time.Time { return start.Add(time.Minute) } - result, err := reader.Query(t.Context(), testQuery(start)) - if err != nil { - t.Fatal(err) - } - if len(client.statements) != 2 || !strings.Contains(client.statements[1], "agents.session.tokens.input") { - t.Fatalf("token query was not issued separately: %#v", client.statements) - } - if len(result.TokenUsage) != 2 || result.TokenUsage[0].InputTokens != 100 || result.TokenUsage[1].OutputTokens != 50 { - t.Fatalf("unexpected token usage history: %+v", result.TokenUsage) - } -} - -func TestReaderRejectsIncompleteSessionTokenUsage(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - rows := &fakeRows{values: [][]any{{ - start.Add(10 * time.Second).UnixNano(), otlpexporter.TokenInputName, float64(100), float64(100), - }}} - if _, _, err := readTokenUsage(rows, start, start.Add(time.Minute)); err == nil { - t.Fatal("incomplete token usage was accepted") - } -} - -func TestReaderKeepsStartedAtJitterInOneAllocationSeries(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - firstStarted := start.Add(-time.Minute).UnixNano() - secondStarted := start.Add(-time.Minute + 1500*time.Millisecond).UnixNano() - baseline := start.Add(-10 * time.Second).UnixNano() - observed := start.Add(10 * time.Second).UnixNano() - client := &fakeClient{rows: &fakeRows{values: [][]any{ - metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.SampleName, 1), - metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.CPUUsageName, 1), - metricRow(baseline, baseline, testAllocation, &firstStarted, "microsandbox", "observed", otlpexporter.CPUCapacityName, 2), - metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.SampleName, 1), - metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.CPUUsageName, 3), - metricRow(observed, observed, testAllocation, &secondStarted, "microsandbox", "observed", otlpexporter.CPUCapacityName, 2), - }}} - reader, err := newReader(client, testCapabilities(), time.Second) - if err != nil { - t.Fatal(err) - } - reader.now = func() time.Time { return start.Add(time.Minute) } - result, err := reader.Query(t.Context(), testQuery(start)) - if err != nil { - t.Fatal(err) - } - if len(result.Series) != 1 || !result.Series[0].StartedAt.Equal(time.Unix(0, firstStarted)) { - t.Fatalf("start-time jitter split one allocation: %+v", result.Series) - } - if len(result.Series[0].Points) != 1 || result.Series[0].Points[0].CPUUtilizationRatio == nil || *result.Series[0].Points[0].CPUUtilizationRatio != .05 { - t.Fatalf("allocation CPU interval was not retained: %+v", result.Series[0].Points) - } -} - -func runtimehistoryResponseValidation(query runtimehistory.Query, result runtimehistory.Result) error { - response := runtimehistory.Response{ - Capabilities: testCapabilities(), Scope: query.Scope, - Requested: runtimehistory.Range{Start: query.Start, End: query.End, MaxPoints: query.MaxPoints}, Resolution: query.Step, - GeneratedAt: result.GeneratedAt, RetainedFrom: result.RetainedFrom, Coverage: result.Coverage, Series: result.Series, - } - return response.Validate(result.GeneratedAt) -} - -func TestReaderFailsClosedOnBackendAndConflictingRows(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - client := &fakeClient{err: errors.New("secret backend detail")} - reader, err := newReader(client, testCapabilities(), time.Second) - if err != nil { - t.Fatal(err) - } - reader.now = func() time.Time { return start.Add(time.Minute) } - if _, err := reader.Query(t.Context(), testQuery(start)); err == nil || strings.Contains(err.Error(), "secret") { - t.Fatalf("backend failure was not sanitized: %v", err) - } - - resolved := start.Add(10 * time.Second).UnixNano() - started := start.Add(-time.Minute).UnixNano() - rows := metricRow(resolved, resolved, testAllocation, &started, "docker", "observed", otlpexporter.SampleName, 1) - rows[len(rows)-1] = float64(2) - client = &fakeClient{rows: &fakeRows{values: [][]any{rows}}} - reader, err = newReader(client, testCapabilities(), time.Second) - if err != nil { - t.Fatal(err) - } - reader.now = func() time.Time { return start.Add(time.Minute) } - if _, err := reader.Query(t.Context(), testQuery(start)); err == nil { - t.Fatal("conflicting duplicate metric values were accepted") - } -} - -func TestReadSamplesRejectsMalformedProjectedRecords(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - resolved := start.Add(10 * time.Second).UnixNano() - started := start.Add(-time.Minute).UnixNano() - tests := []struct { - name string - rows [][]any - }{ - { - name: "observed marker without observed timestamp", - rows: [][]any{metricRow(resolved, -1, testAllocation, nil, "docker", "observed", otlpexporter.SampleName, 1)}, - }, - { - name: "unavailable marker with observed timestamp", - rows: [][]any{metricRow(resolved, resolved, testAllocation, nil, "docker", "unavailable", otlpexporter.SampleName, 1)}, - }, - { - name: "unsupported marker with incarnation timestamp", - rows: [][]any{metricRow(resolved, -1, testAllocation, &started, "docker", "unsupported", otlpexporter.SampleName, 1)}, - }, - { - name: "unavailable resource metric", - rows: [][]any{metricRow(resolved, -1, testAllocation, nil, "docker", "unavailable", otlpexporter.MemoryUsageName, 1)}, - }, - { - name: "observed resource metric without incarnation fence", - rows: [][]any{metricRow(resolved, resolved, testAllocation, nil, "docker", "observed", otlpexporter.MemoryUsageName, 1)}, - }, - { - name: "invalid sample marker value", - rows: [][]any{metricRow(resolved, resolved, testAllocation, &started, "docker", "observed", otlpexporter.SampleName, 2)}, - }, - { - name: "resource metrics without sample marker", - rows: [][]any{metricRow(resolved, resolved, testAllocation, &started, "docker", "observed", otlpexporter.MemoryUsageName, 1)}, - }, - } - for _, test := range tests { - t.Run(test.name, func(t *testing.T) { - if _, _, err := readSamples(&fakeRows{values: test.rows}, start, start.Add(time.Minute)); err == nil { - t.Fatal("malformed projection rows were accepted") - } - }) - } -} - -func TestReaderRejectsUnscopedQueriesAndClosesClient(t *testing.T) { - client := &fakeClient{rows: &fakeRows{}} - reader, err := newReader(client, testCapabilities(), time.Second) - if err != nil { - t.Fatal(err) - } - query := testQuery(time.Now().UTC().Truncate(time.Second).Add(-time.Minute)) - query.TenantID = "" - if _, err := reader.Query(t.Context(), query); err == nil || client.statement != "" { - t.Fatal("unscoped query reached ClickHouse") - } - if err := reader.Close(); err != nil || !client.closed { - t.Fatal("Reader did not close ClickHouse client") - } -} - -func TestOpenRequiresExplicitClickHouseTransportSecurity(t *testing.T) { - config := Config{ - Address: "clickhouse.example.test:9000", Database: "runtime_history", Username: "reader", Password: "must-not-leak", - DialTimeout: time.Second, QueryTimeout: time.Second, Capabilities: testCapabilities(), - } - if _, err := Open(t.Context(), config); err == nil || strings.Contains(err.Error(), "must-not-leak") { - t.Fatalf("implicit plaintext ClickHouse configuration was accepted or leaked: %v", err) - } - config.Insecure = true - reader, err := Open(t.Context(), config) - if err != nil { - t.Fatal(err) - } - if err := reader.Close(); err != nil { - t.Fatal(err) - } -} - -func TestAggregateSeriesUsesProviderObservationOrder(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - query := testQuery(start) - startedAt := start.Add(-time.Minute) - firstObserved, secondObserved := start.Add(5*time.Second), start.Add(15*time.Second) - firstResolved, secondResolved := start.Add(20*time.Second), start.Add(10*time.Second) - samples := []*rawSample{ - { - resolvedAt: secondResolved, observedAt: &secondObserved, allocation: testAllocation, startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, - metrics: map[string]float64{otlpexporter.SampleName: 1, otlpexporter.CPUUsageName: 3, otlpexporter.CPUCapacityName: 2, otlpexporter.MemoryUsageName: 200}, - }, - { - resolvedAt: firstResolved, observedAt: &firstObserved, allocation: testAllocation, startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, - metrics: map[string]float64{otlpexporter.SampleName: 1, otlpexporter.CPUUsageName: 1, otlpexporter.CPUCapacityName: 2, otlpexporter.MemoryUsageName: 100}, - }, - } - points, err := aggregateSeries(query, samples) - if err != nil { - t.Fatal(err) - } - if len(points) != 1 || points[0].CPUUtilizationRatio == nil || *points[0].CPUUtilizationRatio != .1 || points[0].MemoryUsageBytes == nil || *points[0].MemoryUsageBytes != 200 { - t.Fatalf("provider observation order was not preserved: %+v", points) - } -} - -func TestAggregateSeriesLeavesCPUUsageGapWhenEitherEndpointLacksCapacity(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - query := testQuery(start) - startedAt := start.Add(-time.Minute) - firstObserved, secondObserved, thirdObserved := start.Add(-5*time.Second), start.Add(5*time.Second), start.Add(15*time.Second) - samples := []*rawSample{ - { - resolvedAt: firstObserved, observedAt: &firstObserved, allocation: testAllocation, startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, - metrics: map[string]float64{otlpexporter.SampleName: 1, otlpexporter.CPUUsageName: 1, otlpexporter.CPUCapacityName: 2}, - }, - { - resolvedAt: secondObserved, observedAt: &secondObserved, allocation: testAllocation, startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, - metrics: map[string]float64{otlpexporter.SampleName: 1, otlpexporter.CPUUsageName: 3}, - }, - { - resolvedAt: thirdObserved, observedAt: &thirdObserved, allocation: testAllocation, startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, - metrics: map[string]float64{otlpexporter.SampleName: 1, otlpexporter.CPUUsageName: 5, otlpexporter.CPUCapacityName: 2}, - }, - } - points, err := aggregateSeries(query, samples) - if err != nil { - t.Fatal(err) - } - if len(points) != 1 || points[0].CPUUtilizationRatio != nil { - t.Fatalf("missing CPU capacity was synthesized into utilization: %+v", points) - } -} - -func TestAggregateIgnoresLookbackOnlySeriesForSeriesLimit(t *testing.T) { - start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) - query := testQuery(start) - query.MaximumSeries = 64 - raw := make([]*rawSample, 0, 66) - for range 65 { - startedAt := start.Add(-time.Minute) - observedAt := start.Add(-2 * time.Second) - raw = append(raw, &rawSample{ - resolvedAt: start.Add(-time.Second), observedAt: &observedAt, allocation: uuid.NewString(), startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, metrics: map[string]float64{otlpexporter.SampleName: 1}, - }) - } - startedAt := start.Add(-time.Minute) - observedAt := start.Add(5 * time.Second) - raw = append(raw, &rawSample{ - resolvedAt: start.Add(6 * time.Second), observedAt: &observedAt, allocation: uuid.NewString(), startedAt: &startedAt, - provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, metrics: map[string]float64{otlpexporter.SampleName: 1}, - }) - result, err := aggregate(query, start.Add(time.Minute), raw) - if err != nil { - t.Fatal(err) - } - if len(result.Series) != 1 || len(result.Series[0].Points) != 1 { - t.Fatalf("lookback-only incarnations consumed returned-series budget: %+v", result.Series) - } -} diff --git a/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go b/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go new file mode 100644 index 000000000..c34d2841c --- /dev/null +++ b/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go @@ -0,0 +1,318 @@ +// Package postgresreader stores and reads bounded periodic Runtime telemetry. +package postgresreader + +import ( + "errors" + "math" + "sort" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" +) + +const ( + CPUUsageName = "cpu_usage_seconds" + CPUCapacityName = "cpu_capacity_cores" + MemoryUsageName = "memory_usage_bytes" + MemoryLimitName = "memory_limit_bytes" +) + +type rawSample struct { + resolvedAt time.Time + observedAt *time.Time + allocation string + startedAt *time.Time + provider string + status runtimeobs.Status + hasSample bool + metrics map[string]float64 +} + +type rawTokenUsage struct { + resolvedAt time.Time + inputTokens, outputTokens *uint64 +} + +func aggregateTokenUsage(query runtimehistory.Query, raw []rawTokenUsage) ([]runtimehistory.TokenUsagePoint, error) { + latest := map[int]*rawTokenUsage{} + for _, usage := range raw { + index := bucketIndex(query, usage.resolvedAt) + if index < 0 { + continue + } + previous, ok := latest[index] + if !ok || usage.resolvedAt.After(previous.resolvedAt) { + value := usage + latest[index] = &value + } + } + result := make([]runtimehistory.TokenUsagePoint, 0, len(latest)) + for _, index := range sortedIndexes(latest) { + usage := latest[index] + start, end := bucketBounds(query, index) + result = append(result, runtimehistory.TokenUsagePoint{ + Start: start, End: end, SampledAt: usage.resolvedAt, + InputTokens: *usage.inputTokens, OutputTokens: *usage.outputTokens, + }) + } + return result, nil +} + +type coverageAggregate struct { + first, last time.Time + observations int + observed, unavailable int +} + +type pointAggregate struct { + coverageAggregate + cpuContributors, memoryContributors int + cpuUsageDelta, cpuCapacitySeconds float64 + cpuCapacity *float64 + memoryUsage, memoryLimit *uint64 +} + +type seriesAggregate struct { + allocation string + startedAt time.Time + provider string + samples []*rawSample +} + +func aggregate(query runtimehistory.Query, generatedAt time.Time, raw []*rawSample) (runtimehistory.Result, error) { + coverage := map[int]*coverageAggregate{} + series := map[string]*seriesAggregate{} + for _, sample := range raw { + if !sample.hasSample { + continue + } + if !sample.resolvedAt.Before(query.Start) { + index := bucketIndex(query, sample.resolvedAt) + if index >= 0 { + value := coverage[index] + if value == nil { + value = &coverageAggregate{} + coverage[index] = value + } + addCoverage(value, sample) + } + } + if sample.allocation == "" || sample.startedAt == nil || sample.provider == "" { + continue + } + parsedAllocation, allocationErr := uuid.Parse(sample.allocation) + if allocationErr != nil || parsedAllocation == uuid.Nil || parsedAllocation.String() != sample.allocation { + return runtimehistory.Result{}, errors.New("invalid Runtime history allocation") + } + key := sample.allocation + value := series[key] + if value == nil { + value = &seriesAggregate{allocation: sample.allocation, startedAt: *sample.startedAt, provider: sample.provider} + series[key] = value + } else if value.provider != sample.provider { + return runtimehistory.Result{}, errors.New("conflicting Runtime history provider") + } else if sample.startedAt.Before(value.startedAt) { + value.startedAt = *sample.startedAt + } + value.samples = append(value.samples, sample) + } + result := runtimehistory.Result{GeneratedAt: generatedAt, Coverage: make([]runtimehistory.CoveragePoint, 0, len(coverage)), Series: make([]runtimehistory.Series, 0, len(series))} + for _, index := range sortedIndexes(coverage) { + value := coverage[index] + start, end := bucketBounds(query, index) + result.Coverage = append(result.Coverage, runtimehistory.CoveragePoint{ + Start: start, End: end, FirstObservedAt: value.first, LastObservedAt: value.last, + ObservationCount: value.observations, ObservedCount: value.observed, UnavailableCount: value.unavailable, + }) + } + keys := make([]string, 0, len(series)) + for key := range series { + keys = append(keys, key) + } + sort.Strings(keys) + totalPoints := len(result.Coverage) + for _, key := range keys { + value := series[key] + points, err := aggregateSeries(query, value.samples) + if err != nil { + return runtimehistory.Result{}, err + } + if len(points) == 0 { + continue + } + if len(result.Series) >= query.MaximumSeries { + return runtimehistory.Result{}, errors.New("Runtime history series limit exceeded") + } + totalPoints += len(points) + if len(points) > query.MaxPoints || totalPoints > query.MaximumTotalPoints { + return runtimehistory.Result{}, errors.New("Runtime history point limit exceeded") + } + result.Series = append(result.Series, runtimehistory.Series{ + Scope: query.Scope, AllocationID: value.allocation, StartedAt: value.startedAt, ProviderType: value.provider, Points: points, + }) + } + return result, nil +} + +func aggregateSeries(query runtimehistory.Query, samples []*rawSample) ([]runtimehistory.Point, error) { + samples = append([]*rawSample(nil), samples...) + sort.SliceStable(samples, func(i, j int) bool { + if samples[i].observedAt == nil { + return samples[j].observedAt != nil + } + if samples[j].observedAt == nil { + return false + } + if samples[i].observedAt.Equal(*samples[j].observedAt) { + return samples[i].resolvedAt.Before(samples[j].resolvedAt) + } + return samples[i].observedAt.Before(*samples[j].observedAt) + }) + points := map[int]*pointAggregate{} + var previousUsage, previousCapacity *float64 + var previousAt time.Time + var previousStartedAt *time.Time + for _, sample := range samples { + if !sameOptionalTime(previousStartedAt, sample.startedAt) { + previousUsage, previousCapacity = nil, nil + previousAt = time.Time{} + } + previousStartedAt = sample.startedAt + usage, hasUsage := sample.metrics[CPUUsageName] + capacity, hasCapacity := sample.metrics[CPUCapacityName] + if hasCapacity && capacity <= 0 { + return nil, errors.New("invalid Runtime history CPU capacity") + } + if sample.hasSample && !sample.resolvedAt.Before(query.Start) { + index := bucketIndex(query, sample.resolvedAt) + if index >= 0 { + point := points[index] + if point == nil { + point = &pointAggregate{} + points[index] = point + } + addCoverage(&point.coverageAggregate, sample) + cpuContributed := false + if hasCapacity { + value := capacity + point.cpuCapacity = &value + cpuContributed = true + } + if hasUsage && hasCapacity && sample.observedAt != nil && previousUsage != nil && previousCapacity != nil && !previousAt.IsZero() && sample.observedAt.After(previousAt) && usage >= *previousUsage { + point.cpuUsageDelta += usage - *previousUsage + point.cpuCapacitySeconds += sample.observedAt.Sub(previousAt).Seconds() * capacity + cpuContributed = true + } + if cpuContributed { + point.cpuContributors++ + } + memoryContributed := false + if rawValue, ok := sample.metrics[MemoryUsageName]; ok { + value, err := safeUint64(rawValue) + if err != nil { + return nil, err + } + point.memoryUsage = &value + memoryContributed = true + } + if rawValue, ok := sample.metrics[MemoryLimitName]; ok { + value, err := safeUint64(rawValue) + if err != nil || value == 0 { + return nil, errors.New("invalid Runtime history memory limit") + } + point.memoryLimit = &value + memoryContributed = true + } + if memoryContributed { + point.memoryContributors++ + } + } + } + if hasUsage { + value := usage + previousUsage = &value + if sample.observedAt != nil { + previousAt = *sample.observedAt + } else { + previousAt = time.Time{} + } + if hasCapacity { + value := capacity + previousCapacity = &value + } else { + previousCapacity = nil + } + } + } + result := make([]runtimehistory.Point, 0, len(points)) + for _, index := range sortedIndexes(points) { + value := points[index] + start, end := bucketBounds(query, index) + point := runtimehistory.Point{ + Start: start, End: end, FirstObservedAt: value.first, LastObservedAt: value.last, + ObservationCount: value.observations, ObservedCount: value.observed, UnavailableCount: value.unavailable, + CPUContributorCount: value.cpuContributors, MemoryContributorCount: value.memoryContributors, + CPUCapacityCores: value.cpuCapacity, MemoryUsageBytes: value.memoryUsage, MemoryLimitBytes: value.memoryLimit, + } + if value.cpuCapacitySeconds > 0 { + ratio := value.cpuUsageDelta / value.cpuCapacitySeconds + point.CPUUtilizationRatio = &ratio + } + result = append(result, point) + } + return result, nil +} + +func safeUint64(value float64) (uint64, error) { + const maxSafeInteger = uint64(1<<53 - 1) + if value < 0 || value > float64(maxSafeInteger) || math.Trunc(value) != value { + return 0, errors.New("invalid Runtime history byte value") + } + return uint64(value), nil +} + +func addCoverage(value *coverageAggregate, sample *rawSample) { + value.observations++ + if sample.status == runtimeobs.StatusObserved { + value.observed++ + } else { + value.unavailable++ + } + if value.first.IsZero() || sample.resolvedAt.Before(value.first) { + value.first = sample.resolvedAt + } + if value.last.IsZero() || sample.resolvedAt.After(value.last) { + value.last = sample.resolvedAt + } +} + +func bucketIndex(query runtimehistory.Query, value time.Time) int { + if value.Before(query.Start) || !value.Before(query.End) { + return -1 + } + return int(value.Sub(query.Start) / query.Step) +} + +func bucketBounds(query runtimehistory.Query, index int) (time.Time, time.Time) { + start := query.Start.Add(time.Duration(index) * query.Step) + end := start.Add(query.Step) + if end.After(query.End) { + end = query.End + } + return start, end +} + +func sortedIndexes[T any](values map[int]*T) []int { + result := make([]int, 0, len(values)) + for index := range values { + result = append(result, index) + } + sort.Ints(result) + return result +} + +func sameOptionalTime(left, right *time.Time) bool { + return left == nil && right == nil || left != nil && right != nil && left.Equal(*right) +} diff --git a/services/agents-api/internal/runtimehistory/postgresreader/aggregate_test.go b/services/agents-api/internal/runtimehistory/postgresreader/aggregate_test.go new file mode 100644 index 000000000..95ae8761a --- /dev/null +++ b/services/agents-api/internal/runtimehistory/postgresreader/aggregate_test.go @@ -0,0 +1,94 @@ +package postgresreader + +import ( + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" + "testing" + "time" +) + +func TestAggregateSeriesUsesProviderObservationOrder(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + query := testQuery(start) + startedAt := start.Add(-time.Minute) + firstObserved, secondObserved := start.Add(5*time.Second), start.Add(15*time.Second) + firstResolved, secondResolved := start.Add(20*time.Second), start.Add(10*time.Second) + samples := []*rawSample{ + { + resolvedAt: secondResolved, observedAt: &secondObserved, allocation: testAllocation, startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, + metrics: map[string]float64{CPUUsageName: 3, CPUCapacityName: 2, MemoryUsageName: 200}, + }, + { + resolvedAt: firstResolved, observedAt: &firstObserved, allocation: testAllocation, startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, + metrics: map[string]float64{CPUUsageName: 1, CPUCapacityName: 2, MemoryUsageName: 100}, + }, + } + points, err := aggregateSeries(query, samples) + if err != nil { + t.Fatal(err) + } + if len(points) != 1 || points[0].CPUUtilizationRatio == nil || *points[0].CPUUtilizationRatio != .1 || points[0].MemoryUsageBytes == nil || *points[0].MemoryUsageBytes != 200 { + t.Fatalf("provider observation order was not preserved: %+v", points) + } +} + +func TestAggregateSeriesLeavesCPUUsageGapWhenEitherEndpointLacksCapacity(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + query := testQuery(start) + startedAt := start.Add(-time.Minute) + firstObserved, secondObserved, thirdObserved := start.Add(-5*time.Second), start.Add(5*time.Second), start.Add(15*time.Second) + samples := []*rawSample{ + { + resolvedAt: firstObserved, observedAt: &firstObserved, allocation: testAllocation, startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, + metrics: map[string]float64{CPUUsageName: 1, CPUCapacityName: 2}, + }, + { + resolvedAt: secondObserved, observedAt: &secondObserved, allocation: testAllocation, startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, + metrics: map[string]float64{CPUUsageName: 3}, + }, + { + resolvedAt: thirdObserved, observedAt: &thirdObserved, allocation: testAllocation, startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, + metrics: map[string]float64{CPUUsageName: 5, CPUCapacityName: 2}, + }, + } + points, err := aggregateSeries(query, samples) + if err != nil { + t.Fatal(err) + } + if len(points) != 1 || points[0].CPUUtilizationRatio != nil { + t.Fatalf("missing CPU capacity was synthesized into utilization: %+v", points) + } +} + +func TestAggregateIgnoresLookbackOnlySeriesForSeriesLimit(t *testing.T) { + start := time.Date(2026, 9, 23, 8, 0, 0, 0, time.UTC) + query := testQuery(start) + query.MaximumSeries = 64 + raw := make([]*rawSample, 0, 66) + for range 65 { + startedAt := start.Add(-time.Minute) + observedAt := start.Add(-2 * time.Second) + raw = append(raw, &rawSample{ + resolvedAt: start.Add(-time.Second), observedAt: &observedAt, allocation: uuid.NewString(), startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, metrics: map[string]float64{}, + }) + } + startedAt := start.Add(-time.Minute) + observedAt := start.Add(5 * time.Second) + raw = append(raw, &rawSample{ + resolvedAt: start.Add(6 * time.Second), observedAt: &observedAt, allocation: uuid.NewString(), startedAt: &startedAt, + provider: "docker", status: runtimeobs.StatusObserved, hasSample: true, metrics: map[string]float64{}, + }) + result, err := aggregate(query, start.Add(time.Minute), raw) + if err != nil { + t.Fatal(err) + } + if len(result.Series) != 1 || len(result.Series[0].Points) != 1 { + t.Fatalf("lookback-only incarnations consumed returned-series budget: %+v", result.Series) + } +} diff --git a/services/agents-api/internal/runtimehistory/postgresreader/exporter.go b/services/agents-api/internal/runtimehistory/postgresreader/exporter.go new file mode 100644 index 000000000..dc8459c83 --- /dev/null +++ b/services/agents-api/internal/runtimehistory/postgresreader/exporter.go @@ -0,0 +1,97 @@ +package postgresreader + +import ( + "context" + "errors" + "math" + "regexp" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" +) + +var providerPattern = regexp.MustCompile(`^[a-z][a-z0-9_]{0,31}$`) + +func (r *Reader) Export(ctx context.Context, record runtimeobs.ExportRecord) error { + if record.CollectionSource != runtimeobs.CollectionSourcePeriodic { + return nil + } + if err := validateRecord(record); err != nil { + return err + } + if record.ResolvedAt.Before(r.now().Add(-retention)) { + return nil + } + // Counters without a provider incarnation cannot safely participate in rates. + if record.Sample != nil && record.Sample.StartedAt == nil { + value := *record.Sample + value.CPUUsageSecondsTotal = nil + value.CPUCapacityCores = nil + value.MemoryUsageBytes = nil + value.MemoryLimitBytes = nil + record.Sample = &value + } + ctx, cancel := context.WithTimeout(ctx, r.queryTimeout) + defer cancel() + if err := r.store.InsertRuntimeHistorySample(ctx, record); err != nil { + return errors.New("persist Runtime history") + } + return nil +} + +func validateRecord(record runtimeobs.ExportRecord) error { + invalid := errors.New("invalid Runtime history sample") + for _, value := range []string{record.TenantID, record.SessionID, record.EnvironmentID} { + if !validID(value) { + return invalid + } + } + if record.AllocationID != "" && !validID(record.AllocationID) { + return invalid + } + if record.Mode != runtimeobs.ModeManaged || record.CollectionSource != runtimeobs.CollectionSourcePeriodic || !validTime(record.ResolvedAt) || record.ProviderType != "" && !providerPattern.MatchString(record.ProviderType) { + return invalid + } + if record.Status != runtimeobs.StatusObserved && record.Status != runtimeobs.StatusUnavailable && record.Status != runtimeobs.StatusUnsupported { + return invalid + } + if (record.Status == runtimeobs.StatusObserved) != (record.Sample != nil) { + return invalid + } + if value := record.Sample; value != nil { + if !validTime(value.ObservedAt) || value.ObservedAt.After(record.ResolvedAt) { + return invalid + } + if value.StartedAt != nil && (!validTime(*value.StartedAt) || value.StartedAt.After(value.ObservedAt) || record.AllocationID == "" || record.ProviderType == "") { + return invalid + } + if value.CPUUsageSecondsTotal != nil && (!validFloat(*value.CPUUsageSecondsTotal) || *value.CPUUsageSecondsTotal < 0) { + return invalid + } + if value.CPUCapacityCores != nil && (!validFloat(*value.CPUCapacityCores) || *value.CPUCapacityCores <= 0) { + return invalid + } + if value.MemoryUsageBytes != nil && *value.MemoryUsageBytes > maxSafeInteger { + return invalid + } + if value.MemoryLimitBytes != nil && (*value.MemoryLimitBytes == 0 || *value.MemoryLimitBytes > maxSafeInteger) { + return invalid + } + } + if value := record.TokenUsage; value != nil && (value.InputTokens > maxSafeInteger || value.OutputTokens > maxSafeInteger) { + return invalid + } + return nil +} + +const maxSafeInteger = uint64(1<<53 - 1) + +func validID(value string) bool { + parsed, err := uuid.Parse(value) + return err == nil && parsed != uuid.Nil && parsed.String() == value +} +func validTime(value time.Time) bool { + return !value.IsZero() && value.UnixNano() >= 0 && value.Equal(time.Unix(0, value.UnixNano())) +} +func validFloat(value float64) bool { return !math.IsNaN(value) && !math.IsInf(value, 0) } diff --git a/services/agents-api/internal/runtimehistory/postgresreader/reader.go b/services/agents-api/internal/runtimehistory/postgresreader/reader.go new file mode 100644 index 000000000..7cced116c --- /dev/null +++ b/services/agents-api/internal/runtimehistory/postgresreader/reader.go @@ -0,0 +1,156 @@ +package postgresreader + +import ( + "context" + "errors" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" +) + +// This bound accommodates a full 24 hours at five-second cadence, independently +// of the requested chart resolution, including the CPU lookback. +const maximumRawSamples = 20_000 +const retention = 7 * 24 * time.Hour + +type sampleStore interface { + InsertRuntimeHistorySample(context.Context, runtimeobs.ExportRecord) error + ListRuntimeHistorySamples(context.Context, string, string, string, int64, int64, int32) ([]runtimeobs.ExportRecord, error) + PruneRuntimeHistorySamples(context.Context, int64) (int64, error) +} + +type Config struct { + Capabilities runtimehistory.Capabilities + QueryTimeout time.Duration +} + +type Reader struct { + store sampleStore + capabilities runtimehistory.Capabilities + queryTimeout time.Duration + now func() time.Time +} + +func New(store sampleStore, config Config) (*Reader, error) { + if store == nil || config.QueryTimeout <= 0 || config.Capabilities.Validate() != nil || config.Capabilities.Retention != retention || config.Capabilities.MaximumRange > 24*time.Hour { + return nil, errors.New("invalid Runtime history PostgreSQL configuration") + } + capabilities := config.Capabilities + capabilities.Metrics = append([]runtimehistory.Metric(nil), capabilities.Metrics...) + return &Reader{store: store, capabilities: capabilities, queryTimeout: config.QueryTimeout, now: time.Now}, nil +} + +func (r *Reader) Capabilities() runtimehistory.Capabilities { + value := r.capabilities + value.Metrics = append([]runtimehistory.Metric(nil), value.Metrics...) + return value +} + +func (r *Reader) Query(ctx context.Context, query runtimehistory.Query) (runtimehistory.Result, error) { + if err := r.validateQuery(query); err != nil { + return runtimehistory.Result{}, err + } + generatedAt := r.now().UTC() + lookback := r.capabilities.MinimumStep + if r.capabilities.SampleInterval > 0 { + lookback = 2 * r.capabilities.SampleInterval + } + rawStart := query.Start.Add(-lookback) + if oldest := generatedAt.Add(-retention); rawStart.Before(oldest) { + rawStart = oldest + } + if rawStart.UnixNano() < 0 { + rawStart = time.Unix(0, 0).UTC() + } + queryCtx, cancel := context.WithTimeout(ctx, r.queryTimeout) + defer cancel() + records, err := r.store.ListRuntimeHistorySamples(queryCtx, query.TenantID, query.SessionID, query.EnvironmentID, rawStart.UnixNano(), query.End.UnixNano(), maximumRawSamples+1) + if err != nil { + return runtimehistory.Result{}, errors.New("query Runtime history") + } + if len(records) > maximumRawSamples { + return runtimehistory.Result{}, errors.New("Runtime history raw result exceeds limit") + } + raw := make([]*rawSample, 0, len(records)) + tokens := make([]rawTokenUsage, 0, len(records)) + for _, record := range records { + if validateRecord(record) != nil || record.TenantID != query.TenantID || record.SessionID != query.SessionID || record.EnvironmentID != query.EnvironmentID || record.ResolvedAt.Before(rawStart) || !record.ResolvedAt.Before(query.End) { + return runtimehistory.Result{}, errors.New("invalid Runtime history stored sample") + } + sample := &rawSample{resolvedAt: record.ResolvedAt, allocation: record.AllocationID, provider: record.ProviderType, status: record.Status, hasSample: true, metrics: map[string]float64{}} + if value := record.Sample; value != nil { + sample.observedAt = &value.ObservedAt + sample.startedAt = value.StartedAt + if value.CPUUsageSecondsTotal != nil { + sample.metrics[CPUUsageName] = *value.CPUUsageSecondsTotal + } + if value.CPUCapacityCores != nil { + sample.metrics[CPUCapacityName] = *value.CPUCapacityCores + } + if value.MemoryUsageBytes != nil { + sample.metrics[MemoryUsageName] = float64(*value.MemoryUsageBytes) + } + if value.MemoryLimitBytes != nil { + sample.metrics[MemoryLimitName] = float64(*value.MemoryLimitBytes) + } + } + raw = append(raw, sample) + if value := record.TokenUsage; value != nil && !record.ResolvedAt.Before(query.Start) { + tokens = append(tokens, rawTokenUsage{resolvedAt: record.ResolvedAt, inputTokens: &value.InputTokens, outputTokens: &value.OutputTokens}) + } + } + result, err := aggregate(query, generatedAt, raw) + if err != nil { + return runtimehistory.Result{}, err + } + for _, metric := range r.capabilities.Metrics { + if metric == runtimehistory.MetricTokens { + result.TokenUsage, err = aggregateTokenUsage(query, tokens) + break + } + } + return result, err +} + +// Prune expires old telemetry even when no Runtime is being sampled. Each call +// deletes at most 4,096 rows in small lock-skipping batches under one deadline. +func (r *Reader) Prune(ctx context.Context) error { + ctx, cancel := context.WithTimeout(ctx, r.queryTimeout) + defer cancel() + before := r.now().Add(-retention).UnixNano() + for range 16 { + count, err := r.store.PruneRuntimeHistorySamples(ctx, before) + if err != nil { + return errors.New("prune Runtime history") + } + if count < 256 { + return nil + } + } + return nil +} + +func (r *Reader) validateQuery(query runtimehistory.Query) error { + for _, value := range []string{query.TenantID, query.SessionID, query.EnvironmentID} { + parsed, err := uuid.Parse(value) + if err != nil || parsed == uuid.Nil || parsed.String() != value { + return errors.New("invalid Runtime history PostgreSQL scope") + } + } + if query.Start.IsZero() || query.End.IsZero() || !query.End.After(query.Start) || query.Start.Nanosecond() != 0 || query.End.Nanosecond() != 0 || + query.Step < r.capabilities.MinimumStep || query.Step%time.Second != 0 || query.Retention != r.capabilities.Retention || + query.MaxPoints < 2 || query.MaxPoints > r.capabilities.MaximumPoints || query.MaximumSeries != r.capabilities.MaximumSeries || + query.MaximumTotalPoints != r.capabilities.MaximumTotalPoints || query.End.Sub(query.Start) > r.capabilities.MaximumRange { + return errors.New("invalid Runtime history PostgreSQL query") + } + buckets := query.End.Sub(query.Start) / query.Step + if query.End.Sub(query.Start)%query.Step != 0 { + buckets++ + } + if buckets > time.Duration(query.MaxPoints) { + return errors.New("Runtime history PostgreSQL query exceeds point budget") + } + return nil +} diff --git a/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go b/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go new file mode 100644 index 000000000..fbe54271e --- /dev/null +++ b/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go @@ -0,0 +1,195 @@ +package postgresreader + +import ( + "context" + "errors" + "strings" + "testing" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" +) + +const ( + testTenant = "11111111-1111-4111-8111-111111111111" + testSession = "22222222-2222-4222-8222-222222222222" + testEnvironment = "33333333-3333-4333-8333-333333333333" + testAllocation = "44444444-4444-4444-8444-444444444444" +) + +func testCapabilities() runtimehistory.Capabilities { + return runtimehistory.Capabilities{ + CollectionMode: runtimehistory.CollectionPeriodic, SampleInterval: 30 * time.Second, + Retention: 7 * 24 * time.Hour, MinimumStep: 30 * time.Second, MaximumRange: 24 * time.Hour, + MaximumPoints: 1_000, MaximumSeries: 64, MaximumTotalPoints: 10_000, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory}, + } +} + +func testQuery(start time.Time) runtimehistory.Query { + return runtimehistory.Query{ + Scope: runtimehistory.Scope{TenantID: testTenant, SessionID: testSession, EnvironmentID: testEnvironment}, + Start: start, End: start.Add(time.Minute), Step: 30 * time.Second, Retention: 7 * 24 * time.Hour, + MaxPoints: 2, MaximumSeries: 64, MaximumTotalPoints: 10_000, + } +} + +type fakeStore struct { + records []runtimeobs.ExportRecord + written []runtimeobs.ExportRecord + err error + limit int32 + scope runtimehistory.Scope + start, end int64 + pruneCalls int + pruneCount int64 +} + +func (s *fakeStore) InsertRuntimeHistorySample(_ context.Context, r runtimeobs.ExportRecord) error { + s.written = append(s.written, r) + return s.err +} +func (s *fakeStore) ListRuntimeHistorySamples(_ context.Context, tenant, session, environment string, start, end int64, limit int32) ([]runtimeobs.ExportRecord, error) { + s.scope = runtimehistory.Scope{TenantID: tenant, SessionID: session, EnvironmentID: environment} + s.start = start + s.end = end + s.limit = limit + return s.records, s.err +} +func (s *fakeStore) PruneRuntimeHistorySamples(context.Context, int64) (int64, error) { + s.pruneCalls++ + return s.pruneCount, s.err +} +func testReader(t *testing.T, s *fakeStore, now time.Time) *Reader { + t.Helper() + caps := testCapabilities() + caps.Metrics = append(caps.Metrics, runtimehistory.MetricTokens) + r, err := New(s, Config{Capabilities: caps, QueryTimeout: time.Second}) + if err != nil { + t.Fatal(err) + } + r.now = func() time.Time { return now } + return r +} +func observedRecord(at, started time.Time, cpu float64) runtimeobs.ExportRecord { + capacity := 2.0 + memory := uint64(123) + return runtimeobs.ExportRecord{TenantID: testTenant, SessionID: testSession, EnvironmentID: testEnvironment, AllocationID: testAllocation, + Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusObserved, CollectionSource: runtimeobs.CollectionSourcePeriodic, + ResolvedAt: at, Sample: &runtimeobs.Sample{ObservedAt: at, StartedAt: &started, CPUUsageSecondsTotal: &cpu, CPUCapacityCores: &capacity, MemoryUsageBytes: &memory}, + TokenUsage: &runtimeobs.TokenUsage{InputTokens: 12, OutputTokens: 3}} +} +func TestReaderDenseInputBudgetIndependentOfChartPoints(t *testing.T) { + end := time.Now().UTC().Truncate(time.Second) + start := end.Add(-24 * time.Hour) + started := start.Add(-time.Hour) + s := &fakeStore{} + for at := start; at.Before(end); at = at.Add(5 * time.Second) { + s.records = append(s.records, observedRecord(at, started, at.Sub(start).Seconds())) + } + r := testReader(t, s, end) + for _, points := range []int{2, 60, 1000} { + q := testQuery(start) + q.End = end + q.MaxPoints = points + q.Step = time.Duration((86400+points-1)/points) * time.Second + result, err := r.Query(t.Context(), q) + if err != nil { + t.Fatal(err) + } + if s.limit != maximumRawSamples+1 || s.scope != q.Scope || s.start != start.Add(-time.Minute).UnixNano() { + t.Fatal("read bounds changed", s) + } + count := 0 + for _, point := range result.Coverage { + count += point.ObservationCount + } + if count != 17280 || len(result.Series) != 1 || len(result.TokenUsage) != len(result.Coverage) || len(result.Coverage) > points { + t.Fatal("dense downsampling dropped samples", count, len(result.Coverage)) + } + } + s.records = make([]runtimeobs.ExportRecord, maximumRawSamples+1) + if _, err := r.Query(t.Context(), testQuery(start)); err == nil { + t.Fatal("raw overflow accepted") + } +} +func TestReaderFailsClosedAndScopesEveryRead(t *testing.T) { + start := time.Now().UTC().Truncate(time.Second).Add(-time.Minute) + s := &fakeStore{err: errors.New("private backend address")} + r := testReader(t, s, start.Add(time.Minute)) + if _, err := r.Query(t.Context(), testQuery(start)); err == nil || strings.Contains(err.Error(), "private") { + t.Fatal(err) + } + s.err = nil + s.records = []runtimeobs.ExportRecord{observedRecord(start, start.Add(-time.Minute), 0)} + s.records[0].TenantID = testSession + if _, err := r.Query(t.Context(), testQuery(start)); err == nil { + t.Fatal("foreign row accepted") + } + q := testQuery(start) + q.EnvironmentID = "" + if _, err := r.Query(t.Context(), q); err == nil { + t.Fatal("unscoped query accepted") + } +} +func TestExporterPersistsOnlyPeriodicAndPreservesUnknownMetrics(t *testing.T) { + now := time.Now().UTC() + s := &fakeStore{} + r := testReader(t, s, now) + record := observedRecord(now.Add(-time.Second), now.Add(-time.Hour), 0) + record.CollectionSource = runtimeobs.CollectionSourceOnRead + if err := r.Export(t.Context(), record); err != nil || len(s.written) != 0 { + t.Fatal(err) + } + record.CollectionSource = runtimeobs.CollectionSourcePeriodic + record.Sample.StartedAt = nil + if err := r.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + if len(s.written) != 1 || s.written[0].Sample.CPUUsageSecondsTotal != nil || s.written[0].TokenUsage.InputTokens != 12 { + t.Fatal(s.written) + } + record.Sample = nil + record.Status = runtimeobs.StatusUnavailable + if err := r.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + if s.written[1].Sample != nil || s.written[1].TokenUsage == nil { + t.Fatal("unavailable observation lost canonical usage") + } + record.ResolvedAt = now.Add(-retention - time.Second) + if err := r.Export(t.Context(), record); err != nil || len(s.written) != 2 { + t.Fatal("expired record persisted", err) + } +} +func TestPruneHasBoundedWorkAndSanitizedFailure(t *testing.T) { + s := &fakeStore{pruneCount: 256} + r := testReader(t, s, time.Now()) + if err := r.Prune(t.Context()); err != nil || s.pruneCalls != 16 { + t.Fatal(err, s.pruneCalls) + } + s.err = errors.New("secret") + if err := r.Prune(t.Context()); err == nil || strings.Contains(err.Error(), "secret") { + t.Fatal(err) + } +} +func TestRestartFencesCPUWithoutSplittingAllocation(t *testing.T) { + start := time.Now().UTC().Truncate(time.Second).Add(-time.Minute) + firstStart := start.Add(-time.Hour) + nextStart := start.Add(15 * time.Second) + s := &fakeStore{records: []runtimeobs.ExportRecord{observedRecord(start, firstStart, 1), observedRecord(start.Add(10*time.Second), firstStart, 11), observedRecord(start.Add(30*time.Second), nextStart, 100), observedRecord(start.Add(40*time.Second), nextStart, 110)}} + r := testReader(t, s, start.Add(time.Minute)) + result, err := r.Query(t.Context(), testQuery(start)) + if err != nil { + t.Fatal(err) + } + if len(result.Series) != 1 || !result.Series[0].StartedAt.Equal(firstStart) || len(result.Series[0].Points) != 2 { + t.Fatal(result) + } + for _, point := range result.Series[0].Points { + if point.CPUUtilizationRatio == nil || *point.CPUUtilizationRatio != 0.5 { + t.Fatal("CPU bridged native restart", point) + } + } +} diff --git a/services/agents-api/internal/store/runtime_history.go b/services/agents-api/internal/store/runtime_history.go new file mode 100644 index 000000000..36c21fc7e --- /dev/null +++ b/services/agents-api/internal/store/runtime_history.go @@ -0,0 +1,132 @@ +package store + +import ( + "context" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/db/sqlc" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/google/uuid" + "github.com/jackc/pgx/v5/pgtype" +) + +// InsertRuntimeHistorySample copies a sanitized periodic observation. Canonical +// Session/Turn Usage remains the accounting authority; these counters are only +// sampled chart values. Deleted or mismatched owners cannot create orphan rows. +func (s *Store) InsertRuntimeHistorySample(ctx context.Context, record runtimeobs.ExportRecord) error { + tenant, err := parseID(record.TenantID) + if err != nil { + return err + } + session, err := parseID(record.SessionID) + if err != nil { + return err + } + environment, err := parseID(record.EnvironmentID) + if err != nil { + return err + } + params := sqlc.InsertRuntimeHistorySampleParams{ + TenantID: tenant, SessionID: session, EnvironmentID: environment, + ResolvedAtNs: record.ResolvedAt.UnixNano(), ProviderType: record.ProviderType, Status: string(record.Status), + } + if record.AllocationID != "" { + params.AllocationID, err = parseID(record.AllocationID) + if err != nil { + return err + } + } + if sample := record.Sample; sample != nil { + params.ObservedAtNs = pgtype.Int8{Int64: sample.ObservedAt.UnixNano(), Valid: true} + if sample.StartedAt != nil { + params.StartedAtNs = pgtype.Int8{Int64: sample.StartedAt.UnixNano(), Valid: true} + } + params.CpuUsageSeconds = historyFloat(sample.CPUUsageSecondsTotal) + params.CpuCapacityCores = historyFloat(sample.CPUCapacityCores) + params.MemoryUsageBytes = historyInteger(sample.MemoryUsageBytes) + params.MemoryLimitBytes = historyInteger(sample.MemoryLimitBytes) + } + if usage := record.TokenUsage; usage != nil { + params.InputTokens = historyInteger(&usage.InputTokens) + params.OutputTokens = historyInteger(&usage.OutputTokens) + } + return s.queries.InsertRuntimeHistorySample(ctx, params) +} + +func (s *Store) ListRuntimeHistorySamples(ctx context.Context, tenantID, sessionID, environmentID string, startNS, endNS int64, limit int32) ([]runtimeobs.ExportRecord, error) { + tenant, err := parseID(tenantID) + if err != nil { + return nil, err + } + session, err := parseID(sessionID) + if err != nil { + return nil, err + } + environment, err := parseID(environmentID) + if err != nil { + return nil, err + } + rows, err := s.queries.ListRuntimeHistorySamples(ctx, sqlc.ListRuntimeHistorySamplesParams{ + TenantID: tenant, SessionID: session, EnvironmentID: environment, StartNs: startNS, EndNs: endNS, RowLimit: limit, + }) + if err != nil { + return nil, err + } + records := make([]runtimeobs.ExportRecord, 0, len(rows)) + for _, row := range rows { + record := runtimeobs.ExportRecord{ + TenantID: tenantID, SessionID: sessionID, EnvironmentID: environmentID, + Mode: runtimeobs.ModeManaged, CollectionSource: runtimeobs.CollectionSourcePeriodic, + ResolvedAt: time.Unix(0, row.ResolvedAtNs).UTC(), ProviderType: row.ProviderType, Status: runtimeobs.Status(row.Status), + } + if row.AllocationID.Valid { + record.AllocationID = uuid.UUID(row.AllocationID.Bytes).String() + } + if row.ObservedAtNs.Valid { + record.Sample = &runtimeobs.Sample{ + ObservedAt: time.Unix(0, row.ObservedAtNs.Int64).UTC(), + CPUUsageSecondsTotal: historyFloatPointer(row.CpuUsageSeconds), CPUCapacityCores: historyFloatPointer(row.CpuCapacityCores), + MemoryUsageBytes: historyIntegerPointer(row.MemoryUsageBytes), MemoryLimitBytes: historyIntegerPointer(row.MemoryLimitBytes), + } + if row.StartedAtNs.Valid { + value := time.Unix(0, row.StartedAtNs.Int64).UTC() + record.Sample.StartedAt = &value + } + } + if row.InputTokens.Valid && row.OutputTokens.Valid { + record.TokenUsage = &runtimeobs.TokenUsage{InputTokens: uint64(row.InputTokens.Int64), OutputTokens: uint64(row.OutputTokens.Int64)} + } + records = append(records, record) + } + return records, nil +} + +func (s *Store) PruneRuntimeHistorySamples(ctx context.Context, beforeNS int64) (int64, error) { + return s.queries.PruneRuntimeHistorySamples(ctx, beforeNS) +} + +func historyFloat(value *float64) pgtype.Float8 { + if value == nil { + return pgtype.Float8{} + } + return pgtype.Float8{Float64: *value, Valid: true} +} +func historyInteger(value *uint64) pgtype.Int8 { + if value == nil { + return pgtype.Int8{} + } + return pgtype.Int8{Int64: int64(*value), Valid: true} +} +func historyFloatPointer(value pgtype.Float8) *float64 { + if !value.Valid { + return nil + } + return &value.Float64 +} +func historyIntegerPointer(value pgtype.Int8) *uint64 { + if !value.Valid { + return nil + } + result := uint64(value.Int64) + return &result +} diff --git a/services/agents-api/migrations/000056_runtime_history.sql b/services/agents-api/migrations/000056_runtime_history.sql new file mode 100644 index 000000000..b4737a233 --- /dev/null +++ b/services/agents-api/migrations/000056_runtime_history.sql @@ -0,0 +1,36 @@ +-- +goose Up +-- Periodic telemetry is a bounded projection, never execution or Usage authority. +CREATE TABLE runtime_history_samples ( + tenant_id uuid NOT NULL, + session_id uuid NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + environment_id uuid NOT NULL REFERENCES environments(id) ON DELETE CASCADE, + resolved_at_ns bigint NOT NULL CHECK (resolved_at_ns >= 0), + allocation_id uuid, + provider_type text NOT NULL, + status text NOT NULL CHECK (status IN ('observed', 'unavailable', 'unsupported')), + observed_at_ns bigint, + started_at_ns bigint, + cpu_usage_seconds double precision, + cpu_capacity_cores double precision, + memory_usage_bytes bigint, + memory_limit_bytes bigint, + input_tokens bigint, + output_tokens bigint, + PRIMARY KEY (tenant_id, session_id, environment_id, resolved_at_ns), + CHECK ((input_tokens IS NULL) = (output_tokens IS NULL)), + CHECK (input_tokens >= 0 AND input_tokens <= 9007199254740991), + CHECK (output_tokens >= 0 AND output_tokens <= 9007199254740991), + CHECK (observed_at_ns >= 0 AND observed_at_ns <= resolved_at_ns), + CHECK (started_at_ns >= 0 AND started_at_ns <= observed_at_ns), + CHECK ((status = 'observed') = (observed_at_ns IS NOT NULL)), + CHECK (started_at_ns IS NULL OR observed_at_ns IS NOT NULL), + CHECK (cpu_usage_seconds >= 0 AND cpu_usage_seconds < 'Infinity'::double precision), + CHECK (cpu_capacity_cores > 0 AND cpu_capacity_cores < 'Infinity'::double precision), + CHECK (memory_usage_bytes >= 0 AND memory_usage_bytes <= 9007199254740991), + CHECK (memory_limit_bytes > 0 AND memory_limit_bytes <= 9007199254740991), + CHECK (started_at_ns IS NOT NULL OR (cpu_usage_seconds IS NULL AND cpu_capacity_cores IS NULL AND memory_usage_bytes IS NULL AND memory_limit_bytes IS NULL)) +); +CREATE INDEX runtime_history_samples_retention_idx ON runtime_history_samples (resolved_at_ns); + +-- +goose Down +DROP TABLE runtime_history_samples; From 564b59ff9d0a6edfbd500efd93cb8231683e8a3d Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:45:33 +0800 Subject: [PATCH 44/51] Remove unused ClickHouse dependency closure --- go.mod | 10 ---------- go.sum | 22 ---------------------- 2 files changed, 32 deletions(-) diff --git a/go.mod b/go.mod index 5b8fdafe7..026a7c9c3 100644 --- a/go.mod +++ b/go.mod @@ -4,7 +4,6 @@ go 1.25.13 require ( github.com/BurntSushi/toml v1.6.0 - github.com/ClickHouse/clickhouse-go/v2 v2.47.0 github.com/containerd/errdefs v1.0.0 github.com/go-chi/chi/v5 v5.3.2 github.com/google/uuid v1.6.0 @@ -59,24 +58,15 @@ require ( ) require ( - github.com/ClickHouse/ch-go v0.73.0 // indirect - github.com/andybalholm/brotli v1.2.2 // indirect github.com/cenkalti/backoff/v5 v5.0.3 // indirect - github.com/go-faster/city v1.0.1 // indirect - github.com/go-faster/errors v0.7.1 // indirect github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 // indirect github.com/jackc/pgpassfile v1.0.0 // indirect github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect github.com/jackc/puddle/v2 v2.2.2 // indirect - github.com/klauspost/compress v1.19.1 // indirect github.com/mfridman/interpolate v0.0.2 // indirect - github.com/paulmach/orb v0.13.0 // indirect - github.com/pierrec/lz4/v4 v4.1.27 // indirect github.com/rogpeppe/go-internal v1.15.0 // indirect github.com/sethvargo/go-retry v0.4.0 // indirect - github.com/shopspring/decimal v1.4.0 // indirect go.uber.org/multierr v1.11.0 // indirect - go.yaml.in/yaml/v3 v3.0.4 // indirect golang.org/x/sync v0.22.0 // indirect golang.org/x/text v0.41.0 // indirect golang.org/x/time v0.15.0 // indirect diff --git a/go.sum b/go.sum index 8e008635b..d0481a593 100644 --- a/go.sum +++ b/go.sum @@ -1,13 +1,7 @@ github.com/BurntSushi/toml v1.6.0 h1:dRaEfpa2VI55EwlIW72hMRHdWouJeRF7TPYhI+AUQjk= github.com/BurntSushi/toml v1.6.0/go.mod h1:ukJfTF/6rtPPRCnwkur4qwRxa8vTRFBF0uk2lLoLwho= -github.com/ClickHouse/ch-go v0.73.0 h1:jsHiGRbQ3sz+gekvDFJF29LWDo5dzbJm5s1h8TWVP2M= -github.com/ClickHouse/ch-go v0.73.0/go.mod h1:wkFIxrqlXeRJ9cn3r5Fz5Qen9jl5aTMPuGZeuJpANNY= -github.com/ClickHouse/clickhouse-go/v2 v2.47.0 h1:ZDAzrnKSOPTIsm4tdUNfrii2yc8dk4SVRLC77BR7Z5Q= -github.com/ClickHouse/clickhouse-go/v2 v2.47.0/go.mod h1:sPj7C7UYQ2MWHcfX+4eGN6nwnCqwUKfgO6PcwKpd6K8= github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERoyfY= github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU= -github.com/andybalholm/brotli v1.2.2 h1:HzTuoo2ErYQqf5qvcJInB8uvqSVxRttzkFexPWtnceM= -github.com/andybalholm/brotli v1.2.2/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY= github.com/cenkalti/backoff/v5 v5.0.3 h1:ZN+IMa753KfX5hd8vVaMixjnqRZ3y8CuJKRKj1xcsSM= github.com/cenkalti/backoff/v5 v5.0.3/go.mod h1:rkhZdG3JZukswDf7f0cwqPNk4K0sa+F97BxZthm/crw= github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs= @@ -31,10 +25,6 @@ github.com/felixge/httpsnoop v1.1.0 h1:3YtUj32ZZkqZtt3sZZsClsymw/QDuVfpNhoA31zeO github.com/felixge/httpsnoop v1.1.0/go.mod h1:Zqxgdd+1Rkcz8euOqdr7lqgCRJztwr5hp9vDSi5UZCE= github.com/go-chi/chi/v5 v5.3.2 h1:5YQkICvTCSZ25hoRsyJazN0scjzKGiu4VAUc7H1o1nY= github.com/go-chi/chi/v5 v5.3.2/go.mod h1:R+tYY2hNuVUUjxoPtqUdgBqevM9s9njzkTLutVsOCto= -github.com/go-faster/city v1.0.1 h1:4WAxSZ3V2Ws4QRDrscLEDcibJY8uf41H6AhXDrNDcGw= -github.com/go-faster/city v1.0.1/go.mod h1:jKcUJId49qdW3L1qKHH/3wPeUstCVpVSXTM6vO3VcTw= -github.com/go-faster/errors v0.7.1 h1:MkJTnDoEdi9pDabt1dpWf7AA8/BaSYZqibYyhZ20AYg= -github.com/go-faster/errors v0.7.1/go.mod h1:5ySTjWFiphBs07IKuiL69nxdfd5+fzh1u7FPGZP2quo= github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A= github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8= github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY= @@ -62,8 +52,6 @@ github.com/jackc/pgx/v5 v5.10.0 h1:VhSvgU2jSli8o3AqIEOTJr7rZwAEUVo4E4XhR94Zfr0= github.com/jackc/pgx/v5 v5.10.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4= github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo= github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4= -github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk= -github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ= github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE= github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk= github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY= @@ -88,10 +76,6 @@ github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8 github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM= github.com/opencontainers/image-spec v1.1.1 h1:y0fUlFfIZhPF1W537XOLg0/fcx6zcHCJwooC2xJA040= github.com/opencontainers/image-spec v1.1.1/go.mod h1:qpqAh3Dmcf36wStyyWU+kCeDgrGnAve2nCC8+7h8Q0M= -github.com/paulmach/orb v0.13.0 h1:r7n7mQGGF+cj/CbcivEj9J3HGK+XR+yXnvzRdq9saIw= -github.com/paulmach/orb v0.13.0/go.mod h1:6scRWINywA2Jf05dcjOfLfxrUIMECvTSG2MVbRLxu/k= -github.com/pierrec/lz4/v4 v4.1.27 h1:+PhzhWDrjRj89TH2sw43nE3+4+W8lSxIuQadEHZyjUk= -github.com/pierrec/lz4/v4 v4.1.27/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4= github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM= github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= github.com/pressly/goose/v3 v3.27.3 h1:pIglVHjw99r4e/hDHHwbl9vfOsDMqUokfkXo6+n/RxA= @@ -106,8 +90,6 @@ github.com/segmentio/encoding v0.5.4 h1:OW1VRern8Nw6ITAtwSZ7Idrl3MXCFwXHPgqESYfv github.com/segmentio/encoding v0.5.4/go.mod h1:HS1ZKa3kSN32ZHVZ7ZLPLXWvOVIiZtyJnO1gPH1sKt0= github.com/sethvargo/go-retry v0.4.0 h1:9qy1OoIAxBL+gBYnkTnTnWle5wlfsXQlwRzIbbpdqPw= github.com/sethvargo/go-retry v0.4.0/go.mod h1:tvsjdKG6xfiCx4LSiUZ06kcv38xvdVQwv8R6/VnnVWg= -github.com/shopspring/decimal v1.4.0 h1:bxl37RwXBklmTi0C79JfXCEBD1cqqHt0bbgBAGFp81k= -github.com/shopspring/decimal v1.4.0/go.mod h1:gawqmDU56v4yIKSwfBSFip1HdCCXN8/+DMd9qYNcwME= github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI= github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg= @@ -123,8 +105,6 @@ github.com/tidwall/pretty v1.2.1 h1:qjsOFOWWQl+N3RsoF5/ssm1pHmJJwhjlSbZ51I6wMl4= github.com/tidwall/pretty v1.2.1/go.mod h1:ITEVvHYasfjBbM0u2Pg8T2nJnzm8xPwvNhhsoaGGjNU= github.com/tidwall/sjson v1.2.5 h1:kLy8mja+1c9jlljvWTlSazM7cKDRfJuR/bOJhcY5NcY= github.com/tidwall/sjson v1.2.5/go.mod h1:Fvgq9kS/6ociJEDnK0Fk1cpYF4FIW6ZF7LAe+6jwd28= -github.com/xyproto/randomstring v1.0.5 h1:YtlWPoRdgMu3NZtP45drfy1GKoojuR7hmRcnhZqKjWU= -github.com/xyproto/randomstring v1.0.5/go.mod h1:rgmS5DeNXLivK7YprL0pY+lTuhNQW3iGxZ18UQApw/E= github.com/yosida95/uritemplate/v3 v3.0.2 h1:Ed3Oyj9yrmi9087+NczuL5BwkIc4wvTb5zIM+UJPGz4= github.com/yosida95/uritemplate/v3 v3.0.2/go.mod h1:ILOh0sOhIJR3+L/8afwt/kE++YT040gmv5BQTMR2HP4= go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64= @@ -149,8 +129,6 @@ go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpu go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk= go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= go.uber.org/multierr v1.11.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y= -go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc= -go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg= golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M= golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis= golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To= From 78e3e2493bca75c420c99ba02968b5be7a0b85be Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:47:23 +0800 Subject: [PATCH 45/51] Align history schema regression with canonical token metrics --- contracts/agents-api/v1/runtime_history_test.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/contracts/agents-api/v1/runtime_history_test.go b/contracts/agents-api/v1/runtime_history_test.go index bd986d68d..70e8f06b3 100644 --- a/contracts/agents-api/v1/runtime_history_test.go +++ b/contracts/agents-api/v1/runtime_history_test.go @@ -22,14 +22,14 @@ func TestRuntimeHistoryOpenAPICollectionLimits(t *testing.T) { } definitions := openAPIMap(t, document, "definitions") - assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistoryCapabilities", "metrics", 2) + assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistoryCapabilities", "metrics", 3) assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistory", "series", 1000) assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistoryCoverage", "buckets", 10000) assertOpenAPIArrayLimit(t, definitions, "v1.RuntimeHistorySeries", "points", 10000) capabilities := openAPIMap(t, definitions, "v1.RuntimeHistoryCapabilities") metrics := openAPIMap(t, openAPIMap(t, capabilities, "properties"), "metrics") metricItems := openAPIMap(t, metrics, "items") - assertOpenAPIStrings(t, metricItems, "enum", []string{"cpu", "memory"}) + assertOpenAPIStrings(t, metricItems, "enum", []string{"cpu", "memory", "tokens"}) series := openAPIMap(t, definitions, "v1.RuntimeHistorySeries") providerType := openAPIMap(t, openAPIMap(t, series, "properties"), "provider_type") if actual, ok := providerType["pattern"].(string); !ok || actual != "^[a-z][a-z0-9_]{0,31}$" { From 20da3b8d20b9219deb7a069b2fc59de2ef157555 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:47:28 +0800 Subject: [PATCH 46/51] Verify PostgreSQL history and preserve explicit observation gaps --- .../postgresreader/aggregate.go | 26 ++- .../postgresreader/reader_test.go | 22 ++ .../store/runtime_history_acceptance_test.go | 205 ++++++++++++++++++ 3 files changed, 247 insertions(+), 6 deletions(-) create mode 100644 services/agents-api/internal/store/runtime_history_acceptance_test.go diff --git a/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go b/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go index c34d2841c..776753036 100644 --- a/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go +++ b/services/agents-api/internal/runtimehistory/postgresreader/aggregate.go @@ -118,6 +118,15 @@ func aggregate(query runtimehistory.Query, generatedAt time.Time, raw []*rawSamp } value.samples = append(value.samples, sample) } + // Unavailable samples retain allocation identity even without a provider + // timestamp. Keep them in that allocation's coverage and rate sequence. + for _, sample := range raw { + if sample.startedAt == nil && sample.allocation != "" { + if value := series[sample.allocation]; value != nil && !sample.resolvedAt.Before(value.startedAt) { + value.samples = append(value.samples, sample) + } + } + } result := runtimehistory.Result{GeneratedAt: generatedAt, Coverage: make([]runtimehistory.CoveragePoint, 0, len(coverage)), Series: make([]runtimehistory.Series, 0, len(series))} for _, index := range sortedIndexes(coverage) { value := coverage[index] @@ -159,16 +168,17 @@ func aggregate(query runtimehistory.Query, generatedAt time.Time, raw []*rawSamp func aggregateSeries(query runtimehistory.Query, samples []*rawSample) ([]runtimehistory.Point, error) { samples = append([]*rawSample(nil), samples...) sort.SliceStable(samples, func(i, j int) bool { - if samples[i].observedAt == nil { - return samples[j].observedAt != nil + left, right := samples[i].resolvedAt, samples[j].resolvedAt + if samples[i].observedAt != nil { + left = *samples[i].observedAt } - if samples[j].observedAt == nil { - return false + if samples[j].observedAt != nil { + right = *samples[j].observedAt } - if samples[i].observedAt.Equal(*samples[j].observedAt) { + if left.Equal(right) { return samples[i].resolvedAt.Before(samples[j].resolvedAt) } - return samples[i].observedAt.Before(*samples[j].observedAt) + return left.Before(right) }) points := map[int]*pointAggregate{} var previousUsage, previousCapacity *float64 @@ -230,6 +240,10 @@ func aggregateSeries(query runtimehistory.Query, samples []*rawSample) ([]runtim } } } + if !hasUsage || !hasCapacity { + previousUsage, previousCapacity = nil, nil + previousAt = time.Time{} + } if hasUsage { value := usage previousUsage = &value diff --git a/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go b/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go index fbe54271e..c6f4b1171 100644 --- a/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go +++ b/services/agents-api/internal/runtimehistory/postgresreader/reader_test.go @@ -193,3 +193,25 @@ func TestRestartFencesCPUWithoutSplittingAllocation(t *testing.T) { } } } + +func TestUnavailableObservationBreaksCPUContinuity(t *testing.T) { + start := time.Now().UTC().Truncate(time.Second).Add(-time.Minute) + started := start.Add(-time.Hour) + unavailable := observedRecord(start.Add(20*time.Second), started, 0) + unavailable.Sample = nil + unavailable.Status = runtimeobs.StatusUnavailable + s := &fakeStore{records: []runtimeobs.ExportRecord{observedRecord(start.Add(5*time.Second), started, 1), unavailable, observedRecord(start.Add(35*time.Second), started, 100)}} + r := testReader(t, s, start.Add(time.Minute)) + result, err := r.Query(t.Context(), testQuery(start)) + if err != nil { + t.Fatal(err) + } + if len(result.Series) != 1 || len(result.Series[0].Points) != 2 || result.Series[0].Points[0].UnavailableCount != 1 { + t.Fatal("outage missing from allocation coverage", result) + } + for _, point := range result.Series[0].Points { + if point.CPUUtilizationRatio != nil { + t.Fatal("CPU bridged unavailable sample", point) + } + } +} diff --git a/services/agents-api/internal/store/runtime_history_acceptance_test.go b/services/agents-api/internal/store/runtime_history_acceptance_test.go new file mode 100644 index 000000000..23cc3aa7f --- /dev/null +++ b/services/agents-api/internal/store/runtime_history_acceptance_test.go @@ -0,0 +1,205 @@ +package store_test + +import ( + "context" + "encoding/json" + "testing" + "time" + + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimehistory/postgresreader" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/runtimeobs" + "github.com/MiniMax-AI-Dev/parsar/services/agents-api/internal/store" + "github.com/google/uuid" + "github.com/jackc/pgx/v5/pgxpool" +) + +func historyCapabilities() runtimehistory.Capabilities { + return runtimehistory.Capabilities{CollectionMode: runtimehistory.CollectionPeriodic, SampleInterval: 30 * time.Second, + Retention: 7 * 24 * time.Hour, MinimumStep: 30 * time.Second, MaximumRange: 24 * time.Hour, MaximumPoints: 1000, MaximumSeries: 64, MaximumTotalPoints: 10000, + Metrics: []runtimehistory.Metric{runtimehistory.MetricCPU, runtimehistory.MetricMemory, runtimehistory.MetricTokens}} +} +func historyBackend(t *testing.T, s *store.Store) *postgresreader.Reader { + t.Helper() + r, err := postgresreader.New(s, postgresreader.Config{Capabilities: historyCapabilities(), QueryTimeout: 5 * time.Second}) + if err != nil { + t.Fatal(err) + } + return r +} +func historyOwner(t *testing.T, s *store.Store) runtimehistory.Scope { + t.Helper() + tenant := uuid.NewString() + session, err := s.CreateSession(t.Context(), tenant, store.CreateSessionInput{Creator: store.FixtureCreator(), Engine: "codex", IdempotencyKey: uuid.NewString(), Configuration: json.RawMessage(`{"agent":{"model":"fixture-model"},"environment":{"type":"openai_hosted","workspace_directory":"/workspace","capability_directories":[]}}`)}) + if err != nil { + t.Fatal(err) + } + environment, err := s.GetSessionEnvironment(t.Context(), tenant, session.ID) + if err != nil { + t.Fatal(err) + } + return runtimehistory.Scope{TenantID: tenant, SessionID: session.ID, EnvironmentID: environment.ID} +} +func historyRecord(scope runtimehistory.Scope, allocation string, started, at time.Time, cpu float64, input uint64) runtimeobs.ExportRecord { + capacity := 2.0 + memory := uint64(512) + return runtimeobs.ExportRecord{TenantID: scope.TenantID, SessionID: scope.SessionID, EnvironmentID: scope.EnvironmentID, AllocationID: allocation, Mode: runtimeobs.ModeManaged, ProviderType: "docker", Status: runtimeobs.StatusObserved, CollectionSource: runtimeobs.CollectionSourcePeriodic, ResolvedAt: at, + Sample: &runtimeobs.Sample{ObservedAt: at, StartedAt: &started, CPUUsageSecondsTotal: &cpu, CPUCapacityCores: &capacity, MemoryUsageBytes: &memory}, TokenUsage: &runtimeobs.TokenUsage{InputTokens: input, OutputTokens: input / 2}} +} +func historyQuery(scope runtimehistory.Scope, start, end time.Time, points int) runtimehistory.Query { + step := time.Duration((int64(end.Sub(start)/time.Second)+int64(points)-1)/int64(points)) * time.Second + if step < 30*time.Second { + step = 30 * time.Second + } + return runtimehistory.Query{Scope: scope, Start: start, End: end, Step: step, Retention: 7 * 24 * time.Hour, MaxPoints: points, MaximumSeries: 64, MaximumTotalPoints: 10000} +} + +func TestPostgresRuntimeHistoryAcceptance(t *testing.T) { + s, pool := store.NewTestStore(t) + scope := historyOwner(t, s) + foreign := historyOwner(t, s) + reader := historyBackend(t, s) + end := time.Now().UTC().Truncate(time.Second) + start := end.Add(-3 * time.Minute) + started := start.Add(-time.Minute + 123*time.Nanosecond) + allocation := uuid.NewString() + records := []runtimeobs.ExportRecord{ + historyRecord(scope, allocation, started, start.Add(5*time.Second), 0, 10), + historyRecord(scope, allocation, started, start.Add(35*time.Second), 30, 20), + historyRecord(scope, allocation, started, start.Add(95*time.Second), 90, 40), + } + unavailable := historyRecord(scope, allocation, started, start.Add(155*time.Second), 100, 50) + unavailable.Status = runtimeobs.StatusUnavailable + unavailable.Sample = nil + records = append(records, unavailable) + for _, record := range records { + if err := reader.Export(t.Context(), record); err != nil { + t.Fatal(err) + } + } + // Retry is idempotent; on-read and mismatched owners do not enter periodic history. + if err := reader.Export(t.Context(), records[0]); err != nil { + t.Fatal(err) + } + onRead := historyRecord(scope, allocation, started, start.Add(65*time.Second), 60, 30) + onRead.CollectionSource = runtimeobs.CollectionSourceOnRead + if err := reader.Export(t.Context(), onRead); err != nil { + t.Fatal(err) + } + forged := historyRecord(scope, allocation, started, start.Add(75*time.Second), 70, 35) + forged.TenantID = foreign.TenantID + if err := reader.Export(t.Context(), forged); err != nil { + t.Fatal(err) + } + // Another valid scope shares timestamps and cannot leak into the authorized read. + if err := reader.Export(t.Context(), historyRecord(foreign, uuid.NewString(), started, start.Add(65*time.Second), 999, 999)); err != nil { + t.Fatal(err) + } + freshPool, err := pgxpool.NewWithConfig(t.Context(), pool.Config()) + if err != nil { + t.Fatal(err) + } + defer freshPool.Close() + reader = historyBackend(t, store.New(freshPool)) + q := historyQuery(scope, start, end, 6) + result, err := reader.Query(t.Context(), q) + if err != nil { + t.Fatal(err) + } + if len(result.Coverage) != 4 || len(result.Series) != 1 || len(result.TokenUsage) != 4 { + t.Fatalf("durable samples or gaps changed: %+v", result) + } + if !result.Series[0].StartedAt.Equal(started) { + t.Fatal("nanosecond incarnation identity was truncated", result.Series[0].StartedAt, started) + } + if result.Coverage[3].UnavailableCount != 1 || result.TokenUsage[3].InputTokens != 50 { + t.Fatal("resource failure lost canonical usage") + } + if result.Series[0].Points[1].CPUUtilizationRatio == nil || *result.Series[0].Points[1].CPUUtilizationRatio != 0.5 { + t.Fatal("CPU delta changed", result.Series[0].Points) + } + response := runtimehistory.Response{Capabilities: historyCapabilities(), Scope: scope, Requested: runtimehistory.Range{Start: start, End: end, MaxPoints: 6}, Resolution: q.Step, GeneratedAt: result.GeneratedAt, Coverage: result.Coverage, Series: result.Series, TokenUsage: result.TokenUsage} + if err := response.Validate(time.Now()); err != nil { + t.Fatal(err) + } + for _, bad := range []runtimehistory.Scope{{TenantID: foreign.TenantID, SessionID: scope.SessionID, EnvironmentID: scope.EnvironmentID}, {TenantID: scope.TenantID, SessionID: foreign.SessionID, EnvironmentID: scope.EnvironmentID}, {TenantID: scope.TenantID, SessionID: scope.SessionID, EnvironmentID: foreign.EnvironmentID}} { + got, err := reader.Query(t.Context(), historyQuery(bad, start, end, 6)) + if err != nil || len(got.Coverage) != 0 || len(got.TokenUsage) != 0 { + t.Fatal("scope isolation failed", got, err) + } + } + // Chart snapshots do not alter canonical Session Usage. + session, err := s.GetSession(t.Context(), scope.TenantID, scope.SessionID) + if err != nil || len(session.Usage) != 0 && string(session.Usage) != "null" { + t.Fatal("telemetry became accounting authority", string(session.Usage), err) + } +} + +func TestPostgresRuntimeHistoryDenseReadAndBoundedRetention(t *testing.T) { + s, pool := store.NewTestStore(t) + scope := historyOwner(t, s) + reader := historyBackend(t, s) + end := time.Now().UTC().Truncate(time.Second) + start := end.Add(-24 * time.Hour) + started := start.Add(-time.Hour + 987*time.Nanosecond) + base := historyRecord(scope, uuid.NewString(), started, start, 0, 10) + if err := reader.Export(t.Context(), base); err != nil { + t.Fatal(err) + } + // Seed the equivalent of a full day of five-second periodic samples in one + // fixture statement; production inserts still go through the exporter. + _, err := pool.Exec(t.Context(), `INSERT INTO runtime_history_samples + (tenant_id,session_id,environment_id,resolved_at_ns,allocation_id,provider_type,status,observed_at_ns,started_at_ns,cpu_usage_seconds,cpu_capacity_cores,memory_usage_bytes,input_tokens,output_tokens) + SELECT tenant_id,session_id,environment_id,resolved_at_ns+n*5000000000,allocation_id,provider_type,status,observed_at_ns+n*5000000000,started_at_ns,n::double precision,cpu_capacity_cores,memory_usage_bytes,input_tokens+n,output_tokens+n + FROM runtime_history_samples CROSS JOIN generate_series(1,17279) n WHERE tenant_id=$1 AND session_id=$2`, scope.TenantID, scope.SessionID) + if err != nil { + t.Fatal(err) + } + for _, points := range []int{2, 60, 1000} { + got, err := reader.Query(t.Context(), historyQuery(scope, start, end, points)) + if err != nil { + t.Fatal(err) + } + count := 0 + for _, point := range got.Coverage { + count += point.ObservationCount + } + if count != 17280 || len(got.Coverage) > points || len(got.TokenUsage) != len(got.Coverage) { + t.Fatal("raw sample count tied to chart budget", count, len(got.Coverage), points) + } + } + // Expired rows remain invisible before physical pruning runs. + old := historyRecord(scope, base.AllocationID, started.Add(-8*24*time.Hour), start.Add(-8*24*time.Hour), 0, 1) + if err := s.InsertRuntimeHistorySample(t.Context(), old); err != nil { + t.Fatal(err) + } + _, err = pool.Exec(t.Context(), `INSERT INTO runtime_history_samples + (tenant_id,session_id,environment_id,resolved_at_ns,allocation_id,provider_type,status,observed_at_ns,started_at_ns) + SELECT tenant_id,session_id,environment_id,resolved_at_ns-n,allocation_id,provider_type,'unavailable',NULL,NULL + FROM runtime_history_samples CROSS JOIN generate_series(1,4100) n WHERE tenant_id=$1 AND session_id=$2 AND resolved_at_ns=$3`, scope.TenantID, scope.SessionID, old.ResolvedAt.UnixNano()) + if err != nil { + t.Fatal(err) + } + got, err := reader.Query(t.Context(), historyQuery(scope, old.ResolvedAt.Add(-time.Second), old.ResolvedAt.Add(time.Minute), 3)) + if err != nil || len(got.Coverage) != 0 { + t.Fatal("expired samples visible", got, err) + } + if err := reader.Prune(t.Context()); err != nil { + t.Fatal(err) + } + var remaining int + if err := pool.QueryRow(t.Context(), `SELECT count(*) FROM runtime_history_samples WHERE tenant_id=$1 AND resolved_at_ns<$2`, scope.TenantID, end.Add(-7*24*time.Hour).UnixNano()).Scan(&remaining); err != nil || remaining != 5 { + t.Fatal("cleanup exceeded bounded batch", remaining, err) + } + if err := reader.Prune(t.Context()); err != nil { + t.Fatal(err) + } + if err := pool.QueryRow(t.Context(), `SELECT count(*) FROM runtime_history_samples WHERE tenant_id=$1 AND resolved_at_ns<$2`, scope.TenantID, end.Add(-7*24*time.Hour).UnixNano()).Scan(&remaining); err != nil || remaining != 0 { + t.Fatal("retention cleanup incomplete", remaining, err) + } + cancelled, cancel := context.WithCancel(t.Context()) + cancel() + if _, err := reader.Query(cancelled, historyQuery(scope, start, end, 60)); err == nil { + t.Fatal("cancelled query succeeded") + } +} From 1db70b5acb2ab46e4daaaff2d0f04ced0a1db859 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:49:11 +0800 Subject: [PATCH 47/51] Wait for deterministic live Runtime browser samples --- apps/web/e2e/agents-lifecycle.spec.ts | 12 +++++------- 1 file changed, 5 insertions(+), 7 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index bd72d98b7..b7a95dee6 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2631,10 +2631,6 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand metrics: [], }), })); - await page.goto("/"); - const dashboard = page.locator(".dashboard-page"); - await expect(dashboard.getByRole("heading", { name: "Dashboard", exact: true })).toBeVisible(); - await page.route("**/v1/agents/sessions*", async (route) => { if (new URL(route.request().url()).pathname !== "/v1/agents/sessions") return route.fallback(); await route.fulfill({ status: 200, contentType: "application/json", body: JSON.stringify(list(runtimeSessions)) }); @@ -2677,8 +2673,11 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand }); }); + await page.goto("/"); + const dashboard = page.locator(".dashboard-page"); + await expect(dashboard.getByRole("heading", { name: "Dashboard", exact: true })).toBeVisible(); const refresh = dashboard.getByRole("button", { name: "Refresh Dashboard snapshot" }); - await refresh.click(); + await expect(dashboard.locator(".dashboard-runtime-sample-count")).toContainText("1 sample ·"); await expect(dashboard.getByRole("heading", { name: "CPU usage" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Memory usage" })).toBeVisible(); await expect(dashboard.getByRole("heading", { name: "Compute uptime" })).toBeVisible(); @@ -2688,9 +2687,8 @@ test("renders Runtime telemetry as visual snapshot panels with details on demand await expect(liveRange.getByRole("button", { name: "1h" })).toHaveAttribute("aria-pressed", "true"); await liveRange.getByRole("button", { name: "15m" }).click(); await expect(liveRange.getByRole("button", { name: "15m" })).toHaveAttribute("aria-pressed", "true"); - await page.waitForTimeout(20); await refresh.click(); - await page.waitForTimeout(20); + await expect(dashboard.getByLabel("CPU usage: 2 live samples")).toBeVisible(); await refresh.click(); await expect(dashboard.getByLabel("CPU usage: 3 live samples")).toBeVisible(); await expect(dashboard.getByLabel("Memory usage: 3 live samples")).toBeVisible(); From 0c863786cbcf483a4506c1c9f3012e94eabf0716 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:55:58 +0800 Subject: [PATCH 48/51] Keep historical charts limited to supported measurements --- contracts/agents-api/runtime-history-api.md | 6 +++++- contracts/agents-api/runtime-observability-design.md | 7 +++++-- docs/web/README.md | 3 ++- docs/web/README.zh-CN.md | 3 ++- services/agents-api/runtime-history/README.md | 2 +- 5 files changed, 15 insertions(+), 6 deletions(-) diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md index b394998c9..968a7b66e 100644 --- a/contracts/agents-api/runtime-history-api.md +++ b/contracts/agents-api/runtime-history-api.md @@ -91,7 +91,7 @@ period and only from qualified periodic cadence. Resource `series` are keyed only by `allocation_id`. This keeps one continuous Dashboard lifecycle when a provider pauses, restores, restarts, or replaces its underlying compute without replacing the durable allocation. `started_at` remains -the earliest retained provider start estimate for compatible uptime display; it is +the earliest retained provider start estimate; it is not series identity: ```json @@ -168,3 +168,7 @@ or malformed data reject the entire response with a 502 client projection error. - History availability does not imply current Runtime readiness. - Current observations and Durable history have separate freshness and retention semantics and must remain separately labelled in Web. + +Compute uptime is available from current observations only. Retained allocation +series can span compute restarts and unavailable intervals; their earliest start +is not a per-bucket compute start and must not be used to draw an uptime history. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index 0a383e235..d55958231 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -186,7 +186,10 @@ measurement or lifecycle state. | Busy duration | Turn `started_at` to `completed_at` or now | Time model work has been active. | | Idle duration | future durable `idle_since` | Not available in the current design. | -Container restart resets compute uptime but not allocation age. Dashboard labels +Container restart resets compute uptime but not allocation age. Live CPU deltas +require the same known compute start as well as the same allocation. Retained +charts show CPU, memory and tokens; uptime stays in the current/Live view because +the history contract does not supply each bucket's compute start. Dashboard labels must not collapse these values into one generic Runtime duration. ## 9. Collection behavior @@ -341,7 +344,7 @@ Environment scope plus a bounded start, exclusive end, server-selected step, and total point budget. Provider-native identity is never a query input. Reader results remain divided by allocation. Provider `started_at` values are -retained only as compatible display metadata and never split one durable allocation +retained as metadata and never split one durable allocation into multiple Dashboard series. Every bucket reports explicit observation coverage and nullable CPU/memory values. CPU utilization may be derived only from ordered cumulative counters diff --git a/docs/web/README.md b/docs/web/README.md index ee7eaa86c..6f8ee1041 100644 --- a/docs/web/README.md +++ b/docs/web/README.md @@ -38,7 +38,8 @@ PostgreSQL database, with no separate monitoring stack. Web discovers periodic history and offers 1-hour, 6-hour and 24-hour ranges; reload restores data through Core. API-only deployments without a sampler retain the browser-local Live view. Token throughput uses snapshots of canonical Session Usage, preserving missing data. -Web never connects to a database or receives telemetry credentials. +Web never connects to a database or receives telemetry credentials. Compute uptime +is shown only in current/Live observations; History keeps CPU, memory and tokens. When a provider reports only cumulative CPU time, Web derives interval utilization only across adjacent samples from the same verified Runtime incarnation; restarts and counter regressions create diff --git a/docs/web/README.zh-CN.md b/docs/web/README.zh-CN.md index 6ead8d1a9..e679a53a7 100644 --- a/docs/web/README.zh-CN.md +++ b/docs/web/README.zh-CN.md @@ -33,7 +33,8 @@ Dashboard 是默认首页,集中展示当前 Agent 和 Session 结果,并加 Core 默认将周期采样存入现有 PostgreSQL,无需额外监控服务。页面通过 Core 查询 1 小时、6 小时和 24 小时历史,刷新后仍可读取。未启用执行 worker 的部署保留 浏览器本地 Live 视图。Token 速率来自 Session 的实际用量快照,缺失数据不按零计算。 -Web 不直接访问数据库,也不接收监控凭据。 +Web 不直接访问数据库,也不接收监控凭据。运行时长只在当前/Live 视图展示; +历史图表保留 CPU、内存和 Token。 ![Runtime 监控实时趋势](images/runtime-dashboard.png) diff --git a/services/agents-api/runtime-history/README.md b/services/agents-api/runtime-history/README.md index 6dc77e6ed..2e6cba438 100644 --- a/services/agents-api/runtime-history/README.md +++ b/services/agents-api/runtime-history/README.md @@ -29,7 +29,7 @@ sampling or export to an existing OTLP receiver. Sampling-only example: ``` Sampling accepts 5–300 seconds; omitted or zero selects 30. Queue capacity defaults -to 256 (maximum 4096); write/export timeout defaults to two seconds (maximum30). +to 256 (maximum 4096); write/export timeout defaults to two seconds (maximum 30). A server without an execution worker advertises on-read collection and does not claim periodic coverage. Retained queries remain available through Core. From e5b324775f569781e6e84432d3ce252933aeea8e Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:56:15 +0800 Subject: [PATCH 49/51] Clarify retained history start metadata --- contracts/agents-api/v1/runtime_history.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/contracts/agents-api/v1/runtime_history.go b/contracts/agents-api/v1/runtime_history.go index 93304c5a6..49cd5ee5f 100644 --- a/contracts/agents-api/v1/runtime_history.go +++ b/contracts/agents-api/v1/runtime_history.go @@ -67,8 +67,8 @@ type RuntimeHistorySeries struct { Points []RuntimeHistoryPoint `json:"points" binding:"required" validate:"max=10000"` } -// RuntimeHistoryTime is a lossless JSON-safe provider start estimate retained -// for compatible uptime display. AllocationID is the series identity. +// RuntimeHistoryTime preserves the earliest retained provider start estimate. +// It is not a per-bucket compute start. AllocationID is the series identity. type RuntimeHistoryTime struct { Seconds int64 `json:"seconds" binding:"required" minimum:"0" maximum:"9007199254740991"` Nanoseconds int `json:"nanoseconds" binding:"required" minimum:"0" maximum:"999999999"` From 0708ec28b3966c2782662fd1ccb118e019484a33 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:56:47 +0800 Subject: [PATCH 50/51] Fence live CPU deltas across restarts and omit derived historical uptime --- apps/web/e2e/agents-lifecycle.spec.ts | 2 ++ .../features/dashboard/DashboardView.test.tsx | 2 +- .../dashboard/RuntimeTrendCharts.test.tsx | 3 +++ .../features/dashboard/RuntimeTrendCharts.tsx | 2 +- .../dashboard/runtime-history.test.ts | 13 +++++++-- .../src/features/dashboard/runtime-history.ts | 4 +-- .../features/dashboard/runtime-trends.test.ts | 27 ++++++++++++++++--- .../src/features/dashboard/runtime-trends.ts | 4 +++ 8 files changed, 46 insertions(+), 11 deletions(-) diff --git a/apps/web/e2e/agents-lifecycle.spec.ts b/apps/web/e2e/agents-lifecycle.spec.ts index b7a95dee6..2dd5df09a 100644 --- a/apps/web/e2e/agents-lifecycle.spec.ts +++ b/apps/web/e2e/agents-lifecycle.spec.ts @@ -2909,6 +2909,8 @@ test("restores retained Runtime history after a Dashboard reload", async ({ page await expect(dashboard.getByRole("group", { name: "Runtime trend source" })).toHaveCount(0); await expect(dashboard.getByLabel(/Durable · 30s; 1 Runtime targets/)).toBeVisible(); await expect(dashboard.getByLabel("Runtime durable-history charts")).toBeVisible(); + await expect(dashboard.getByRole("heading", { name: "Compute uptime", exact: true })).toHaveCount(0); + await expect(dashboard.locator('[data-chart-engine="uplot"]')).toHaveCount(3); await expect(dashboard.getByText("CPU usage durable trend available")).toBeAttached(); await expect(dashboard).toContainText("120 buckets"); await expect(dashboard).toContainText("119/120 observations"); diff --git a/apps/web/src/features/dashboard/DashboardView.test.tsx b/apps/web/src/features/dashboard/DashboardView.test.tsx index 887b61802..3cb3e67a7 100644 --- a/apps/web/src/features/dashboard/DashboardView.test.tsx +++ b/apps/web/src/features/dashboard/DashboardView.test.tsx @@ -256,7 +256,7 @@ describe("Dashboard loaded-result presentation", () => { expect(html).toContain('aria-pressed="true">1h'); expect(html).toContain("CPU usage"); expect(html).toContain("Memory usage"); - expect(html).toContain("Compute uptime"); + expect(html).not.toContain("Compute uptime"); expect(html).toContain("Token throughput"); expect(html).toContain("No retained CPU samples"); expect(html).toContain("0/2 valid points · 0 snapshots · no history is synthesized"); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx index 09ce5d4c0..c29e9be82 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.test.tsx @@ -58,6 +58,7 @@ describe("Runtime live-window chart accessibility", () => { expect(html.match(/data-chart-engine="uplot"/g)).toHaveLength(4); expect(html).not.toContain("Collecting live samples"); + expect(html).toContain("Compute uptime"); }); it("exposes interactive series, point selection, and Grafana-style in-plot range selection", () => { @@ -81,6 +82,8 @@ describe("Runtime live-window chart accessibility", () => { ); expect(html).toContain('aria-label="CPU usage durable history chart"'); + expect(html.match(/data-chart-engine="uplot"/g)).toHaveLength(3); + expect(html).not.toContain("Compute uptime"); expect(html).toContain('aria-label="CPU usage: 2 retained buckets"'); expect(html).not.toContain('aria-label="CPU usage: 2 live samples"'); }); diff --git a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx index c80ac7577..2edf56ca3 100644 --- a/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx +++ b/apps/web/src/features/dashboard/RuntimeTrendCharts.tsx @@ -556,7 +556,7 @@ export function RuntimeTrendCharts({
`${Math.round(value)}%`} rangeStart={oldest} rangeEnd={newest} source={source} bands={[{ from: 0, to: 30, tone: "safe" }, { from: 30, to: 70, tone: "warning" }, { from: 70, to: 100, tone: "danger" }]} ticks={[1, .7, .3, 0]} emptyMessage={durable ? "No retained CPU samples" : undefined} /> formatDashboardBytes(Math.round(value))} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No complete retained memory samples" : undefined} /> - formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained uptime samples" : undefined} /> + {!durable ? formatDashboardDuration(value)} rangeStart={oldest} rangeEnd={newest} source={source} /> : null} `${formatDashboardTokens(Math.round(value))}/min`} rangeStart={oldest} rangeEnd={newest} source={source} emptyMessage={durable ? "No retained token samples" : undefined} />
); diff --git a/apps/web/src/features/dashboard/runtime-history.test.ts b/apps/web/src/features/dashboard/runtime-history.test.ts index ed3d7140e..58cbfcd3c 100644 --- a/apps/web/src/features/dashboard/runtime-history.test.ts +++ b/apps/web/src/features/dashboard/runtime-history.test.ts @@ -129,12 +129,21 @@ describe("Runtime Durable Dashboard history", () => { memoryLimitBytes: 1_024, inputTokensPerMinute: null, outputTokensPerMinute: null, - targets: [{ label: "Durable worker", cpuRatio: .25, uptimeSeconds: 99.5 }], + targets: [{ label: "Durable worker", cpuRatio: .25, uptimeSeconds: null }], }); - expect(samples[1]?.targets[0]?.uptimeSeconds).toBe(129.5); + expect(samples[1]?.targets[0]?.uptimeSeconds).toBeNull(); expect(samples[1]).toMatchObject({ inputTokensPerMinute: 60, outputTokensPerMinute: 20 }); }); + it("does not derive compute uptime from retained allocation starts or unavailable observations", () => { + const source = history(); + source.series[0]!.points[1] = { + ...source.series[0]!.points[1]!, observed_count: 0, unavailable_count: 1, cpu: null, memory: null, + }; + const samples = runtimeDurableTrendSamples([session], [source]); + expect(samples.flatMap((sample) => sample.targets.map((target) => target.uptimeSeconds))).toEqual([null, null]); + }); + it("keeps aggregate memory absent when any queried target has no memory value", () => { const second = { ...session, id: "44444444-4444-4444-8444-444444444444" } as AgentSession; const secondHistory = history({ diff --git a/apps/web/src/features/dashboard/runtime-history.ts b/apps/web/src/features/dashboard/runtime-history.ts index 201b154a3..9200bf92a 100644 --- a/apps/web/src/features/dashboard/runtime-history.ts +++ b/apps/web/src/features/dashboard/runtime-history.ts @@ -114,18 +114,16 @@ export function runtimeDurableTrendSamples( }); } for (const series of history.series) { - const startedAt = series.started_at.seconds + series.started_at.nanoseconds / 1_000_000_000; const targetID = `${history.session_id}:${series.allocation_id}`; const label = titles.get(history.session_id) ?? "Runtime"; for (const point of series.points) { const value = bucket(point.end * 1_000); const observedAt = point.last_observed_at; - const uptime = observedAt === null || observedAt < startedAt ? null : observedAt - startedAt; value.targets.set(targetID, { seriesId: targetID, label, cpuRatio: point.cpu?.utilization_ratio ?? null, - uptimeSeconds: uptime, + uptimeSeconds: null, }); const usage = point.memory?.usage_bytes; const limit = point.memory?.limit_bytes; diff --git a/apps/web/src/features/dashboard/runtime-trends.test.ts b/apps/web/src/features/dashboard/runtime-trends.test.ts index b90825da9..4554c7364 100644 --- a/apps/web/src/features/dashboard/runtime-trends.test.ts +++ b/apps/web/src/features/dashboard/runtime-trends.test.ts @@ -19,7 +19,7 @@ function snapshot(at: number, options: { cpuUsageSecondsTotal?: number; cpuCapacity?: number; memory?: number; - startedAt?: number; + startedAt?: number | null; allocationId?: string; observedAt?: number; } = {}): RuntimeDashboardSnapshot { @@ -75,7 +75,7 @@ function snapshot(at: number, options: { allocation_created_at: observedAt - 180, resolved_at: observedAt, observed_at: observedAt, - started_at: options.startedAt ?? observedAt - 120, + started_at: options.startedAt === undefined ? observedAt - 120 : options.startedAt, cpu: { utilization_ratio: Object.hasOwn(options, "cpuRatio") ? options.cpuRatio ?? null : .25, usage_cores: Object.hasOwn(options, "cpuUsageCores") ? options.cpuUsageCores ?? null : .5, @@ -166,14 +166,15 @@ describe("Runtime live-window trends", () => { expect(replaced.targets[0]?.seriesId).not.toBe(first.targets[0]?.seriesId); }); - it("keeps one allocation across start changes but resets CPU on allocation changes or counter regressions", () => { + it("resets cumulative CPU on start changes, allocation changes or counter regressions", () => { const base = snapshot(60_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 100, startedAt: 0, }); const continued = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 1, })); - expect(continued.at(-1)?.targets[0]?.cpuRatio ?? null).toBe(.5); + expect(continued.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); + expect(continued[1]?.targets[0]?.seriesId).toBe(continued[0]?.targets[0]?.seriesId); for (const next of [ snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 160, startedAt: 0, allocationId: "44444444-4444-4444-8444-444444444444" }), snapshot(120_000, { cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: 10, startedAt: 0 }), @@ -183,6 +184,24 @@ describe("Runtime live-window trends", () => { } }); + it("starts a fresh CPU baseline after a same-allocation restart even when its counter is higher", () => { + const cumulative = (at: number, usage: number, startedAt: number | null) => snapshot(at, { + cpuRatio: null, cpuUsageCores: null, cpuUsageSecondsTotal: usage, startedAt, + }); + let samples = appendRuntimeTrendSample([], cumulative(60_000, 1, 0)); + samples = appendRuntimeTrendSample(samples, cumulative(90_000, 20, 65)); + samples = appendRuntimeTrendSample(samples, cumulative(120_000, 50, 65)); + expect(samples.map((sample) => sample.targets[0]?.cpuRatio ?? null)).toEqual([null, null, .5]); + expect(new Set(samples.map((sample) => sample.targets[0]?.seriesId)).size).toBe(1); + for (const [priorStart, nextStart] of [[null, 0], [0, null], [null, null]] as const) { + const missingFence = appendRuntimeTrendSample( + appendRuntimeTrendSample([], cumulative(60_000, 1, priorStart)), + cumulative(90_000, 20, nextStart), + ); + expect(missingFence.at(-1)?.targets[0]?.cpuRatio ?? null).toBeNull(); + } + }); + it("keeps directly reported CPU continuous across start changes but not stale observations", () => { const base = snapshot(60_000, { cpuRatio: .25, startedAt: 0 }); const continued = appendRuntimeTrendSample(appendRuntimeTrendSample([], base), snapshot(120_000, { cpuRatio: .5, startedAt: 1 })); diff --git a/apps/web/src/features/dashboard/runtime-trends.ts b/apps/web/src/features/dashboard/runtime-trends.ts index d0b65758b..8f9c9ce85 100644 --- a/apps/web/src/features/dashboard/runtime-trends.ts +++ b/apps/web/src/features/dashboard/runtime-trends.ts @@ -21,6 +21,7 @@ export interface RuntimeTrendTarget { export interface RuntimeTrendCPUCandidate extends RuntimeTrendTarget { observedAt: number | null; + startedAt: number | null; allocationKey: string | null; usageSecondsTotal: number | null; capacityCores: number | null; @@ -115,6 +116,7 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT label: sessionTitle(session), cpuRatio: reportedCpuRatio(observation), observedAt: safeInteger(observation.observed_at), + startedAt: safeInteger(observation.started_at), allocationKey: allocationKey(observation), usageSecondsTotal: finiteNonNegative(observation.cpu?.usage_seconds_total), capacityCores: finiteNonNegative(observation.cpu?.capacity_cores), @@ -156,6 +158,7 @@ export function runtimeTrendSample(snapshot: RuntimeDashboardSnapshot): RuntimeT cpuRatio: target.cpuRatio, uptimeSeconds: target.uptimeSeconds, observedAt: target.observedAt, + startedAt: target.startedAt, allocationKey: target.allocationKey, usageSecondsTotal: target.usageSecondsTotal, capacityCores: target.capacityCores, @@ -193,6 +196,7 @@ function cpuRatios(previous: RuntimeTrendSample, next: RuntimeTrendSample): Map< continue; } if ( + prior.startedAt === null || current.startedAt === null || prior.startedAt !== current.startedAt || prior.usageSecondsTotal === null || current.usageSecondsTotal === null || current.usageSecondsTotal < prior.usageSecondsTotal || current.capacityCores === null || current.capacityCores <= 0 From b0b08d0f3b4370032b1d3826f3c36ed378f26022 Mon Sep 17 00:00:00 2001 From: saladday <1203511142@qq.com> Date: Wed, 23 Sep 2026 19:57:59 +0800 Subject: [PATCH 51/51] Select review depth according to change risk --- AGENTS.md | 6 +++--- CONTRIBUTING.md | 11 +++++++---- 2 files changed, 10 insertions(+), 7 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a5ad47173..f48404ad3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -10,8 +10,8 @@ execution truth or expose server-side credentials in the browser. Update architecture and generated contracts with changes. Run `make check` before reporting completion. -After implementation and verification, request an independent blind review with -only requirements, acceptance criteria, boundaries, repository path and baseline. -Fix in-scope blockers and review again. Never use `codex exec` for this review. +Choose independent blind review according to change risk; follow CONTRIBUTING.md +for reviewer context and re-review criteria. Small, verified fixes may use self-review. +Never use `codex exec` as a substitute reviewer. Documentation and code comments are English; user-facing copy may be bilingual. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 128a2f06f..61c614b16 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -20,10 +20,13 @@ hashes. Do not automatically sync or delete the original repository's Core. Develop in an isolated worktree on a feature branch and submit a PR. Do not edit or commit implementation directly on main. An empty repository bootstrap commit -is only the comparison base for the first import PR. After validation, conduct an -independent blind review using only requirements, acceptance criteria, boundaries, -repository path and comparison baseline. Fix in-scope blockers before delivery. -Do not use `codex exec` as a substitute reviewer. +is only the comparison base for the first import PR. Choose review depth by risk. Substantial changes and changes involving security, +shared lifecycle ownership or uncertain cross-package behavior need an independent +blind review after validation. Give the reviewer only requirements, acceptance +criteria, boundaries, repository path and comparison baseline. Small, verified +fixes may use self-review, including focused corrections after a blind review; +repeat independent review when a correction materially changes the design or risk. +Fix in-scope blockers before delivery. Do not use `codex exec` as a substitute reviewer. The Core Web is an administrator console for execution and resource operations; business collaboration remains in Parsar. Environment Template management shares
{hasLine - ? `${title} ${source} trend available` - : source === "live" - ? `${title} collecting live samples; ${validPoints} of 2 valid points from ${samples.length} snapshots` - : `${title} ${emptyMessage}; ${emptyDetail ?? `${validPoints} valid points from ${samples.length} retained buckets`}`}{runtimeChartCaption({ + title, + source, + hasLine, + allSeriesHidden, + validPoints, + sampleCount: samples.length, + emptyMessage, + emptyDetail, + })}
SeriesLatest valueMissing samples