Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ OpenAgentCore is protocol-first and modular. Core orchestrates operations that p
| Boundary | Protocol code | Protocol doc |
| --- | --- | --- |
| Application–Core (`/v1`) | Types in `contracts/agents-api/v1/` and route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/openapi.yaml` | [Agents API guide](docs/api/public-agent-api.md) |
| Web and operators–Core (`/core/v1`) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/core.openapi.yaml` | [Core API](docs/api/README.md#core-api) |
| Web and operators–Core (`/core/v1`) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/core.openapi.yaml` | [Core administration API](contracts/agents-api/admin-api.md) |
| Nodes and daemons–Core (`/api/v1` HTTP routes; the node and daemon wire protocols are separate rows) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/runtime.openapi.yaml` | [Machine connection API](docs/api/README.md#machine-connection-api) |
| Core–Sandbox Provider | `services/core/internal/sandbox/sandbox_provider.go` | [Sandbox Provider guide](docs/sandbox-provider.md) |
| Core–sandbox node | `services/core/internal/sandbox/node/wire.go` | [Node generation protocol](contracts/agents-api/node-generation-protocol.md) |
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Core exposes two APIs:
| API | Path | Used by |
| --- | --- | --- |
| **[Agents API](docs/api/public-agent-api.md)** | `/v1` | Your applications. Same protocol as [OpenAI's Agents API](https://developers.openai.com/api/docs/guides/agents-api/overview) |
| **[Core API](docs/api/README.md#core-api)** | `/core/v1` | Operators, through Web |
| **[Core API](contracts/agents-api/admin-api.md)** | `/core/v1` | Operators, through Web |

Core keeps all state. The Runtime runs the chosen harness inside the Environment.
Each connection is a defined protocol, so any part can be replaced on its own. See
Expand Down
2 changes: 1 addition & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ Core 对外提供两组 API:
| API | 路径 | 调用方 |
| --- | --- | --- |
| **[Agents API](docs/api/public-agent-api.md)** | `/v1` | 你的应用,与 [OpenAI 的 Agents API](https://developers.openai.com/api/docs/guides/agents-api/overview) 协议一致 |
| **[Core API](docs/api/README.md#core-api)** | `/core/v1` | 管理员,通过 Web 调用 |
| **[Core API](contracts/agents-api/admin-api.md)** | `/core/v1` | 管理员,通过 Web 调用 |

所有状态都由 Core 保存;Runtime 在 Environment 中运行所选 Harness。各部件之间都通过既定协议连接,
任何一个都可以单独替换。详见[架构说明](docs/architecture.md)。
Expand Down
1,068 changes: 148 additions & 920 deletions contracts/agents-api/README.md

Large diffs are not rendered by default.

416 changes: 219 additions & 197 deletions contracts/agents-api/admin-api.md

Large diffs are not rendered by default.

171 changes: 45 additions & 126 deletions contracts/agents-api/core-errors.md

Large diffs are not rendered by default.

164 changes: 45 additions & 119 deletions contracts/agents-api/core-metrics.md
Original file line number Diff line number Diff line change
@@ -1,21 +1,12 @@
# Core operational metrics

`GET /core/v1/metrics?range=1h|6h|24h|7d` is a Core-key read. It implements the response shape agreed with Core Web PR #96
(`53dc9d646bc6e3cc2cd9b8bbb353bc53c3513ecf`). It changes neither public `/v1`
resources nor Agent or Sandbox metrics. Project API keys cannot call it.
`GET /core/v1/metrics?range=1h|6h|24h|7d` reports Core's own health: its process, execution queue and slots, PostgreSQL and background jobs. It requires the Core key ([Core administration API](admin-api.md)).

Only `range` is accepted, once; omission defaults to `1h`. Empty, repeated,
unsupported or other query parameters return `400 invalid_request`. A missing
metrics service returns `503 core_metrics_unavailable`. Partial measurement
failures return the usual `200 core.metrics` envelope with `service.status` set
to `degraded` and unavailable fields set to JSON null. No database or native
error text, credentials, bodies, resource IDs or tenant labels are exposed.
`range` is the only parameter, sent at most once; it defaults to `1h`. An empty, repeated or unsupported value, or any other parameter, returns 400 `invalid_request`. When Core has no metrics service, or cannot read it, the route returns 503 `core_metrics_unavailable`. When only some measurements fail, the response is still `200` with `service.status` set to `degraded` and each missing value set to null. The response never contains database or native error text, credentials, bodies, resource IDs or tenant labels.

## Time and missing data

The response's `range` uses UTC RFC 3339 boundaries. Its exclusive `end` is the
most recent complete bucket boundary. The included interval is `[start,end)`;
the current partial bucket is excluded from range aggregates and series.
`range` in the response has UTC RFC 3339 `start` and `end` and `resolution_seconds`. `end` is the most recent complete bucket boundary; the interval is `[start, end)`, so the current partial bucket is never included.

| Range | Bucket size | Buckets |
| --- | --- | --- |
Expand All @@ -24,114 +15,49 @@ the current partial bucket is excluded from range aggregates and series.
| 24h | 900 seconds | 96 |
| 7d | 7200 seconds | 84 |

Current gauges and range aggregates are intentionally different: current worker,
connection pool, Go heap and goroutine values are read when requested; process
CPU, RSS and limits, queue gauges, database size are
sampled every 30 seconds with bounded I/O.
Samples older than 60 seconds are not reported as current. Queue, running and
pool and process series report the highest **observed** value in each bucket, not a claim
that all intermediate peaks were captured. Missing observations and the process's
partial first bucket stay null. Successful periodic ping samples produce linear
interpolated p50/p95; there is no request-triggered ping.

A fixed-size in-process ring retains seven days of 30-second samples and
rejection counts, plus two hours of padding for complete bucket alignment. Restart loses those measurements: no synthetic backfill occurs.
The `execution.unavailable` count is null if the requested interval starts before
this process's observation began; an entirely observed interval with no rejections
is zero. PostgreSQL Turn history remains queryable across process restarts.
An empty queue has a measured count of zero but no oldest age. No started Turns
or successful ping samples means null percentiles, not zero latency.

## Sources

The envelope contains `object`, `range`, `service`, `execution`, `database`,
`jobs` and `process`, matching the typed client from PR #96. All numeric values
and `service.execution_owner` are nullable; lists of complete buckets are always
present.

- `service.revision` is a full source commit injected into `main.buildRevision`
by the standalone builder's `-ldflags`. Manual builds without a valid revision
report null. `started_at` records process initialization. `execution_owner`
reflects the execution worker's existing lease checks, with unknown ownership
represented as null. Measurement or job failures report `degraded`; sandbox reset
is reported separately by the deployment contract, not a service status.
- `execution.slots_in_use` is the worker's active Session reservation set. Its
configured capacity is four; environment input, Turns and file work share it.
It does not count native harness subprocesses. Worker-disabled installations
have zero configured execution slots.
- `queued_turns` and `in_progress_turns` count root rows in `turns`, including
operational state retained for deleted Sessions. Native Subagent views and
pending Environment input reservations are not extra queued root Turns.
`waiting_for_daemon` is the queued subset whose Session device binding is
absent from the actual connected-device registry. `oldest_queued_seconds`
measures the oldest queued row's `created_at`.
- `queue_wait_ms` uses `started_at - created_at`, in milliseconds, for Turns
started in the interval. Each bucket uses its own started Turns, with PostgreSQL
`percentile_cont`. `interrupted` counts failed Turns whose outcome error code is
`execution_interrupted`, using `completed_at` in the interval.
- `unavailable` counts actual HTTP errors emitted with code
`execution_unavailable`, once per rejected response. Other 503 codes and errors
occurring after a stream has already started are not counted. The existing error
writer reports the code; no response/request body capture is involved.
- `database.ping_ms` measures a periodic pool ping, including connection acquisition.
Pool `in_use`, `idle` and `max` come from `pgxpool.Stat()`. `size_bytes` is
`pg_database_size(current_database())`, not host disk usage. Failure to measure
one value does not turn it into zero.
- `process.memory_bytes` is Go `runtime.MemStats.Alloc` (allocated heap bytes),
not RSS or container memory. `goroutines` is `runtime.NumGoroutine()`.
- `process.cpu_cores` is the increase in this process's user plus system CPU
time divided by elapsed sampling time. Linux uses `getrusage(RUSAGE_SELF)`;
subprocess and whole-host CPU are excluded. The first interval is null.
Missing/invalid counters, counter resets, nonpositive elapsed time and gaps
longer than 60 seconds reset the baseline; they never manufacture a zero.
- `process.rss_bytes` is Linux `/proc/self/status` `VmRSS`, converted from KiB
to bytes. `cpu_limit_cores` is the process's cgroup v2 `cpu.max` quota/period,
or `GOMAXPROCS` when its quota is `max` or the actual cgroup root has no quota interface. `memory_limit_bytes` is that cgroup's
finite `memory.max`; `max` is null. Membership and mount information resolve
the process's cgroup, including nested and subtree mounts. These are that
cgroup's configured limits, not whole-host metrics or ancestor-limit discovery.
Unreadable or malformed values remain null. Non-Linux builds report null CPU,
RSS and memory limit, with `GOMAXPROCS` as CPU capacity.
- `process.series` always contains the range's complete buckets with `start`,
`cpu_cores` and `rss_bytes`. Each measurement uses its own observed maximum;
missing samples and the partial first bucket stay null. These samples share
the existing bounded ring and restart gaps. Unsupported process measurements
do not by themselves change execution or database health.
- Execution slots, connected daemons, the connection pool, Go heap and goroutines are read when the request arrives. Process CPU, RSS and limits, queue counts and database size come from a sample Core takes every 30 seconds; a sample older than 60 seconds is not reported as current.
- Each series bucket reports the highest value observed in it, not every intermediate peak. Missing observations and the process's partial first bucket are null.
- Samples and rejection counts live in memory for seven days, plus two hours of padding for bucket alignment. A restart loses them; Core does not backfill. Turn history comes from PostgreSQL and survives restarts.
- `execution.unavailable` is null when the interval starts before this process began observing, and zero for a fully observed interval without rejections.
- An empty queue has a count of zero and a null oldest age. No started Turns or no successful ping samples give null percentiles, not zero latency. Percentiles of periodic pings (p50, p95) are linearly interpolated; a request never triggers a ping.

## Fields

The response has `object: "core.metrics"`, `range`, `service`, `execution`, `database`, `jobs` and `process`. Every numeric value and `service.execution_owner` is nullable; each `series` always lists every complete bucket of the range.

| Field | Meaning |
| --- | --- |
| `service.status` | `running`, or `degraded` when a measurement or job fails, the latest sample is missing or stale, or execution ownership is unknown or Core has execution slots but does not hold the execution lease. A sandbox reset is reported by the [deployment](sandbox-deployment.md), not here |
| `service.revision` | The full source commit injected at build time; null for builds without one |
| `service.started_at` | When the process initialized |
| `service.execution_owner` | Whether this process holds the execution worker's database lease |
| `execution.slots_in_use`, `execution.slots_total` | Active Session reservations of the execution worker, and its capacity: [`core.execution_concurrency`](../../docs/configuration.md#settings), 4 by default. Environment input, Turns and file work share the slots; native Harness subprocesses are not counted. Without a worker both are 0 |
| `execution.queued_turns`, `execution.in_progress_turns` | Root Turns in those states, including Turns of deleted Sessions. Subagent Turns and input reserved for a preparing Environment are not counted |
| `execution.waiting_for_daemon` | Queued Turns whose Session's device is not connected; null without a gateway |
| `execution.oldest_queued_seconds` | Age of the oldest queued Turn, from its `created_at` |
| `execution.connected_daemons` | Runtime daemons connected to Core's gateway; null without a gateway |
| `execution.queue_wait_ms` | p50 and p95 of `started_at - created_at` for Turns started in the interval, by PostgreSQL `percentile_cont` |
| `execution.interrupted` | Failed Turns with error code `execution_interrupted`, by `completed_at` in the interval |
| `execution.unavailable` | HTTP responses sent with error code `execution_unavailable`, counted once each. Other 503 codes and errors after a stream started are not counted |
| `execution.series` | Per bucket: `queued`, `in_progress` and `queue_wait_p95_ms` |
| `database.ping_ms` | p50 and p95 of the periodic pool ping, including connection acquisition |
| `database.pool` | `in_use`, `idle` and `max` connections of the pool |
| `database.size_bytes` | `pg_database_size(current_database())`, not host disk usage |
| `database.series` | Per bucket: `ping_p95_ms` and `pool_in_use` |
| `process.memory_bytes`, `process.goroutines` | Go `runtime.MemStats.Alloc` (allocated heap, not RSS) and `runtime.NumGoroutine()` |
| `process.cpu_cores` | Increase in the process's user plus system CPU time divided by the elapsed sampling time (Linux `getrusage(RUSAGE_SELF)`), excluding subprocesses. Null for the first interval; a missing or reset counter, a nonpositive interval or a gap over 60 seconds restarts the baseline |
| `process.rss_bytes` | Linux `/proc/self/status` `VmRSS`, in bytes |
| `process.cpu_limit_cores` | The process's cgroup v2 `cpu.max` quota divided by its period, or `GOMAXPROCS` when the quota is `max` or the cgroup has no quota interface |
| `process.memory_limit_bytes` | The process's cgroup `memory.max`; null when it is `max` |
| `process.series` | Per bucket: `cpu_cores` and `rss_bytes` |

Core resolves its own cgroup, including nested and subtree mounts, and reports that cgroup's limits, not the host's or an ancestor's. Unreadable or malformed values are null. Non-Linux builds report null CPU, RSS and memory limit, with `GOMAXPROCS` as the CPU limit. Missing process measurements alone do not make the service `degraded`.

## Background jobs

The four bounded IDs are `scheduler`, `runtime_sampler`, `history_cleanup` and
`audit_cleanup`. Each reports `status`, `last_run_at`, `processed` and `failed`.
A not-yet-observed run is unknown; a disabled or stopped loop is stopped. The
last-run time is the completion/observation of the last pass, not its next deadline.

Scheduler processed counts selected Turn/environment work in that poll. Runtime
sampling counts observed and failed targets from its existing sweep result.
Cleanup processed counts confirmed removed rows; an unsuccessful cleanup reports
unknown processed count. Cleanup/scheduler failure counts identify failed passes,
not guessed numbers of lost rows or failed Turns. The actual scheduling, sampling,
retention and execution lifecycles keep their existing owners and timing.

No additional telemetry database, monitoring server, model-provider probe,
message queue, object store, host disk measurement or scheduling mechanism is
introduced. Metrics cannot authorize execution or change resource ownership.

## Implementation rules

The Core-key-only `/core/v1/metrics` contract is documented in
[core-metrics.md](core-metrics.md). Keep this separate from
Agent outcome and Sandbox capacity views. Instrument existing worker and job
owners without changing scheduling, lease or retention behavior. Periodic pool
pings and bounded in-process samples have explicit restart gaps; unknown values
must remain null. Complete UTC buckets exclude the active partial bucket. Root
Turn history is queried read-only from PostgreSQL with native timestamps.
Count `execution_unavailable` at the existing HTTP error writer, once per rejected
response; never record request/response bodies or infer this count from every
503 or failed Turn. Builds inject the source commit with ldflags. No new monitoring
service or storage system is required. Keep the frontend response shape aligned
with the paired console contract.
`jobs` lists `scheduler`, `runtime_sampler`, `history_cleanup` and `audit_cleanup`. Each has `status` (`unknown` before its first run, `ok`, `failing`, or `stopped` when its loop is disabled or ended), `last_run_at` (when the last pass finished), `processed` and `failed`, both describing the last pass.

Core process CPU, RSS and cgroup limits are sampled by the existing 30-second
Core metrics loop; Go heap and goroutine reads retain their meaning. Process
series use the same bounded ring and complete buckets, with null first CPU
intervals and restart gaps. Do not substitute host usage for process usage.
- The scheduler's `processed` counts the Turn and Environment work it selected in its last poll.
- The Runtime sampler's `processed` and `failed` count the targets its last sweep observed and failed to observe.
- A cleanup job's `processed` counts the rows it removed.
- For the scheduler and the cleanup jobs, a failed pass sets `processed` to null and `failed` to 1; `failed` never estimates lost rows or failed Turns.
Loading
Loading