Skip to content

Governance config: one-time bootstrap load + two-level lazy TTL (stop fetching GET /v1/governance/{agent} on every run start) #114

Description

@susheem-k

Problem

tokenops_run() calls governance_config_for(service) on every run start / agent hop unless an explicit governor= is passed (src/tokenops/control/run.py:224-226).

  • Embedded mode routes through governance_cache.py (process cache keyed (store_path, agent), invalidated on Store governance writes) — effectively free after the first call.
  • Remote / HTTP mode routes client.governance_config_for -> HttpStore.governance_config_for -> GET /v1/governance/{agent} with no cache. Every run start is a blocking round trip to the plane for config that only changes on Admin edits.

Under the HTTP hard-dependency (see #55, #60), and given TokenOps runs per turn, this adds a plane round trip to every turn purely for near-static config.

Desired behavior

Client-side governance-config cache that:

  1. Loads once on application bootstrap (or lazily on first use) and is held in an in-memory store keyed by (base_url, agent).
  2. Two-level TTL, both evaluated lazily on request arrival (when a run hits the agent) — not on a background timer:
    • Soft TTL: past it, the cached value is still served for this run, and a background refresh is kicked off to rebuild from a fresh GET /v1/governance/{agent}. No request blocks.
    • Hard TTL: past it, the cache entry is considered stale-unsafe — the next request blocks on a synchronous rebuild before building the Governor.
  3. On a successful fetch (foreground or background), the entry is replaced and both TTL clocks reset.
  4. Explicit invalidation still works (Admin edit -> clear_governance_config_cache, or a plane-push/webhook later).
  5. Fail-open on background refresh error (keep serving the last good value until hard TTL); fail-closed on hard-TTL refresh error (surface it — do not run ungoverned, cf. Fail-closed when a run has no governed dispatch, instead of silently running ungoverned #109).

Notes / scope

  • Generalize governance_cache.py so it wraps the HTTP path, not just the embedded Store path.
  • TTLs configurable via env (e.g. TOKENOPS_GOVERNANCE_SOFT_TTL_S, TOKENOPS_GOVERNANCE_HARD_TTL_S) with sane defaults (soft ~60s, hard ~600s — tune).
  • Background refresh: a single-flight worker per (base_url, agent) so concurrent runs don't stampede the plane.
  • Re-key the existing cache from (store_path, agent) to (base_url, agent) in remote mode.
  • Keep the deep-copy-on-read semantics (build_governor mutates its input dict).

Acceptance sketch

  • Steady-state: 0 GET /v1/governance/{agent} calls on the run hot path; refreshes happen in the background after soft TTL.
  • A config change in Admin is picked up within soft TTL (background) or forced at hard TTL.
  • Plane outage shorter than hard TTL is invisible to runs.

Related: #55, #60, #109.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions