You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
What has been built (issues #10–#23, the G1–G14 hardening epic #24):
OIDC + hashed API keys + roles + scope isolation + bind guards · token-bucket rate limiting · LLM circuit breakers with token budgets · HMAC-signed outbound webhooks with retry · ingest idempotency keys + content dedup · per-scope retention with a scheduled purge worker · OpenTelemetry tracing · Prometheus /metrics · ingest backpressure · five source adapters (file, CloudWatch, Datadog, Loki, k8s) · versioned v1 JSON schemas · generated Go and Python clients · a web UI · 39 environment variables.
What sits underneath it:
GROUP BY fingerprint · severity_weight + log(count) + log(change_ratio) + log(services)·0.5 · 12 regexes · a 30-minute proximity check.
That ratio is the strategic problem. raglogs has the operational surface of a product people depend on, wrapped around an analysis core that has never been verified to work on a second incident.
Why this is a real cost, not an aesthetic complaint
Every subsystem above is a permanent maintenance surface that competes for attention with explanation quality.
Anything that makes existing features actually work as documented.
Anything that improves explanation quality, with an eval delta attached.
Lift the freeze when make eval reports a defensible improvement over the baseline arm on RCAEval and the generated OTel-demo corpus.
Suggested working agreement to add to AGENTS.md
Any PR that changes src/core/ — normalization, clustering, evidence, explain, timeline, retrieval — must include the eval delta in its description: metric before, metric after, on which corpus. "No change" is an acceptable answer; "not measured" is not.
Note
This is a policy issue, not a code issue — close it by agreeing (or explicitly disagreeing, which is also a fine outcome as long as it's a decision rather than a drift). The underlying observation stands regardless of what gets decided: the platform has outrun the engine, and only one of those two is the product.
The ratio problem
What has been built (issues #10–#23, the G1–G14 hardening epic #24):
OIDC + hashed API keys + roles + scope isolation + bind guards · token-bucket rate limiting · LLM circuit breakers with token budgets · HMAC-signed outbound webhooks with retry · ingest idempotency keys + content dedup · per-scope retention with a scheduled purge worker · OpenTelemetry tracing · Prometheus
/metrics· ingest backpressure · five source adapters (file, CloudWatch, Datadog, Loki, k8s) · versioned v1 JSON schemas · generated Go and Python clients · a web UI · 39 environment variables.What sits underneath it:
GROUP BY fingerprint·severity_weight + log(count) + log(change_ratio) + log(services)·0.5· 12 regexes · a 30-minute proximity check.That ratio is the strategic problem. raglogs has the operational surface of a product people depend on, wrapped around an analysis core that has never been verified to work on a second incident.
Why this is a real cost, not an aesthetic complaint
Proposal
Freeze net-new platform surface until raglogs demonstrates measurable lift over the trivial baseline on at least two independent corpora.
Frozen for now:
Explicitly still allowed:
Lift the freeze when
make evalreports a defensible improvement over the baseline arm on RCAEval and the generated OTel-demo corpus.Suggested working agreement to add to
AGENTS.mdNote
This is a policy issue, not a code issue — close it by agreeing (or explicitly disagreeing, which is also a fine outcome as long as it's a decision rather than a drift). The underlying observation stands regardless of what gets decided: the platform has outrun the engine, and only one of those two is the product.
Part of #74.