Skip to content

Policy: freeze net-new platform surface until eval shows lift over the baseline #88

Description

@leo-aa88

The ratio problem

What has been built (issues #10–#23, the G1–G14 hardening epic #24):

OIDC + hashed API keys + roles + scope isolation + bind guards · token-bucket rate limiting · LLM circuit breakers with token budgets · HMAC-signed outbound webhooks with retry · ingest idempotency keys + content dedup · per-scope retention with a scheduled purge worker · OpenTelemetry tracing · Prometheus /metrics · ingest backpressure · five source adapters (file, CloudWatch, Datadog, Loki, k8s) · versioned v1 JSON schemas · generated Go and Python clients · a web UI · 39 environment variables.

What sits underneath it:

GROUP BY fingerprint · severity_weight + log(count) + log(change_ratio) + log(services)·0.5 · 12 regexes · a 30-minute proximity check.

That ratio is the strategic problem. raglogs has the operational surface of a product people depend on, wrapped around an analysis core that has never been verified to work on a second incident.

Why this is a real cost, not an aesthetic complaint

  1. Every subsystem above is a permanent maintenance surface that competes for attention with explanation quality.
  2. Several already contain configuration that does nothing (SEVERITY_WEIGHT_* settings are ignored by cluster scoring (dead config) #67, and the local-embeddings issue in this epic) — they were built and then not exercised.
  3. Integration tests for all of it have never run in CI (see the CI issue in this epic), so the confidence they provide is partly illusory.
  4. The hardening epic Epic: Hardening & Integration — safe, always-on incident service #24 shipped 14 capabilities in roughly two days of commit history. That pace is only possible when nothing is being validated against reality.

Proposal

Freeze net-new platform surface until raglogs demonstrates measurable lift over the trivial baseline on at least two independent corpora.

Frozen for now:

  • New source adapters (five is plenty to prove the abstraction).
  • New auth modes, new API surface, new client languages.
  • New operational subsystems.

Explicitly still allowed:

Lift the freeze when make eval reports a defensible improvement over the baseline arm on RCAEval and the generated OTel-demo corpus.

Suggested working agreement to add to AGENTS.md

Any PR that changes src/core/ — normalization, clustering, evidence, explain, timeline, retrieval — must include the eval delta in its description: metric before, metric after, on which corpus. "No change" is an acceptable answer; "not measured" is not.

Note

This is a policy issue, not a code issue — close it by agreeing (or explicitly disagreeing, which is also a fine outcome as long as it's a decision rather than a drift). The underlying observation stands regardless of what gets decided: the platform has outrun the engine, and only one of those two is the product.


Part of #74.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationepicTracking issue spanning multiple featuresqualityExplanation quality / core analysis engine

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions