Skip to content

Redesign trigger detection: rare-event correlation and service linkage, not 12 regexes + a 30-minute window #82

Description

@leo-aa88

Scope narrowed (2026-09-22). The architectural direction for this redesign is established:
docs/design-trigger-detection.md maps #82's asks onto the multi-modal RCA machinery, which is
owned by #118 (now landed). What #82 criticized, however, still exists in production code — the
legacy 12-regex TRIGGER_PATTERNS detector is still authoritative — so this issue stays open,
narrowed to that remaining gap and no longer a broad redesign epic:

Remaining acceptance criterion: replace the legacy 12-regex TRIGGER_PATTERNS detector with the
intended rare-event-correlation + service-linkage mechanism, preserving/finalizing the deferred
30-minute-window behavior and its test. The broader multi-modal architecture is owned by #118 and is
out of scope here.

If the regex replacement is later judged not worth doing, close this explicitly as won't-implement /
superseded-in-direction
, not as completed — the legacy mechanism must actually disappear (or be
consciously abandoned) before this closes.


Problem

raglogs markets "causal timeline reconstruction". What is implemented is temporal ordering plus a regex list.

Trigger detection is 12 hand-written regexes

src/core/normalization/patterns.py:44 — TRIGGER_PATTERNS matches deploy/rollout/release, application/service/pod restart, config reload, migration, queue full, circuit breaker, webhook secret change, token expiry.

Real triggers it cannot see: Helm upgrade succeeded, scaled replicaset to N, feature flag enabled, certificate renewed, failover to replica, ConfigMap updated, node cordoned, autoscaler scaled down, secret rotated, DNS record updated, traffic shifted to canary — plus every trigger phrased in a language or house style the list doesn't anticipate.

And critically: most real triggers write no log line at all. The RCAEval RE2 corpus is 270 cases of CPU hog, memory leak, disk stress, socket stress, network latency, and packet loss — zero of which announce themselves. The current model will find nothing across the entire corpus.

Causality is a 30-minute proximity check

src/core/explain/evidence.py:216:

delta = primary.first_seen - earliest_trigger.timestamp
minutes = int(delta.total_seconds() / 60)
if 0 <= minutes <= 30:
    items.append(f"First error spike occurred {minutes}m after {trigger_label.lower()}")

That is the entire causal claim. No check that the trigger's service relates to the erroring service. No check that the trigger is unusual for this window. No control comparison against a window where nothing broke.

The failure mode is asymmetric and bad

A false trigger match is worth 2 of 8 confidence points and is the sole gate on reaching "high" (src/core/explain/confidence.py:70). So a stray INFO line reading token expired — which matches TRIGGER_PATTERNS — manufactures high confidence for a wrong story. Confidently wrong is worse than "I don't know" for an on-call tool, and the current design has no defence against it.

Direction

Higher-value than adding a 13th regex:

  1. Rare-event correlation. A statistically rare fingerprint in the lookback window is a trigger candidate by definition — this is what baseline_count already measures. A line that appeared 0 times in the last 24h and once, right before the error onset, is a stronger signal than any regex match. This generalizes to trigger phrasings nobody anticipated, in any language.
  2. Require service overlap or dependency linkage. A deploy of service X should not be offered as the trigger for errors confined to unrelated service Y. trace_id is already parsed and indexed and currently unused in the explain path — it is the obvious way to establish real linkage.
  3. Control comparison. src/core/compare/differ.py already implements window diffing. Use it: a trigger candidate that also appears in a comparable healthy window is not a trigger, it is background noise.
  4. Keep the regex list, demote it. It becomes a type classifier (infer_trigger_type already does this well) and a confidence tiebreaker, rather than the detection mechanism.
  5. Separate "trigger found" from "trigger explains this." Report them independently. Being able to say "errors began at 14:09, no correlated change found" honestly is a feature, not a gap.

Acceptance criteria

  • Trigger detection finds candidates on RCAEval RE2 cases, where no trigger line exists — i.e. it is no longer dependent on the change announcing itself.
  • A trigger in an unrelated service is not offered as the cause of errors with no linkage to it.
  • Confidence no longer reaches "high" on the strength of a single unvalidated regex match.
  • Eval lift over the trivial baseline reported before and after, on both RE2 and RE3.
  • Confounded case from the OTel generator (real deploy + injected fault in the same window) selects the correct trigger.

Blocked by

  • Eval harness issue in this epic.
  • RCAEval corpus issue in this epic.

Part of #74.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestqualityExplanation quality / core analysis engine

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions