Skip to content

Flaky mecak8s Prometheus metrics listener test under race suite #2074

Description

@jbeda

Failure

TestTelemetryMetricsAddrServesPrometheus intermittently fails in the full task test:race run:

--- FAIL: TestTelemetryMetricsAddrServesPrometheus (5.05s)
    telemetry_test.go:296: /metrics never responded on 127.0.0.1:36253

The same test passed immediately in isolation with go test -race ./cmd/mecak8s -run '^TestTelemetryMetricsAddrServesPrometheus$' -count=1. This was observed while working on PR #1904; that diff does not touch cmd/mecak8s.

Investigation

cmd/mecak8s/telemetry_test.go selects a free loopback port by binding and closing it before serve binds, with a documented port-handoff race (around lines 125-137). The test starts a fixed 5-second poll immediately after launching the server (around lines 271-296), while the production serve path initializes other listeners before metrics. Those are plausible sources of nondeterminism under aggregate race-suite load; the exact cause of this failure is not established.

Keep the real serve integration assertion that /metrics serves a mecatl_ series. Investigate a deterministic listener/readiness seam (avoiding the free-port handoff) and per-request scrape timeout or context so the overall test deadline is effective. Do not merely raise the sleep/deadline without establishing the cause.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    flakeIntermittent test or CI failure

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions