Skip to content

Timing-sensitive KUnit suites flake on loaded hosts (drm_sched, lib_ratelimit) #52

Description

@tamnd

The 7.2.8 x86_64 defconfig+gk cell with gcc-8.5.0 (cell 34d78cb6) is flaky: boots 2 and 3 reached L8, and boot 1 never finished its KUnit run.

In boot 1, drm_sched_basic_priority_tests passed drm_sched_priorities and then the machine stopped making progress. RCU reported a stall on CPU 1 at 234s, and the NMI backtrace shows CPU 1 spinning in hrtimer_cancel. The call comes from dma_fence_release, through drm_sched_fence_release_scheduled, inside drm_mock_sched_job_signal_timer, which is the timer itself. A timer callback that ends up waiting for its own hrtimer to finish never returns. The same trace was still there at 1119s, when the boot budget ran out.

This looks like a race in the drm_sched mock scheduler and not something gcc-8.5.0 did, since the same build passed the suite twice. TCG on a loaded host makes the timing wide enough to hit it. To check whether it's specific to the GCC, compare how often it shows up in other columns as the stripe fills in. If it shows up in other columns too, drm_sched_basic_priority_tests is a candidate for the known flaky list instead of a per-cell flake.

The log is kunit-1.log in the cell directory.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/bootQEMU rig, gk-init, smoke suites and KUnitkind/findingA measured result worth writing up, such as a new edge or a hole

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions