Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 32 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,6 +259,35 @@ Rules worth knowing before you write one:
code is the risk to design for: `factory.yaml` scores edits to CI or lint
configuration, and test files that lose more lines than they gain, as not
low, so such a "fix" is held for a person.
- **A flaky Actions job can be re-run before a round is spent.**
`publish: {ci: {wait: true, rerun: {budget: 1}}}` re-runs the failed
GitHub Actions jobs on the same head when CI is red on a head jig pushed,
before any repair round (or before the job ends red, with no `on_fail`).
It first waits until nothing on the head is still running, since GitHub
will not re-run a job in a running workflow; then, for each workflow run
the failed jobs belong to, it makes one `gh api -X POST
repos/{owner}/{repo}/actions/runs/{id}/rerun-failed-jobs` request (per run,
not per job: re-running one job would put the run in progress and GitHub
would refuse its siblings), checking the lease and cancellation right
before each. It then judges that same head again, reading each re-run job
as pending until its new run replaces the old red one; if the branch moves
meanwhile, the re-run ends `ci_rerun_head_moved`, so a person's fix is
never reported as jig's flaky pass. The whole re-run — waiting, requests,
and judgement — fits in one CI timeout. GitHub refusing because a
workflow run is in progress is waited out and does not spend the budget;
any other refusal is recorded and falls through to repair or the end. A
re-run GitHub accepted that never finishes, or whose CI cannot be read, is
recorded and leaves CI red as it was, so repair still runs. If
any red check is not an Actions job (a commit status, another app), no
re-run happens. Re-runs never move the branch. The budget, 1–3, is per
attempt: a repair round's new head gets only what is left. Every re-run is
in the result's `ci_reruns` (`attempt`, `head`, `jobs`, `outcome`), and a
pass after one sets `ci_flaky` — it is reported as flaky, never as a clean
pass. Re-run waits run on wall-clock time, so time spent on them is time a
later repair round no longer has under the attempt's ceiling. The
publish-only retry never re-runs. `gh` needs permission to
re-run Actions jobs on the repository. `factory.yaml` and
`factory-parallel.yaml` both declare `rerun: {budget: 1}`.
- **A read-only review panel can run in parallel.** `parallel:
[review-correctness, review-security, review-maintainability]` runs those
phases at once, as one step of the chain. It is opt-in and narrow: one
Expand All @@ -274,8 +303,9 @@ Rules worth knowing before you write one:
and charges only its own budget, and then the whole group runs again, so
earlier approvals are re-judged. The group runs at most 1 + the sum of its
members' budgets times. `examples/definitions/factory-parallel.yaml` is the
stock factory with its panel grouped; `factory.yaml` itself stays
sequential until the parallel panel has been watched on real work.
stock factory with its panel grouped — the same phases, roles, and
`publish` block, CI re-runs and repair included; `factory.yaml` itself
stays sequential until the parallel panel has been watched on real work.
- **Validation happens before anything runs.** `jig def validate <file>`
is the same check the store applies at save time, offline.

Expand Down
10 changes: 10 additions & 0 deletions cmd/jig/ci_repair_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,16 @@ func (g *e2eGateway) FailedCheckLogs(_ context.Context, _ string, checks []worke
return checks
}

// ActionsJobRun and RerunFailedJobs are never reached: the definition
// declares no re-runs, and its checks are not Actions jobs.
func (g *e2eGateway) ActionsJobRun(context.Context, string, int64) (int64, error) {
return 0, fmt.Errorf("unexpected workflow run lookup")
}

func (g *e2eGateway) RerunFailedJobs(context.Context, string, int64) error {
return fmt.Errorf("unexpected re-run request")
}

// scriptCIRepairRuntime scripts the chain's build and the repair round's.
func scriptCIRepairRuntime(t *testing.T) {
t.Helper()
Expand Down
4 changes: 4 additions & 0 deletions cmd/jig/factory_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,10 @@ func TestTheFactoryRepairsRedCIThroughItsWholeReviewPanel(t *testing.T) {
spec.Publish.CI.OnFail.Budget < 1 {
t.Fatalf("publish.ci = %+v, want on_fail repairing through build", spec.Publish.CI)
}
// One re-run tells a flaky job from a broken one before a round is spent.
if spec.Publish.CI.Rerun == nil || spec.Publish.CI.Rerun.Budget != 1 {
t.Fatalf("publish.ci.rerun = %+v, want budget 1", spec.Publish.CI.Rerun)
}
after := map[string]bool{}
seen := false
for _, phase := range spec.Phases {
Expand Down
43 changes: 26 additions & 17 deletions docs/dogfooding.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,11 +21,13 @@ The report is only as good as the ledger underneath it. Run the trial on a
| `jig report` and `GET /api/report` (PR #24) | there is no report. `jig report --help` must print its usage. |
| Spend on every agent phase exit (PR #25) | spend omits every phase that did not pass. It also overcounts every resumed session, because Claude Code's `total_cost_usd` is a running total per session. |
| Publish history across retries (PR #22) | a publish-only retry erases the stop code it replaces, so the *retried, still red* rate reads zero. |
| Flaky-check re-runs (`publish.ci.rerun`, U4 of the plan) | optional. Without it, the *Re-runs* section is all zeros and section 3's masking threshold does not apply. |
| Flaky-check re-runs (`publish.ci.rerun`, PR #28) | optional. Without it, the *Re-runs* section is all zeros and section 3's masking threshold does not apply. |

This runbook describes re-runs as the plan specifies them (R5–R7, KTD5). The
PR that implements them had not landed when this was written, so check the
README's `publish.ci` bullets for the syntax as it shipped.
This runbook describes re-runs as they shipped (R5–R7). One change from the
plan's KTD5: jig re-runs the failed jobs of each workflow run with one
`rerun-failed-jobs` request per run, not one request per job, because
re-running one job puts its run in progress and GitHub then refuses its
siblings. The README's `publish.ci` bullets have the full behaviour.

Build once in your jig clone, run every command below from that clone, and
use the same binary for `serve`, `worker`, and `report`:
Expand Down Expand Up @@ -94,15 +96,22 @@ JOB=$(gh run view "$RUN" --repo "$REPO" --json jobs --jq '.jobs[0].databaseId')
gh api --allow-escape-sequences -H "Accept: application/vnd.github+json" \
"repos/$REPO/actions/jobs/$JOB/logs" | tail -c 2000

# The per-job re-run (KTD5). It really re-runs that job: CI minutes are
# spent and the job's secrets are exercised, which is exactly what jig will do.
# The job-to-run lookup jig makes before a re-run. It must print $RUN.
gh api -H "Accept: application/vnd.github+json" "repos/$REPO/actions/jobs/$JOB" --jq .run_id

# A re-run, which needs the same Actions write permission as the
# rerun-failed-jobs request jig sends per workflow run. It really re-runs
# the job: CI minutes are spent and its secrets are exercised, which is
# exactly what jig will do. (rerun-failed-jobs itself needs a run with a
# failed job, which a green main does not have.)
gh api -X POST "repos/$REPO/actions/jobs/$JOB/rerun"
```

The first call must print log text, not an error. The second must exit 0.
If it fails with HTTP 403, the token cannot re-run jobs: fix it before
declaring any `rerun` policy. Both calls matter because the two CI repair
gate bugs showed up only against real `gh`. The fakes never saw them.
The first call must print log text, not an error. The second must print the
run id. The third must exit 0. If it fails with HTTP 403, the token cannot
re-run Actions jobs: fix it before declaring any `rerun` policy. These calls
matter because the two CI repair gate bugs showed up only against real
`gh`. The fakes never saw them.

### 1.3 CI is green on `main`

Expand Down Expand Up @@ -169,17 +178,17 @@ Copy `factory.yaml` and edit what its header says to edit (the test
command, the builder's `writes` allowlist, and the risk classifier's
paths), then save it:

To re-run failed Actions jobs before a repair round is spent (requires U4),
the copy's `publish` block gains one line, a re-run budget of 1 to 3. A pass
after a re-run publishes, and the report counts it as flaky, never as a
clean first pass:
The stock `factory.yaml` re-runs failed Actions jobs once before a repair
round is spent: its `publish` block declares a re-run budget (1 to 3). A
pass after a re-run publishes, and the report counts it as flaky, never as a
clean first pass. Delete the `rerun` line to trial without re-runs:

```yaml
publish:
hold_when: "risk != low"
ci:
wait: true
rerun: {budget: 1} # the added line
rerun: {budget: 1} # delete to disable re-runs
on_fail: {run: build, budget: 2}
timeout: 30m
```
Expand Down Expand Up @@ -250,8 +259,8 @@ The worst case is roughly:
|---|---|---|
| `T_chain` | plan → build → test → reviewers → risk, until the PR is opened. Measure it on your first jobs (the UI lane's timeline). A round re-runs everything from `build` on, so it costs about as much. | measure |
| `K` | `publish.ci.on_fail.budget`, the repair rounds | 2 |
| `R` | `publish.ci.rerun.budget`, the re-runs per attempt (0 without U4) | 1 once U4 lands |
| `T_ci` | `publish.ci.timeout`. A re-run first waits for every pending check on the head, so a slow sibling job counts against it too. | 30m |
| `R` | `publish.ci.rerun.budget`, the re-runs per attempt (0 without a `rerun` line) | 1 |
| `T_ci` | `publish.ci.timeout`. One re-run (waiting for every pending check on the head, the requests, and the judgement) fits in one `T_ci`, so a slow sibling job counts against it too. | 30m |

With the stock values (K=2, R=1, T_ci=30m), the CI waits alone can take 2
hours, which leaves about 40 minutes for each of the three chain runs. If
Expand Down
5 changes: 5 additions & 0 deletions examples/definitions/factory-parallel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,10 @@
# accepted_unpublished, as red CI did before; the publish retry judges CI
# but never repairs.
#
# THE CI RE-RUN. `rerun: {budget: 1}` re-runs the failed GitHub Actions jobs
# once, on the same head, before a repair round is spent: a job that passes
# on re-run publishes, reported as flaky; one still red goes to repair.
#
# THE ROSTER. Builders and reviewers are `opus`; the planner and the extra
# high-risk reviewer are `sonnet`. `effort` is medium for building to a plan
# and high for reviewers hunting what the build missed. `budget_usd` caps what
Expand Down Expand Up @@ -400,5 +404,6 @@ publish:
hold_when: "risk != low"
ci:
wait: true
rerun: {budget: 1}
on_fail: {run: build, budget: 2}
timeout: 30m
5 changes: 5 additions & 0 deletions examples/definitions/factory.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,10 @@
# accepted_unpublished, as red CI did before; the publish retry judges CI
# but never repairs.
#
# THE CI RE-RUN. `rerun: {budget: 1}` re-runs the failed GitHub Actions jobs
# once, on the same head, before a repair round is spent: a job that passes
# on re-run publishes, reported as flaky; one still red goes to repair.
#
# THE ROSTER. Builders and reviewers are `opus`; the planner and the extra
# high-risk reviewer are `sonnet`. `effort` is medium for building to a plan
# and high for reviewers hunting what the build missed. `budget_usd` caps what
Expand Down Expand Up @@ -387,5 +391,6 @@ publish:
hold_when: "risk != low"
ci:
wait: true
rerun: {budget: 1}
on_fail: {run: build, budget: 2}
timeout: 30m
36 changes: 35 additions & 1 deletion internal/protocol/definition.go
Original file line number Diff line number Diff line change
Expand Up @@ -136,11 +136,24 @@ type PublishSpec struct {
// CISpec opts a definition into waiting for the pull request's CI after
// publish. Timeout is a Go duration ("30m"); empty means DefaultCITimeout.
// OnFail, when declared, repairs red CI inside the attempt instead of ending
// it there.
// it there. Rerun, when declared, re-runs failed GitHub Actions jobs on the
// same head before any repair round, to tell a flaky check from a real one.
type CISpec struct {
Wait bool `yaml:"wait"`
Timeout string `yaml:"timeout"`
OnFail *CIRepairSpec `yaml:"on_fail"`
Rerun *CIRerunSpec `yaml:"rerun"`
}

// CIRerunSpec is the flaky-check re-run policy:
//
// rerun: {budget: N}
//
// Budget counts re-runs per attempt, not per head: a repair round's new head
// gets only what the earlier heads left. A re-run never changes the branch,
// and a pass after one is reported as flaky, never as a clean pass.
type CIRerunSpec struct {
Budget int `yaml:"budget"`
}

// CIRepairSpec is the CI repair loop:
Expand Down Expand Up @@ -168,6 +181,10 @@ const (
// every phase after the repair phase, reviewers included, so a
// larger budget is mostly a larger bill for a fix that is not converging.
MaxCIRepairRounds = 3

// MaxCIReruns caps publish.ci.rerun's budget. Re-running more often than
// this stops telling a flaky check from a broken one and starts hiding it.
MaxCIReruns = 3
)

// WaitsForCI reports whether the definition gates acceptance on CI.
Expand Down Expand Up @@ -456,9 +473,26 @@ func (spec *DefinitionSpec) validatePublish(phases map[string]PhaseSpec) error {
timeout, MinCITimeout, MaxCITimeout)
}
}
if err := spec.validateCIRerun(); err != nil {
return err
}
return spec.validateCIRepair(phases)
}

func (spec *DefinitionSpec) validateCIRerun() error {
ci := spec.Publish.CI
if ci == nil || ci.Rerun == nil {
return nil
}
if !ci.Wait {
return fmt.Errorf("publish: ci: rerun re-runs red CI checks, which needs wait: true")
}
if ci.Rerun.Budget < 1 || ci.Rerun.Budget > MaxCIReruns {
return fmt.Errorf("publish: ci: rerun: budget %d is outside 1..%d", ci.Rerun.Budget, MaxCIReruns)
}
return nil
}

func (spec *DefinitionSpec) validateCIRepair(phases map[string]PhaseSpec) error {
ci := spec.Publish.CI
if ci == nil || ci.OnFail == nil {
Expand Down
48 changes: 48 additions & 0 deletions internal/protocol/definition_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -525,6 +525,36 @@ phases:
"resume_from")
}

func TestPublishCIRerunIsBoundedAndNeedsAWait(t *testing.T) {
const base = `
name: ci-rerun
roster:
builder: {model: opus, system_prompt: s, user_prompt: u}
phases:
- {name: build, kind: agent, owner: builder}
`
if spec := mustParse(t, base+"publish: {ci: {wait: true}}\n"); spec.Publish.CI.Rerun != nil {
t.Fatalf("rerun = %+v, want nil when undeclared", spec.Publish.CI.Rerun)
}
for budget := 1; budget <= MaxCIReruns; budget++ {
spec := mustParse(t, base+fmt.Sprintf("publish: {ci: {wait: true, rerun: {budget: %d}}}\n", budget))
if spec.Publish.CI.Rerun == nil || spec.Publish.CI.Rerun.Budget != budget {
t.Fatalf("rerun = %+v, want budget %d", spec.Publish.CI.Rerun, budget)
}
}
if MaxCIReruns != 3 {
t.Fatalf("MaxCIReruns = %d, want 3", MaxCIReruns)
}
mustReject(t, base+"publish: {ci: {wait: false, rerun: {budget: 1}}}\n",
"publish: ci: rerun", "wait: true")
mustReject(t, base+"publish: {ci: {rerun: {budget: 1}}}\n",
"publish: ci: rerun", "wait: true")
mustReject(t, base+"publish: {ci: {wait: true, rerun: {}}}\n",
"publish: ci: rerun: budget 0", "outside 1..3")
mustReject(t, base+"publish: {ci: {wait: true, rerun: {budget: 4}}}\n",
"publish: ci: rerun: budget 4", "outside 1..3")
}

func TestNegativeRoleBudgetIsRejected(t *testing.T) {
spec := mustParse(t, `
name: budget
Expand Down Expand Up @@ -579,6 +609,24 @@ func TestAParallelGroupOfConsecutiveReadOnlyReviewersValidates(t *testing.T) {
}
}

// A parallel panel and the CI re-run and repair policies are independent:
// declared together they validate, and each keeps its own rules.
func TestAParallelGroupValidatesWithCIRerunsAndRepair(t *testing.T) {
const publish = "publish: {ci: {wait: true, rerun: {budget: 2}, on_fail: {run: build, budget: 1}}}\n"
spec := mustParse(t, parallelPanel("", panelPhases, "parallel: [review-a, review-b, review-c]\n"+publish))
if _, _, ok := spec.ParallelRange(); !ok {
t.Fatal("the group was lost next to a publish block")
}
if spec.Publish.CI.Rerun == nil || spec.Publish.CI.Rerun.Budget != 2 || spec.Publish.CI.OnFail == nil {
t.Fatalf("publish.ci = %+v, want the re-run and repair policies kept", spec.Publish.CI)
}
mustReject(t, parallelPanel("", panelPhases,
"parallel: [review-a, review-b, review-c]\npublish: {ci: {wait: true, rerun: {budget: 4}}}\n"),
"publish: ci: rerun: budget 4")
mustReject(t, parallelPanel("", panelPhases,
"parallel: [build, review-a]\n"+publish), "parallel")
}

// R8: every rule that keeps a group's members concurrent-safe is enforced at
// save time, naming what broke it.
func TestAParallelGroupIsRejectedUnlessItsMembersAreConcurrentSafe(t *testing.T) {
Expand Down
16 changes: 16 additions & 0 deletions internal/worker/claiming.go
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,17 @@ type attemptLease struct {

cancelled chan struct{}
cancelOnce sync.Once

// lost is the control plane's verdict that this lease can never renew
// again, once any heartbeat has received one.
lost error
}

// lostVerdict reports the lease-lost verdict a heartbeat received, or nil.
func (l *attemptLease) lostVerdict() error {
l.mutex.Lock()
defer l.mutex.Unlock()
return l.lost
}

func newAttemptLease(client *Client, attemptID, token string) *attemptLease {
Expand All @@ -60,6 +71,11 @@ func newAttemptLease(client *Client, attemptID, token string) *attemptLease {
func (l *attemptLease) heartbeat(ctx context.Context) error {
response, err := l.client.Heartbeat(ctx, l.attemptID, protocol.HeartbeatRequest{LeaseToken: l.token})
if err != nil {
if leaseLost(err) {
l.mutex.Lock()
l.lost = err
l.mutex.Unlock()
}
return err
}
l.mutex.Lock()
Expand Down
Loading
Loading