From b25aa671729bf5e6855643a3e070ca7e01ca5bf4 Mon Sep 17 00:00:00 2001 From: Olivier Cots Date: Wed, 2 Sep 2026 15:35:38 +0200 Subject: [PATCH] ci: opt this repo's docs build into the GPU upgrade pass (#885 part 2) Phase F of the plan, the last piece: `CTActions/documentation.yml` (#71, merged) and CTBase's `AmbiguousDescription` fix this page relies on (#941, #945) are both in. This just wires the two new inputs and adds the concurrency group they require. - `gpu_runner: '["occidata"]'` -- opts into `build-gpu`. The existing GitHub-hosted `build` job is untouched and still deploys first, every time; `build-gpu` only runs once `build` has already published (`needs: build` in the reusable) and can never fail the workflow (`continue-on-error: true`) -- on success it redeploys with real GPU numbers, on failure or a long SLURM queue the already-published site is simply left as is. `build-gpu`'s own `if` restricts it to push/tag, so PR-preview docs builds are unaffected regardless. - `gpu_timeout_minutes: 90` -- real margin over the 2265s (~38 min) measured on a same-day build (Phase D probe, 2026-09-02), not the reusable's blanket 120 default, budgeting for the slower case right after occidata-runner-maintenance.yml's Monday cache purge. - `concurrency: { group: documentation-${{ github.ref }}, cancel-in-progress: true }` -- new, and load-bearing now that a GPU pass can run behind a push for up to 90 minutes: without it, two pushes to `main` in quick succession could let an older run's GPU pass finish (and redeploy) after a newer run's CPU build has already published, showing older content over newer. Tags carry distinct refs, so a release build is never cancelled by an unrelated push to `main`. Closes #885. Co-Authored-By: Claude Opus 5 --- .github/workflows/Documentation.yml | 26 +++++++++++++++++++++++++- 1 file changed, 25 insertions(+), 1 deletion(-) diff --git a/.github/workflows/Documentation.yml b/.github/workflows/Documentation.yml index f131174a8..671d79b33 100644 --- a/.github/workflows/Documentation.yml +++ b/.github/workflows/Documentation.yml @@ -8,7 +8,18 @@ on: - 'v[0-9]+\.[0-9]+\.[0-9]+' pull_request: types: [labeled, synchronize, reopened] - + +# A GPU-attempt job now runs behind every push/tag build (see `gpu_runner` below) and +# can take up to `gpu_timeout_minutes`. Without this, two pushes to `main` in quick +# succession could leave an older run's GPU pass finishing — and redeploying — after a +# newer run's CPU build has already published, showing older content over newer. +# `cancel-in-progress` cancels the whole older run, its queued/running GPU job +# included. Tags carry distinct refs (the tag name, not `main`), so a release build is +# never cancelled by an unrelated push to `main`, and vice versa. +concurrency: + group: documentation-${{ github.ref }} + cancel-in-progress: true + jobs: call: # A 'labeled' event fires once per label added, and re-evaluates against the PR's @@ -23,5 +34,18 @@ jobs: uses: control-toolbox/CTActions/.github/workflows/documentation.yml@main with: use_ct_registry: true + # GPU-backed docs build (#885 part 2): publish then upgrade. The GitHub-hosted + # `build` job (unaffected by this) always deploys first; `build-gpu` runs only + # once `build` has already published and can only ever improve the already-live + # site — see CTActions#71 and .reports/campaign/P-gpu-docs-executable.md for the + # full design and the feasibility probe this is based on (Phase D, 2026-09-02: + # real GPU on occidata, a `:gpu` solve in 205s, full docs build 2265s, push via + # GITHUB_TOKEN/HTTPS confirmed). `build-gpu`'s own `if` already restricts it to + # push/tag — inert on PR-preview builds regardless of what's passed here. + gpu_runner: '["occidata"]' + # Measured cold build on occidata: 2265s (~38 min). Budget with real margin — + # this is the number to get right at the top of a Monday, right after the + # occidata-runner-maintenance.yml cache purge, not the warm-cache case. + gpu_timeout_minutes: 90 secrets: SSH_KEY: ${{ secrets.SSH_KEY }}