Skip to content

Release 5.0.0 - #34

Closed
serialexperimentslainnnn wants to merge 38 commits into
mainfrom
develop
Closed

serialexperimentslainnnn wants to merge 38 commits into
mainfrom
develop

Conversation

@serialexperimentslainnnn

Copy link
Copy Markdown
Owner

Pull request

Summary

What does this PR change and why? One short paragraph is fine.

Related issue

Closes #

Type of change

  • Bug fix
  • New feature
  • Refactor (no behavioural change)
  • Docs / build / CI
  • Security fix

Risk and rollback

Risk: what breaks if this is wrong, and for whom? (none is a valid
answer for docs-only changes — say so rather than leaving it blank.)

Rollback: how is this undone once released? Reverting the commit is not a
rollback for a published plugin — a user on the bad version stays there until
they update. If the change touches persisted settings, the transcript format,
or the permission surface, say what happens to a user who already ran it.

Checklist

  • PR targets the develop branch (or main only for hotfixes).
  • Commits follow Conventional Commits (the commit-msg hook enforces it —
    install once with git config core.hooksPath .githooks).
  • ./gradlew test verifyPlugin buildPlugin passes locally.
  • verifyPlugin is Compatible across the declared range (251 → 263.*)
    and reports no new internal-API usage (@ApiStatus.Internal).
    The CDN download is unreliable here; use
    -PlocalIdePath=<dir>[,<dir>…] with locally-extracted IDEs.
  • No new deprecated or scheduled-for-removal IntelliJ Platform APIs.
  • Tests added or updated for the new behaviour — src/test/kotlin/… for
    Kotlin, src/test/frontend/… (npm test) for anything under
    src/main/resources/jcef/.
  • Protocol changes: ./gradlew checkDrift is green and the baseline in
    scripts/drift-baseline.properties matches what was verified.
  • New dependency? Its licence is compatible with GPL-3.0-only and it is
    recorded in THIRD-PARTY-NOTICES.md if it
    ships in the artifact.
  • User-visible changes are documented in CHANGELOG.md
    and RELEASE_NOTES.md under Unreleased.
  • No secrets, tokens, conversation transcripts, or personal absolute
    paths in the diff or commit messages.
  • Follows the conventions in CONTRIBUTING.md, the
    architectural contract in CLAUDE.md, and the recorded
    decisions in docs/adr/.

How was this tested?

  • Unit tests (./gradlew test) and frontend tests (npm test)
  • Manual sandbox (./gradlew runIde) — describe the scenarios you
    exercised.
  • Smoke test on a real IDE install — describe.
  • UI changes only: driven with the keyboard alone, with the focus ring
    visible on every control touched. Automated checks catch roughly half of
    real accessibility barriers and none of the judgement calls, so this one
    is not delegable to a tool.

Notes for reviewers

Anything tricky, follow-up work, or open questions.

serialexperimentslainnnn and others added 30 commits August 5, 2026 16:43
Audited the repository against git-workflow-standards. Adds the governance
the standard requires and that this repo lacked, minus anything CI-bound
(no CI gate exists yet - a local lab is planned).

- commitlint.config.mjs + a VERSIONED .githooks/commit-msg gate, so the
  rule travels with the repo instead of living in one laptop. The hook
  self-tests against a known-good message first and degrades to a warning
  if the toolchain is broken: a hook that fails closed on its own bugs
  gets bypassed with --no-verify, and then protects nothing.
- .gitattributes: this plugin ships and is tested on Windows, and
  publishes signed zips. Without a declared policy a clone with
  autocrlf=true rewrites every file and invalidates artifact signatures.
  Vendored frontend bundles marked -diff so a bump is reviewable.
- docs/adr/0001: records the two justified deviations (GitFlow for
  installable versioned software; GPG-on-YubiKey for real revocation and
  because the same key signs the distributed artifacts) and corrects a
  real violation - v4.3.2 and v4.4.1 were each re-cut and force-pushed
  three times. Published tags are now immutable; a post-tag mistake is
  fixed by the next patch version.

CHANGELOG generation is deliberately still open: the tooling silently
skips non-Conventional commits, so the message gate has to land first.

Refs: git-workflow-standards §3.2, §4.1, §5.2, §5.3
The published zip redistributes third-party code - marked 12.0.0 (MIT),
DOMPurify 3.0.11, highlight.js 11.9.0 (BSD-3-Clause) vendored into the
plugin jar, and kotlinx.serialization 1.7.3 (Apache-2.0) as its own jars.
MIT, BSD-3-Clause and Apache-2.0 all require the copyright notice and
licence text to be preserved ON REDISTRIBUTION, and a file sitting in the
Git repository does not accompany the binary a user installs from the
Marketplace. This was unmet.

processResources now packages THIRD-PARTY-NOTICES.md, LICENSE and the
three licence texts into META-INF/ of the plugin jar - verified present in
the built artifact, not just in the source tree. Generated at build time
from one root-level source of truth rather than a checked-in copy, so the
notices cannot drift from the files they describe.

Every licence was verified by reading the LICENSE of the exact shipped
version, never the manifest or a badge. That found a real one: DOMPurify
3.0.11 is dual "Apache-2.0 OR MPL-2.0". An OR expression is a choice the
redistributor must make and record; leaving it unstated is an unmade
decision. Apache-2.0 is selected, with the reasoning recorded in the
notices file. kotlinx.serialization ships no META-INF/LICENSE in its jars,
so its licence was read from the project source instead of inferred.

Also corrects the documented JDK path across CLAUDE.md, README,
CONTRIBUTING and docs/: ~/.local/jdks/jdk-21.0.11+10 no longer exists on
this machine (it is now ~/.jdks/jbr-21.0.11), and gradlew failed with
"JAVA_HOME is set to an invalid directory" rather than anything a build
log filter would obviously catch.

Refs: opensource-licensing-standards §3.5, §3.6, §5.4, §8
Audited the JCEF web UI against WCAG 2.2 AA. This is a conformance
requirement, not a nice-to-have: the plugin is distributed to consumers in
the EU, where Directive (EU) 2019/882 has applied since 28 June 2025 (in
Spain, Ley 11/2023).

4.1.3 Status Messages (AA) - the transcript streams without ever moving
focus, so a screen-reader user got NO signal that Claude started, finished,
or was blocked on them: the turn simply stopped, silently. Adds a polite
live region declared in the static shell (it has to exist in the DOM before
text is written into it, or the first change is never announced) plus
CC.announce. Only turn EDGES are announced, never per token - a region
updated on every delta talks over itself and gets switched off, which is
worse than silence. The most important case is a pending permission card:
it appears without taking focus, so it was previously undetectable.

2.4.7 Focus Visible / 1.4.11 Non-text Contrast (AA) - the stylesheet
suppressed the default outline in five places. Four had some replacement;
.find-input had NONE, so focus there was invisible. Adds a :focus-visible
baseline covering the controls built as <span role="button"> (the code-block
Copy control, menu items, tool cards), which get nothing for free because
they are not native buttons, and honours forced-colors instead of overriding
the system's focus colour.

prefers-reduced-motion was already correct (a universal reset covering all
eight keyframe animations) and is now pinned by a test so it cannot decay
into a hand-picked subset as animations are added.

Refactor: the frontend test harness hand-copied the shell DOM and had
already drifted - shell.html gained the live region, the harness did not, so
tests exercised a DOM the product does not have. It now extracts the body
from the real shell.html. A mount point added to the shell reaches tests
automatically; one removed breaks the tests that relied on it.

Scope stated honestly in the test file: automated checks catch ~40-57% of
real barriers and none of the semantic judgements. These pin the structural
guarantees that regress silently. They do not certify conformance - that
still needs keyboard-only and screen-reader passes by a person.

Refs: accessibility-standards §3.2, §3.4, §3.8, §4.1
`@anthropic-ai/claude-agent-sdk` sat in `dependencies` while CLAUDE.md has
always said it is "protocol reference only, not distributed" — and it is: the
published artifact is jars and inlined web assets, with zero `node_modules`
entries (verifiable with `unzip -l`). The mismatch produced seven permanent
npm-audit findings (three high) against code no user ever receives, which is
worse than useless: a backlog of alarms that cannot be acted on trains you to
ignore the one that matters.

Moving it to `devDependencies` makes the declaration match reality, so
`npm audit --omit=dev` — the distributed scope — now reports zero. The full
tree still reports the same seven; they are build-time tooling and are handled
as maintenance. SECURITY.md states that triage boundary explicitly, with the
command to verify it rather than trust it.

`checkDrift` was the constraint on this change: it reads the SDK out of
`node_modules` and runs `npm update`. Both survive, because `npm install`
installs devDependencies — only `--omit=dev` would break it, and nothing uses
it. Verified green, and it reported real drift while doing so, so the baseline
moves with it: SDK 0.3.220 → 0.3.222, binary 2.1.220 → 2.1.222, protocol
surface unchanged.

Also corrects the package manifest itself: it declared `"license": "ISC"` on a
GPL-3.0-only repository and lacked `"private": true`, i.e. it was publishable
to npm with the wrong licence. And the supported-versions table still said 2.x.
The control was strong and undocumented: nothing stated what it is defending
against, which makes it unreviewable. Coverage cannot be argued, and every
discussion of a proposed bypass restarts from first principles.

ADR 0002 states the trust model — user trusted, `claude` binary trusted as
software but untrusted as a channel, everything it relays untrusted — and runs
STRIDE over the three surfaces that actually exist: the binary as a child
process, third-party MCP servers, and model-returned content.

The third is the one worth having written down. We do not attempt to detect
prompt injection; content-level detection is unsolved and a control built on it
would be a liability. The document says so explicitly and records where the
defence actually sits instead: the guard judges the tool call and never the
reasoning behind it, so an injection that succeeds completely still has to ask
to read the key and still gets the same answer. Also records why MCP servers
are denied rather than asked about, and why caller trust is an allowlist —
names are attacker-supplied, and 4.4.0 is the cautionary tale of that list
going stale (it failed safe, which is why the allowlist was the right shape,
but it still failed).

Non-goals are as load-bearing as the goals, so they are listed: we do not
defend against the user, against a compromised host, or against every
obfuscation. That last one draws the line triage keeps blurring — failing to
*recognise* a path is a pattern gap; enforcement of a match, once made, is
absolute.

Claims verified against the code before writing, not from memory:
`binaryPermissionMode` rewriting acceptEdits/bypassPermissions to `default`,
and the CSP's `default-src 'none'` / `connect-src 'none'` / hash-pinned scripts.
SECURITY.md links here so a reporter can tell a finding from a stated position
before writing the email.
Until now the entire quality bar for ~13k lines of Kotlin and ~3.6k lines of
shipped JavaScript rested on review. That is the thing the standards say to
mechanise: if formatting is discussed in a review, a formatter is missing.

Adds and configures:

- detekt 1.23.8 (config/detekt/detekt.yml). Every non-default setting carries
  its reasoning AT the setting, not in a message nobody will find again. Three
  rules were re-calibrated rather than suppressed: TooManyFunctions counts only
  the public surface (the default penalised extracting helpers, which is the
  exact fix the complexity rules demand); LongParameterList ignores defaulted
  parameters (one absent from the call site is not part of the problem the rule
  exists for); MagicNumber skips the colour palette (a 0xRRGGBB inside a
  property named DIFF_ADDED_BG is that name's definition).
- Spotless 8.9.0 / ktlint 1.8.0. max-line-length and function-naming are
  disabled here and left to detekt: both tools enforced them, only detekt can
  scope an exception to the test tree, and running both meant the
  stricter-but-blinder one decided.
- ESLint + Prettier over the JCEF frontend, which ships inside the plugin jar
  and had never passed through any tool. no-eval / no-implied-eval / no-new-func
  are errors because the page runs under a hash-pinned CSP with no
  'unsafe-eval': without this gate, code Chromium will silently refuse in a
  user's IDE can still reach main. Vendored marked/DOMPurify/highlight.js are
  excluded — a finding there is not ours to fix.
- koverVerify, gated per package. Risk is not evenly spread: permission/ decides
  whether the agent may read an SSH key, ui/ paints a browser. Thresholds sit
  slightly BELOW what each package measures today, so they catch regression
  rather than invite test-padding. ui/, context/, process/, actions/ and util/
  are excluded with the reason stated rather than gated at a token value.

config/detekt/baseline.xml holds TWO entries, both about ClaudeSession, both
explained inside the file. It started this work at 203; the other 201 were
fixed, not frozen.

A `Static analysis` job runs all of it and is a required check on main and
develop, so none of the above depends on anyone remembering to run it.
Three strands that could not be separated into their own commits: the formatter
rewrote the same files the refactor touched, hunk for hunk, so splitting would
have meant manual surgery across ~100 files for commits that would not build in
isolation. Recorded here rather than pretended otherwise — it also means there
is no pure formatting commit to add to .git-blame-ignore-revs.

REFACTOR — the dispatch tables, split in two levels

ClaudeSession.onEvent was one `when` over 47 event types: 244 lines, cyclomatic
complexity 111, the single function where every protocol concern in the plugin
met. ClaudeEvent now declares seven sealed sub-interfaces and dispatch picks the
group, then the variant. The grouping lives in the TYPE on purpose: a sealed
hierarchy keeps the compiler checking exhaustiveness at BOTH levels, so a new
protocol event nobody handles is a compile error rather than a silently dropped
frame — the property checkDrift exists to protect, and not up for trade against
a complexity threshold. The groups are semantic: they differ in what the host
OWES the binary (a Control frame must be answered or the binary hangs; a Notice
is fire-and-forget). JcefBridge.Msg and JcefChatPanel.onBridgeMessage got the
same treatment. Several `when` chains were dictionaries written as control flow
and are now data (ProtocolParser.parseSystem: 25 arms, 21 identical bar two
names).

FIXED — defects the tooling and live probing surfaced

- Locale-dependent formatting. TokenFormat rendered "1,2k" on a comma-decimal
  machine, and worse, the trailing-.0 test then stopped matching so a flat 1000
  showed as "1,0k". The same bug in JcefTheme.rgba emitted four rgba components
  instead of three, so the browser DISCARDED the declaration and the accent/link
  washes never rendered. Now pinned to Locale.ROOT.
- Diff tabs were persisted into the workspace and could never be restored: they
  are mock:// in-memory previews, and the platform persists open tabs by URL
  without filtering by file system. One workspace had 13 such entries, all named
  "Claude · SKILL.md", producing 26 warnings on every launch. DiffTabCleanup now
  drops them on projectClosingBeforeSave — the one hook that runs BEFORE the
  state is written.
- Every chat animation was dead. The accessibility pass had added a
  prefers-reduced-motion block, and inside JCEF that query reports `reduce`
  regardless of the desktop (measured: it matched while GNOME had animations
  enabled), so it flattened everything. The tell was Vibe Mode's rainbow, which
  survived because it is a setInterval, not a keyframe. Motion is now reduced
  only when the user asks, via a plugin setting, off by default.
- Requests needing a live binary were dropped. The panel is built BEFORE
  session.start() runs, so requestMcp/requestVersion/requestUsage all found
  isRunning() == false and did nothing; each looked like its own bug. They now
  queue until ready.
- sniffMediaType treated any RIFF container as WEBP; it now checks the form type
  that actually identifies the format (WAV and AVI share the RIFF header).

FEAT — plan limits, in both surfaces

get_usage is a control request the plugin has known about since 4.0.1 and never
sent; it returns every rate-limit window at once plus the extra-credit balance.
The dashboard shows one bar per window with its reset countdown; the composer
readout shows a dot per window. One severity scale for both: blue under 65%,
amber under 85%, red above — status hues, never the product accent, because a
quota bar that is coral at 10% and coral at 95% has said nothing. Crossing 65%
and 85% announces once per threshold per window (a warning that repeats every
refresh is one users learn to ignore), and 85% also raises an IDE notification.

ClaudeSession.rateLimits is now per-window: the single field could only ever
hold whichever event arrived last, so a user on a weekly limit saw their
five-hour bar and concluded they had room.

Also adds CC.diagnostics, which reports what the embedded browser actually
resolves to idea.log. The UI is a browser nobody can open devtools on, and
without it a CSS rule that silently does not apply is indistinguishable from a
backend that never sent the state — which cost three wrong diagnoses before it
existed.

Suite: 682 -> 689 JVM, 54 -> 67 frontend.
Documents what the release actually did, including the parts that did not go to
plan — a release note that only lists wins is one nobody can act on.

- CHANGELOG gains two sections: the static-analysis work (203 findings fixed
  rather than frozen, down to 2 recorded exceptions) and the defects it exposed.
- docs/RELEASE_CHECKLIST.md gains a Coverage policy with the MEASURED
  per-package numbers, what is excluded and why, and the Kover 0.9.2 limitation
  behind the current floor+aggregate shape. It also corrects a claim that never
  had a basis: build.gradle.kts cited a ">=90% target documented in
  RELEASE_CHECKLIST.md", that document had never mentioned coverage, and the
  real figure was 53%.
- docs/BACKLOG.md is new. Every entry was PROBED against the real binary rather
  than inferred from the SDK types, including the two that are deliberately not
  worth doing, so nobody investigates them a third time.
- AGENTS.md, CI_SETUP.md, BRANCHING.md and the ADRs updated for the GitHub
  Actions pipeline that replaced the deleted GitLab one.

Known gaps, stated rather than left to be discovered: the risk-ranked human
review is not done; the reduced-motion decision generalises from a measurement
on one machine; and Fable's weekly window cannot be shown at all, because
get_usage returns every per-model window as null (measured, not assumed).
Squash and rebase merging are now disabled at the repository level, leaving
merge commit as the only option. BRANCHING.md said the opposite for `develop`
("squash or merge"), which was worse than saying nothing: it invited the exact
click that destroys signatures.

Both disabled methods REWRITE commits — new SHAs, new committer information —
which invalidates the author's hardware-backed signatures and replaces them with
GitHub's web-flow key. The "signed commits required" rule would still pass, and
that is the trap: the commits are signed, just no longer by the author, which is
the whole property the rule exists to provide.

The cost is recorded too: the merge node itself is GitHub-signed, because the
alternative is a local merge and a direct push, which the pull-request
requirement blocks. Relaxing that to save one commit's provenance would be a far
worse trade. Release provenance therefore rests on the signed tag, not on the
merge node.

Also corrects the status-check lists, which had not been updated when the
`Static analysis` gate was added.
CodeQL flagged two high-severity alerts on the same line of the frontend test
harness: "incomplete multi-character sanitization" and "bad HTML filtering
regexp". Both were correct: the pattern used to strip script tags does not match
a closing tag written with a trailing space, so something that reads like a
sanitiser silently is not one.

The exposure was low. This reads OUR OWN shell.html off disk inside a test, and
the reason script tags are stripped is load control — the harness injects the
app modules itself, in a controlled order — rather than sanitisation. That is a
reason to rank the finding, not a reason to keep it: a regex that looks like
HTML sanitisation is a pattern someone copies to a place where it does matter.

Fixed at the root rather than by widening the pattern. jsdom is already a
dependency of this harness, so the file is parsed and the script elements are
removed as nodes. All 67 frontend tests still pass, which is the check that the
harness still builds the same DOM as the product.

Recorded for the risk review: the two CodeQL contexts the rulesets require
(java-kotlin, javascript-typescript) both PASSED — they report "the analysis
ran". The aggregate CodeQL check, which reports "something was found", is
required by neither branch. A high-severity finding can therefore merge with no
gate objecting.
A regression introduced by this branch made the plugin unusable, and fixing it
surfaced a run of smaller defects that all shared one shape: something failed or
was still loading, and the UI said nothing at all.

THE REGRESSION

JcefChatPanel.pendingUntilReady was declared below the `init` block that uses it.
Kotlin runs property initializers and init blocks in declaration order, so the
list was still null while init ran and the constructor threw. It took the whole
tab with it: no chat could be opened or restored. lastUsage/lastUsageAt had the
same defect and stayed silent, being nullable and primitive. InitOrderContractTest
now fails the build on any property declared after its class's init.

WAITING IS NOW VISIBLE

The binary is launched BEFORE the tab is built (start() only dispatches, so
`claude` boots while JCEF creates its browser), and a boot screen holds the tab
until the process is up. Three states, not two: running, starting, and neither —
that last one is a failed launch and must clear the screen, or a missing binary
leaves the tab covered forever.

Context and cost no longer wait on a clock. The poll timer's initial delay equals
its interval, so the first reading was a minute late; worse, that tick landed
while the binary was still starting and returned early, costing a second minute.
Ready, tab-open and both turn edges now poll directly, and the timer retires at
the end of a turn instead of round-tripping forever on an idle session.

Plan limits were pushed to the dashboard but not to the composer, so the same
number appeared instantly in one place and "later" in the other. Reasoning tokens,
context and output now settle at 0 rather than being omitted until non-zero: an
absent figure is indistinguishable from one that failed to load.

FAILURES THAT WERE HARD TO READ

- The CLI wraps failed tool results in <tool_use_error>...</tool_use_error>
  (verified in claude 2.1.222, which carries the same text unwrapped in a sibling
  field). Rendered verbatim it put raw markup in a native GUI. Stripped only when
  it encloses the whole payload; is_error already conveys the failure.
- A failed tool card stayed collapsed, so the entire message was "the header is
  red", and its output scrolled sideways rather than wrapping, hiding the half
  that says what to do instead. It now opens itself once, and error text wraps.
- ToolSearch was missing from SensitiveGuard.AGENT_TOOLS, so on a session that
  defers tools, the call that loads every other tool's schema landed in the
  untrusted branch. Added, with AskUserQuestion, Mcp and FileRead/Edit/Write.
  Entries are only ever added here: this is a trust allowlist, not an inventory,
  and a missing first-party name is the 4.4.0 hard-DENY incident.
- A markdown link whose href is a path did nothing when clicked: the host handled
  https:// and jb://open and dropped the rest without a sound. Both paths now go
  through one gate (LinkResolver.isOpenable). The scheme test requires two or more
  characters so a Windows drive letter stays a path.
- Copy on a message copied and gave no feedback, which reads as a broken button;
  the shared flash helper is now exported rather than reimplemented, and the
  .copied class finally has a CSS rule — it had never had one.

Suite: 692 -> 694 JVM, 67 -> 84 frontend. verifyPlugin Compatible on IC-251,
IC-252, IU-253, IU-261 and IU-262.
Two CI failures, one of them ours.

OURS — the `Static analysis` job died with an IOException from DriftLiveCheck.
Kover instruments and aggregates EVERY `Test` task in the project, and
`checkDrift` is registered as one, so it had silently become a dependency of
`koverVerify`. That task is on-demand by design: it downloads the latest SDK and
probes a `claude` binary installed on the machine. A runner has no such binary.

The task's own documentation already said "NOT wired into `check`" — it just was
not true of the coverage graph, and nothing checked that it was. It passed
locally for the one reason that makes this class of bug expensive: the
maintainer's machine has the binary, so "on-demand" and "wired in" looked
identical until CI ran it. Verified with `koverVerify --dry-run`, which listed
`:checkDrift` in the graph before and does not after.

NOT OURS — the `JVM tests` job failed resolving
com.jetbrains.intellij.platform:test-framework with a 502 Bad Gateway from
cache-redirector.jetbrains.com. Nothing to fix; it needs a re-run.

Also bumps the IntelliJ Platform Gradle Plugin 2.16.0 -> 2.18.1, which the build
had been warning about on every run.

Stated plainly: the plugin bump is NOT verified locally. A 2.16 -> 2.18 jump
touches the whole build, and CI is the verification here rather than a claim
made in advance.
Reverts the 2.16.0 -> 2.18.1 half of the previous commit. The `checkDrift`
coverage-graph fix in that commit stands; only the plugin version goes back.

On 2.18.1 the headless suite never runs. `ChatSessionManagerHeadlessTest` hangs
before its first assertion, and a thread dump puts the EDT here:

  BasePlatformTestCase.setUp
    -> LightPlatformTestCase.doSetup
      -> IndexingTestUtil.waitUntilIndexesAreReady   (267s and counting)

That is the platform's own fixture waiting for indexing that never completes.
Our code is not on the stack at all — the bump changes which platform
test-framework is resolved, so this is a fixture-level regression rather than
something fixable from this side.

Measured both ways rather than inferred: the same test on 2.16.0 finishes in 19
seconds, and the full gate is green at 694 tests.

The version is now pinned with the reason written AT the setting, because the
build prints an "outdated" warning on every single run and the next person to
see it will otherwise do exactly what I did. Re-attempt it as its own change
with the headless suite as the acceptance test, not as a drive-by on a release
branch.

I shipped that bump unverified and said so in the message; this is what it cost.
The one-line lesson is the boring one: a build-tooling bump is a change like any
other and does not get to skip the suite because it looks like configuration.
Merging `develop` into `main` now publishes. The tag is DERIVED from the version
in build.gradle.kts rather than supplied alongside it, so the two can no longer
disagree — the mismatch the old flow guarded against with a comparison simply
cannot occur.

A merge that does not bump the version publishes nothing: the workflow finds the
tag present, logs a notice and stops. It does not fail. `main` legitimately
receives merges that are not releases, and a red run on each of those is an alarm
people learn to ignore. Published tags stay immutable, which is the correction
ADR 0001 records after v4.3.2 and v4.4.1 were each force-re-cut three times.

Pushing a tag by hand still works, as the escape hatch for re-cutting after a
failed publish without an empty commit on main. A tag this workflow creates does
not re-trigger it, so there is no loop.

The tag is cut AFTER the approval and the publish, not in the guard. Created
earlier it would name a version that was never published when a build fails or an
approval is declined — and since tags here are immutable, that would block the
next attempt. Cutting it last makes it mean "this was published", which is the
only claim it can honestly make once the version, not the tag, is the input.

WHAT THIS COSTS, STATED RATHER THAN GLOSSED

The tag is signed by the CI key, not the maintainer's YubiKey — which cannot sign
inside a runner, and whose non-exportability is exactly what makes it worth
trusting. The chain still ends in hardware because the CI key is certified by it.

But no signature on a release now asserts that a person authorised it. That claim
moves entirely to the two gates around publication: `main` accepts only reviewed
pull requests, and publishing requires an approval from a named reviewer on a
protected environment. BRANCHING.md and SECURITY.md said the opposite and are
corrected — a verification instruction that overstates what it proves is worse
than none, because someone acts on it.

ADR 0001 still describes the tag-triggered flow and needs superseding.
STRICTER — the policy is now a gate rather than a promise

CLAUDE.md says "never ship a deprecated or scheduled-for-removal API — treat it
as a blocker, not a warning". Nothing enforced it. The verifier's default failure
level is COMPATIBILITY_PROBLEMS + INTERNAL_API_USAGES + OVERRIDE_ONLY_API_USAGES,
so a deprecated usage was reported and the job went green anyway. A rule that
lives only in prose is not a rule; DEPRECATED_API_USAGES is what makes the
sentence true. Verified green today, so it lands with no debt to forgive.

EXPERIMENTAL_API_USAGES is deliberately excluded, and that is a decision rather
than an oversight: DiffTabCleanup uses ProjectCloseListener.projectClosingBeforeSave
knowingly, because it is the only hook that runs before the workspace is written.
An experimental API is acceptable with a reason. A deprecated one is not, because
it has an announced removal and this plugin must keep working across 251 → 262.

LEANER — four measured duplications, none of them a weakened gate

- Every commit ran the pipeline TWICE. `push` on topic branches and
  `pull_request` both fire, and the concurrency group keyed on `github.ref`
  differs between them (refs/heads/x vs refs/pull/N/merge), so neither cancelled
  the other. Keyed on the commit now: same SHA, same group, duplicate cancelled.
- Superseded runs were only cancelled for pull requests. Three pushes in a row
  left three full pipelines racing, each spending ten minutes downloading IDEs
  for a commit that had already been replaced.
- The JVM suite ran twice. `koverVerify` depends on `:test`, so putting it in the
  `Static analysis` job re-ran the entire suite on a second runner with a cold
  cache. Coverage is a property OF a test run and now shares its job.
- The plugin was built twice. `verifyPlugin` already produces the distributable;
  `Build plugin` built its own. That was not only wasteful but subtly wrong — the
  bytes being asserted were never the bytes that were verified. It now downloads
  the verified artifact.

Job DISPLAY NAMES are unchanged, deliberately: a ruleset references a required
check by its name, so renaming one does not fail the gate — it silently stops
applying it. No rulesets need reapplying.

Also gives drift.yml the concurrency group it was missing (queue, do not cancel:
a half-written drift report is worse than a late one).

Caught while writing this: the SHA I pinned actions/download-artifact to was
invented. Verified against the API and corrected. Pinning by SHA protects nothing
if the SHA is made up.
Nine open Dependabot PRs, each firing the full pipeline — verifier included, ten
minutes and 1.25 GB of IDE downloads apiece — for dependency bumps that are
reviewed in seconds.

Five of them were SECURITY updates (undici, ip-address, fast-uri, hono, postcss:
exactly the `npm audit` findings). Two things about that stream were not
understood when this file was written, and both are documented at the setting now:

  - security updates ignore `open-pull-requests-limit` entirely, which is how five
    arrived under a limit of three;
  - the existing group did not cover them, because `applies-to` defaults to
    version-updates.

So each ecosystem now has an explicit `applies-to: security-updates` group. These
are transitive devDependencies that are never distributed — `npm audit --omit=dev`
reports 0, and the artifact contains zero node_modules entries — so reviewing them
one at a time bought nothing.

github-actions and gradle had no grouping at all. Actions are pinned by full
commit SHA, so a bump is a one-line change per action; grouping them costs no
review fidelity. Gradle groups minor and patch only: a MAJOR keeps its own PR
deliberately, because that is the ecosystem where a bump can hang the headless
suite — the 2.16 -> 2.18 platform-plugin attempt did exactly that, and it deserved
its own run and its own decision.

Syntax verified against GitHub's Dependabot options reference rather than written
from memory, after inventing an action SHA earlier today.
I edited build.gradle.kts and ran verifyPlugin and `help` against it, but not
spotlessCheck — so `Static analysis` failed on spotlessKotlinGradleCheck for a
purely mechanical reason. The formatter is a gate like any other and running a
subset of the gate is the same as not running it.
…ipelines

ci: release from main, enforce zero deprecations, remove duplicate work
Bumps [org.junit:junit-bom](https://github.com/junit-team/junit-framework) from 5.11.4 to 6.1.2.
- [Release notes](https://github.com/junit-team/junit-framework/releases)
- [Commits](junit-team/junit-framework@r5.11.4...r6.1.2)

---
updated-dependencies:
- dependency-name: org.junit:junit-bom
  dependency-version: 6.1.2
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps the actions group with 1 update: [actions/download-artifact](https://github.com/actions/download-artifact).


Updates `actions/download-artifact` from 7.0.0 to 8.0.1
- [Release notes](https://github.com/actions/download-artifact/releases)
- [Commits](actions/download-artifact@37930b1...3e5f45b)

---
updated-dependencies:
- dependency-name: actions/download-artifact
  dependency-version: 8.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
...

Signed-off-by: dependabot[bot] <support@github.com>
…b_actions/actions-1fa769d870

build(deps): bump actions/download-artifact from 7.0.0 to 8.0.1 in the actions group
Three changes, all aimed at the same thing: the pipeline was spending its time on
work nobody needed.

DEVELOP NO LONGER REQUIRES BRANCHES TO BE UP TO DATE

`strict_required_status_checks_policy` meant every merge into develop invalidated
every other open PR, forcing each to update and re-run the entire suite. With two
grouped Dependabot PRs that is annoying; with five it makes a release impractical,
and the cost is paid on exactly the changes that least deserve scrutiny.

What the setting guards against is real — two PRs that are green apart can break
together — so it is not simply dropped. It is dropped where a second net exists:
CI runs on every push to develop, so a semantic conflict is caught there, on
develop, before anything reaches main. `main` KEEPS the strict policy: one merge
per release, and it is the merge that publishes.

THE VERIFIER RUNS WHERE THE ANSWER MATTERS

~10 minutes and 1.25 GB of IDE downloads, previously on every push to every topic
branch. Now on pull requests and on the protected branches. It remains a required
check, so nothing merges without it. The loss is early detection mid-branch, which
is a genuine cost rather than free savings.

TOPIC BRANCHES WRITE THEIR OWN CACHE

Read-only was blanket-applied to everything but develop and main, so a topic
branch restored the shared cache and saved nothing — every push re-downloaded what
the previous one had already fetched. Read-only now applies to pull requests from
forks only.

Verified against GitHub's cache documentation rather than assumed: "Workflow runs
cannot restore caches created for child branches or sibling branches", and a cache
created on a pull request is written to the merge ref and "can only be restored by
re-runs of the pull request". A topic branch cannot reach what develop reads back;
the isolation is the platform's, not ours.

NB the ruleset change needs ./scripts/apply-rulesets.sh to take effect.
~10 minutes and 1.25 GB of IDE downloads per run, previously paid on every
iteration of a branch nobody was about to merge. It now runs only on pushes to
develop and main — not on topic branches, not on pull requests.

This REQUIRED dropping "Plugin verifier" and "Build plugin" from the required
checks on both rulesets, and that is not a detail: a required check whose job
never runs is never reported, so the pull request would wait forever. The two
changes have to move together or the branch becomes unmergeable.

What still covers a release: release.yml re-runs the FULL gate — verifier
included — on the exact tree being published, behind the environment approval, so
nothing reaches the Marketplace unverified. What is genuinely lost is catching a
binary incompatibility at pull-request time rather than after the merge into
develop. A real regression in feedback latency, accepted deliberately.

Dependabot also moves to monthly, and majors are no longer bot-proposed in any
ecosystem: they are where a bump actually breaks something and where CI is not
enough on its own — the platform-plugin 2.16 -> 2.18 attempt passed CI and hung
the headless suite locally, inside the platform's own test fixture. Security
updates are unaffected by either change.
Corrects the previous commit, which moved the verifier in the wrong direction. It
ran it only on PUSHES to develop and main and dropped it from main's required
checks — so a pull request from develop into main, the merge that publishes,
would not have run it at all. That is precisely the door that has to be guarded.

The policy now matches the intent: topic branches and pull requests into develop
run the fast checks (compile, our own JVM and frontend suites, static analysis,
dependency audit) and iterate quickly. Any pull request targeting main also runs
the plugin verifier and the artifact assertions, both required again on main
alongside the two CodeQL analyses.

The verifier is the only thing that catches a BINARY incompatibility across the
declared 251 -> 262 range — compiling against 252 proves nothing about 262, which
is exactly how the 4.4.1 /login regression shipped. Skipping it on a branch is a
latency trade; skipping it on the way to a release would not be.
Branches now run JVM tests and frontend tests and nothing else. Static analysis
and the dependency audit join the plugin verifier behind the develop -> main
door, where the exhaustive gate belongs.

Required checks match, because they have to: a check whose job does not run is
never reported and would block the pull request forever. develop requires the two
suites; main requires all eight.

The costs, named rather than discovered later:

  - A formatting, detekt, ESLint or coverage failure now lands ON develop and is
    fixed by a follow-up commit, instead of being caught in the pull request. That
    happened today with spotlessKotlinGradleCheck, and the PR is what caught it.
  - The dependency audit no longer runs on the Dependabot PR that proposes a bump.
    It runs once that bump is on develop, and again before it can reach main, so
    nothing ships un-audited — the finding just arrives one merge later.

What this buys is the thing that was actually hurting: a branch iterates in about
three minutes instead of thirteen, and nothing is promoted to main without the
full gate, plus release.yml re-running all of it on the exact published tree.
…e/org.junit-junit-bom-6.1.2

build(deps): bump org.junit:junit-bom from 5.11.4 to 6.1.2
A branch with an open pull request already fires `pull_request` on every push to
it (the `synchronize` event), so keeping a `push` trigger meant two complete
pipelines per commit for identical information. This removes the duplication at
its source instead of relying on the concurrency group to cancel one in time.

The result is the intended shape: a PR into develop runs the two required suites
once, a PR from develop into main runs all eight.

Two things this gives up, recorded because each removes something the current
setup was leaning on:

  - There is no longer a CI run on the push a merge into develop creates. That
    run was the stated justification for dropping the up-to-date requirement on
    develop: two pull requests that are green apart can break together, and
    develop's own run was what would have caught it. It is now caught at the pull
    request into main, where the full gate runs — later, but still before
    anything is published.
  - A branch with no open pull request gets no checks at all, and a pull request
    from a fork is the only path that would ever exercise them for an outside
    contributor.

release.yml is untouched: it carries its own `push: branches: [main]` trigger and
still fires on the merge that publishes.
Points the five toolchain jobs at ghcr.io/serialexperimentslainnnn/cc-ci and drops
the setup-java and setup-node steps, which provisioned inside a container that
already has both. `Build plugin` keeps no container: it only unzips an artifact.

GRADLE_USER_HOME is set per job and MUST match the value in the Dockerfile. If the
two diverge nothing fails — the run simply re-downloads everything the image
already holds, and the image appears to have bought nothing. That silence is the
reason it is stated at the setting rather than assumed.

NOT VERIFIED against a real image: at the time of writing it has not been built or
pushed. Two things have to be true before this can merge, and both fail in ways
that look like something else:

  - the package must be public (or linked to this repo), or every job dies on a
    401 that reads like a wrong image name;
  - the warmed caches must actually be in the image — `docker run --rm IMAGE
    sh -c 'ls /opt/gradle-home/caches'` answers it in seconds.

Also still open: whether gradle/actions/setup-gradle should stay. It restores its
own cache over GRADLE_USER_HOME, so it now layers on top of the baked one. That
may be a useful increment or redundant work; it needs measuring, not guessing.
The package stays private. Each container job authenticates with the GITHUB_TOKEN
the run already has, so there is no new secret to create, store or rotate, and the
credential expires with the job.

`packages: read` is granted per job rather than at the top level, keeping the
default token read-only on everything else. Without it the pull fails with a 401
that reads like a wrong image name rather than a permission problem — which is
exactly the kind of error that gets debugged in the wrong place.

One prerequisite this does NOT remove: the package must be linked to this
repository, or the token has no grant on it. That is done once, from the package
settings, and it is what makes "same account" mean "same permissions" here.
The cleanup step wipes /warmup, and `npm ci` had installed node_modules inside it
— so the image built the frontend dependencies and deleted them moments later.
The warm-up looked like it worked and bought nothing: CI would re-download the
whole tree on every run, silently, because nothing fails when a cache is missing.

Setting npm_config_cache moves the reusable part to /opt/npm-cache, which the
cleanup does not touch. node_modules stays disposable, and that is correct
independently of this bug: it must match the package-lock.json of the commit CI
checks out, not the one that happened to be current when the image was cut.

Found by a question about what that `rm -rf` actually deletes, which is a better
review than reading the line I had just written myself.
The image was already pulled by most jobs; the remaining ones provisioned
their own JDK, Node and Gradle cache and so ran on a toolchain nothing else
had used. Now `Build plugin`, the protocol-drift check and the release gate
use it too, which is the point of having built it.

`gradle/actions/setup-gradle` is removed everywhere rather than set to
read-only, because it was not doing the job it appeared to be doing. The warm
GRADLE_USER_HOME measures 31 GB — 23 GB of extracted IDE transforms under
caches/9.5.1 and 7.7 GB of downloaded IDE artifacts under modules-2 — and an
Actions cache entry is capped at 10 GB per repository. It could only ever have
stored a fraction, evicted it, and re-downloaded the rest next run. The image
has no such ceiling. The trade is explicit and worth stating: refreshing what
CI has cached is now a deliberate rebuild-and-push, not something that drifts
between runs.

Also removes the verifier's `Free disk space` step. Inside a container those
paths are the IMAGE's, not the runner's, so it had been freeing nothing while
looking like this job's safety margin. The margin now comes from the IDEs
being baked: nothing is downloaded or extracted at verify time.

Recorded because it is the failure everyone hits once: the private package
must be granted Read access to this repository in its own settings. The
`packages: read` permission widens what the token may ASK for; it does not
authorise it against a package the repo was never linked to, and without the
link the pull fails with a bare `denied` that reads like a wrong image name.

Not containerised, deliberately: `publish`, which holds the Marketplace token
and the signing key and is not a test, and CodeQL, which is weekly and is
where a container breaks quietly.

Not verified: that the image builds with no network at all. The one attempt
failed on uid mapping, which says nothing about CI, where the container runs
as root.
…code-for-jetbrains into feature/update-pipelines
The CI image was 38.1 GB, 29.1 GB of it the IDEs `verifyPlugin` downloads. Every job in ci.yml
pulls its own copy on its own runner, so `Initialize containers` measured 5m37s on a job whose
actual work is an 8-second vitest run, and 38 GB on a runner with ~25-30 GB free on the root
volume was also flirting with `No space left on device`.

The verifier is the only consumer of those IDEs and it runs only on a pull request from develop
into main, so it now downloads them when it runs. The set also moves — the verifier resolves from
the EAP/RC channels — so the baked copies stopped matching on JetBrains' release schedule, not
ours.

The warm-up itself was not warming anything. `dependencies --configuration compileClasspath`
resolves dependency metadata; it never triggers the artifact transform that EXTRACTS the IntelliJ
Platform, which is where the GB are. Measured: it leaves caches/*/transforms at 179 MB with no
extracted IDE in it. What warmed the platform was the `verifyPlugin` step, so removing that step
alone would have shipped a cold image that re-resolved the platform in every job — and the
`> /dev/null 2>&1 || true` meant nothing would have said so. It is now `testClasses`,
unredirected and without `|| true`, so a failed warm-up fails the image build.

Verified with the network disabled inside the image: `testClasses` compiles in 33s and the 84
frontend tests pass from the baked npm cache. 38.1 GB -> 8.07 GB, 3.61 GB compressed.

Also in this change:

- .dockerignore: the tree reaching `COPY` goes from 2.09 GB to 3.2 MB (node_modules, build/,
  .git). `node_modules` is now removed in the same layer `npm ci` creates it, since a later `rm`
  hides space rather than reclaiming it.
- python3 is installed explicitly. It was arriving transitively through dnf-plugins-core, which
  was itself unused — the Adoptium repo file is written with printf, not dnf config-manager, and
  curl is already in the base image — and bin/fake-claude, the stand-in the integration tests
  drive a real ClaudeSession against, is a `#!/usr/bin/env python3` script. Dropping the unused
  package without naming python3 would have broken the integration suite.
- One dnf transaction with tsflags=nodocs, git-core instead of git, /usr/share/locale removed.
- codeql.yml: the matrix is split so java-kotlin runs in the image and inherits the warm cache,
  dropping a setup-java that handed it a different JDK from the rest of the pipeline, while
  javascript-typescript stays on a bare runner where it is faster. Both display names are
  unchanged: they are required checks in .github/rulesets/main.json, and a renamed job stops
  applying its gate silently.
- No `:latest` anywhere; the image is `:base`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant