Release 5.0.0 - #35
Merged
Merged
Conversation
Audited the repository against git-workflow-standards. Adds the governance the standard requires and that this repo lacked, minus anything CI-bound (no CI gate exists yet - a local lab is planned). - commitlint.config.mjs + a VERSIONED .githooks/commit-msg gate, so the rule travels with the repo instead of living in one laptop. The hook self-tests against a known-good message first and degrades to a warning if the toolchain is broken: a hook that fails closed on its own bugs gets bypassed with --no-verify, and then protects nothing. - .gitattributes: this plugin ships and is tested on Windows, and publishes signed zips. Without a declared policy a clone with autocrlf=true rewrites every file and invalidates artifact signatures. Vendored frontend bundles marked -diff so a bump is reviewable. - docs/adr/0001: records the two justified deviations (GitFlow for installable versioned software; GPG-on-YubiKey for real revocation and because the same key signs the distributed artifacts) and corrects a real violation - v4.3.2 and v4.4.1 were each re-cut and force-pushed three times. Published tags are now immutable; a post-tag mistake is fixed by the next patch version. CHANGELOG generation is deliberately still open: the tooling silently skips non-Conventional commits, so the message gate has to land first. Refs: git-workflow-standards §3.2, §4.1, §5.2, §5.3
The published zip redistributes third-party code - marked 12.0.0 (MIT), DOMPurify 3.0.11, highlight.js 11.9.0 (BSD-3-Clause) vendored into the plugin jar, and kotlinx.serialization 1.7.3 (Apache-2.0) as its own jars. MIT, BSD-3-Clause and Apache-2.0 all require the copyright notice and licence text to be preserved ON REDISTRIBUTION, and a file sitting in the Git repository does not accompany the binary a user installs from the Marketplace. This was unmet. processResources now packages THIRD-PARTY-NOTICES.md, LICENSE and the three licence texts into META-INF/ of the plugin jar - verified present in the built artifact, not just in the source tree. Generated at build time from one root-level source of truth rather than a checked-in copy, so the notices cannot drift from the files they describe. Every licence was verified by reading the LICENSE of the exact shipped version, never the manifest or a badge. That found a real one: DOMPurify 3.0.11 is dual "Apache-2.0 OR MPL-2.0". An OR expression is a choice the redistributor must make and record; leaving it unstated is an unmade decision. Apache-2.0 is selected, with the reasoning recorded in the notices file. kotlinx.serialization ships no META-INF/LICENSE in its jars, so its licence was read from the project source instead of inferred. Also corrects the documented JDK path across CLAUDE.md, README, CONTRIBUTING and docs/: ~/.local/jdks/jdk-21.0.11+10 no longer exists on this machine (it is now ~/.jdks/jbr-21.0.11), and gradlew failed with "JAVA_HOME is set to an invalid directory" rather than anything a build log filter would obviously catch. Refs: opensource-licensing-standards §3.5, §3.6, §5.4, §8
Audited the JCEF web UI against WCAG 2.2 AA. This is a conformance requirement, not a nice-to-have: the plugin is distributed to consumers in the EU, where Directive (EU) 2019/882 has applied since 28 June 2025 (in Spain, Ley 11/2023). 4.1.3 Status Messages (AA) - the transcript streams without ever moving focus, so a screen-reader user got NO signal that Claude started, finished, or was blocked on them: the turn simply stopped, silently. Adds a polite live region declared in the static shell (it has to exist in the DOM before text is written into it, or the first change is never announced) plus CC.announce. Only turn EDGES are announced, never per token - a region updated on every delta talks over itself and gets switched off, which is worse than silence. The most important case is a pending permission card: it appears without taking focus, so it was previously undetectable. 2.4.7 Focus Visible / 1.4.11 Non-text Contrast (AA) - the stylesheet suppressed the default outline in five places. Four had some replacement; .find-input had NONE, so focus there was invisible. Adds a :focus-visible baseline covering the controls built as <span role="button"> (the code-block Copy control, menu items, tool cards), which get nothing for free because they are not native buttons, and honours forced-colors instead of overriding the system's focus colour. prefers-reduced-motion was already correct (a universal reset covering all eight keyframe animations) and is now pinned by a test so it cannot decay into a hand-picked subset as animations are added. Refactor: the frontend test harness hand-copied the shell DOM and had already drifted - shell.html gained the live region, the harness did not, so tests exercised a DOM the product does not have. It now extracts the body from the real shell.html. A mount point added to the shell reaches tests automatically; one removed breaks the tests that relied on it. Scope stated honestly in the test file: automated checks catch ~40-57% of real barriers and none of the semantic judgements. These pin the structural guarantees that regress silently. They do not certify conformance - that still needs keyboard-only and screen-reader passes by a person. Refs: accessibility-standards §3.2, §3.4, §3.8, §4.1
`@anthropic-ai/claude-agent-sdk` sat in `dependencies` while CLAUDE.md has always said it is "protocol reference only, not distributed" — and it is: the published artifact is jars and inlined web assets, with zero `node_modules` entries (verifiable with `unzip -l`). The mismatch produced seven permanent npm-audit findings (three high) against code no user ever receives, which is worse than useless: a backlog of alarms that cannot be acted on trains you to ignore the one that matters. Moving it to `devDependencies` makes the declaration match reality, so `npm audit --omit=dev` — the distributed scope — now reports zero. The full tree still reports the same seven; they are build-time tooling and are handled as maintenance. SECURITY.md states that triage boundary explicitly, with the command to verify it rather than trust it. `checkDrift` was the constraint on this change: it reads the SDK out of `node_modules` and runs `npm update`. Both survive, because `npm install` installs devDependencies — only `--omit=dev` would break it, and nothing uses it. Verified green, and it reported real drift while doing so, so the baseline moves with it: SDK 0.3.220 → 0.3.222, binary 2.1.220 → 2.1.222, protocol surface unchanged. Also corrects the package manifest itself: it declared `"license": "ISC"` on a GPL-3.0-only repository and lacked `"private": true`, i.e. it was publishable to npm with the wrong licence. And the supported-versions table still said 2.x.
The control was strong and undocumented: nothing stated what it is defending against, which makes it unreviewable. Coverage cannot be argued, and every discussion of a proposed bypass restarts from first principles. ADR 0002 states the trust model — user trusted, `claude` binary trusted as software but untrusted as a channel, everything it relays untrusted — and runs STRIDE over the three surfaces that actually exist: the binary as a child process, third-party MCP servers, and model-returned content. The third is the one worth having written down. We do not attempt to detect prompt injection; content-level detection is unsolved and a control built on it would be a liability. The document says so explicitly and records where the defence actually sits instead: the guard judges the tool call and never the reasoning behind it, so an injection that succeeds completely still has to ask to read the key and still gets the same answer. Also records why MCP servers are denied rather than asked about, and why caller trust is an allowlist — names are attacker-supplied, and 4.4.0 is the cautionary tale of that list going stale (it failed safe, which is why the allowlist was the right shape, but it still failed). Non-goals are as load-bearing as the goals, so they are listed: we do not defend against the user, against a compromised host, or against every obfuscation. That last one draws the line triage keeps blurring — failing to *recognise* a path is a pattern gap; enforcement of a match, once made, is absolute. Claims verified against the code before writing, not from memory: `binaryPermissionMode` rewriting acceptEdits/bypassPermissions to `default`, and the CSP's `default-src 'none'` / `connect-src 'none'` / hash-pinned scripts. SECURITY.md links here so a reporter can tell a finding from a stated position before writing the email.
Until now the entire quality bar for ~13k lines of Kotlin and ~3.6k lines of shipped JavaScript rested on review. That is the thing the standards say to mechanise: if formatting is discussed in a review, a formatter is missing. Adds and configures: - detekt 1.23.8 (config/detekt/detekt.yml). Every non-default setting carries its reasoning AT the setting, not in a message nobody will find again. Three rules were re-calibrated rather than suppressed: TooManyFunctions counts only the public surface (the default penalised extracting helpers, which is the exact fix the complexity rules demand); LongParameterList ignores defaulted parameters (one absent from the call site is not part of the problem the rule exists for); MagicNumber skips the colour palette (a 0xRRGGBB inside a property named DIFF_ADDED_BG is that name's definition). - Spotless 8.9.0 / ktlint 1.8.0. max-line-length and function-naming are disabled here and left to detekt: both tools enforced them, only detekt can scope an exception to the test tree, and running both meant the stricter-but-blinder one decided. - ESLint + Prettier over the JCEF frontend, which ships inside the plugin jar and had never passed through any tool. no-eval / no-implied-eval / no-new-func are errors because the page runs under a hash-pinned CSP with no 'unsafe-eval': without this gate, code Chromium will silently refuse in a user's IDE can still reach main. Vendored marked/DOMPurify/highlight.js are excluded — a finding there is not ours to fix. - koverVerify, gated per package. Risk is not evenly spread: permission/ decides whether the agent may read an SSH key, ui/ paints a browser. Thresholds sit slightly BELOW what each package measures today, so they catch regression rather than invite test-padding. ui/, context/, process/, actions/ and util/ are excluded with the reason stated rather than gated at a token value. config/detekt/baseline.xml holds TWO entries, both about ClaudeSession, both explained inside the file. It started this work at 203; the other 201 were fixed, not frozen. A `Static analysis` job runs all of it and is a required check on main and develop, so none of the above depends on anyone remembering to run it.
Three strands that could not be separated into their own commits: the formatter rewrote the same files the refactor touched, hunk for hunk, so splitting would have meant manual surgery across ~100 files for commits that would not build in isolation. Recorded here rather than pretended otherwise — it also means there is no pure formatting commit to add to .git-blame-ignore-revs. REFACTOR — the dispatch tables, split in two levels ClaudeSession.onEvent was one `when` over 47 event types: 244 lines, cyclomatic complexity 111, the single function where every protocol concern in the plugin met. ClaudeEvent now declares seven sealed sub-interfaces and dispatch picks the group, then the variant. The grouping lives in the TYPE on purpose: a sealed hierarchy keeps the compiler checking exhaustiveness at BOTH levels, so a new protocol event nobody handles is a compile error rather than a silently dropped frame — the property checkDrift exists to protect, and not up for trade against a complexity threshold. The groups are semantic: they differ in what the host OWES the binary (a Control frame must be answered or the binary hangs; a Notice is fire-and-forget). JcefBridge.Msg and JcefChatPanel.onBridgeMessage got the same treatment. Several `when` chains were dictionaries written as control flow and are now data (ProtocolParser.parseSystem: 25 arms, 21 identical bar two names). FIXED — defects the tooling and live probing surfaced - Locale-dependent formatting. TokenFormat rendered "1,2k" on a comma-decimal machine, and worse, the trailing-.0 test then stopped matching so a flat 1000 showed as "1,0k". The same bug in JcefTheme.rgba emitted four rgba components instead of three, so the browser DISCARDED the declaration and the accent/link washes never rendered. Now pinned to Locale.ROOT. - Diff tabs were persisted into the workspace and could never be restored: they are mock:// in-memory previews, and the platform persists open tabs by URL without filtering by file system. One workspace had 13 such entries, all named "Claude · SKILL.md", producing 26 warnings on every launch. DiffTabCleanup now drops them on projectClosingBeforeSave — the one hook that runs BEFORE the state is written. - Every chat animation was dead. The accessibility pass had added a prefers-reduced-motion block, and inside JCEF that query reports `reduce` regardless of the desktop (measured: it matched while GNOME had animations enabled), so it flattened everything. The tell was Vibe Mode's rainbow, which survived because it is a setInterval, not a keyframe. Motion is now reduced only when the user asks, via a plugin setting, off by default. - Requests needing a live binary were dropped. The panel is built BEFORE session.start() runs, so requestMcp/requestVersion/requestUsage all found isRunning() == false and did nothing; each looked like its own bug. They now queue until ready. - sniffMediaType treated any RIFF container as WEBP; it now checks the form type that actually identifies the format (WAV and AVI share the RIFF header). FEAT — plan limits, in both surfaces get_usage is a control request the plugin has known about since 4.0.1 and never sent; it returns every rate-limit window at once plus the extra-credit balance. The dashboard shows one bar per window with its reset countdown; the composer readout shows a dot per window. One severity scale for both: blue under 65%, amber under 85%, red above — status hues, never the product accent, because a quota bar that is coral at 10% and coral at 95% has said nothing. Crossing 65% and 85% announces once per threshold per window (a warning that repeats every refresh is one users learn to ignore), and 85% also raises an IDE notification. ClaudeSession.rateLimits is now per-window: the single field could only ever hold whichever event arrived last, so a user on a weekly limit saw their five-hour bar and concluded they had room. Also adds CC.diagnostics, which reports what the embedded browser actually resolves to idea.log. The UI is a browser nobody can open devtools on, and without it a CSS rule that silently does not apply is indistinguishable from a backend that never sent the state — which cost three wrong diagnoses before it existed. Suite: 682 -> 689 JVM, 54 -> 67 frontend.
Documents what the release actually did, including the parts that did not go to plan — a release note that only lists wins is one nobody can act on. - CHANGELOG gains two sections: the static-analysis work (203 findings fixed rather than frozen, down to 2 recorded exceptions) and the defects it exposed. - docs/RELEASE_CHECKLIST.md gains a Coverage policy with the MEASURED per-package numbers, what is excluded and why, and the Kover 0.9.2 limitation behind the current floor+aggregate shape. It also corrects a claim that never had a basis: build.gradle.kts cited a ">=90% target documented in RELEASE_CHECKLIST.md", that document had never mentioned coverage, and the real figure was 53%. - docs/BACKLOG.md is new. Every entry was PROBED against the real binary rather than inferred from the SDK types, including the two that are deliberately not worth doing, so nobody investigates them a third time. - AGENTS.md, CI_SETUP.md, BRANCHING.md and the ADRs updated for the GitHub Actions pipeline that replaced the deleted GitLab one. Known gaps, stated rather than left to be discovered: the risk-ranked human review is not done; the reduced-motion decision generalises from a measurement on one machine; and Fable's weekly window cannot be shown at all, because get_usage returns every per-model window as null (measured, not assumed).
Squash and rebase merging are now disabled at the repository level, leaving
merge commit as the only option. BRANCHING.md said the opposite for `develop`
("squash or merge"), which was worse than saying nothing: it invited the exact
click that destroys signatures.
Both disabled methods REWRITE commits — new SHAs, new committer information —
which invalidates the author's hardware-backed signatures and replaces them with
GitHub's web-flow key. The "signed commits required" rule would still pass, and
that is the trap: the commits are signed, just no longer by the author, which is
the whole property the rule exists to provide.
The cost is recorded too: the merge node itself is GitHub-signed, because the
alternative is a local merge and a direct push, which the pull-request
requirement blocks. Relaxing that to save one commit's provenance would be a far
worse trade. Release provenance therefore rests on the signed tag, not on the
merge node.
Also corrects the status-check lists, which had not been updated when the
`Static analysis` gate was added.
CodeQL flagged two high-severity alerts on the same line of the frontend test harness: "incomplete multi-character sanitization" and "bad HTML filtering regexp". Both were correct: the pattern used to strip script tags does not match a closing tag written with a trailing space, so something that reads like a sanitiser silently is not one. The exposure was low. This reads OUR OWN shell.html off disk inside a test, and the reason script tags are stripped is load control — the harness injects the app modules itself, in a controlled order — rather than sanitisation. That is a reason to rank the finding, not a reason to keep it: a regex that looks like HTML sanitisation is a pattern someone copies to a place where it does matter. Fixed at the root rather than by widening the pattern. jsdom is already a dependency of this harness, so the file is parsed and the script elements are removed as nodes. All 67 frontend tests still pass, which is the check that the harness still builds the same DOM as the product. Recorded for the risk review: the two CodeQL contexts the rulesets require (java-kotlin, javascript-typescript) both PASSED — they report "the analysis ran". The aggregate CodeQL check, which reports "something was found", is required by neither branch. A high-severity finding can therefore merge with no gate objecting.
A regression introduced by this branch made the plugin unusable, and fixing it surfaced a run of smaller defects that all shared one shape: something failed or was still loading, and the UI said nothing at all. THE REGRESSION JcefChatPanel.pendingUntilReady was declared below the `init` block that uses it. Kotlin runs property initializers and init blocks in declaration order, so the list was still null while init ran and the constructor threw. It took the whole tab with it: no chat could be opened or restored. lastUsage/lastUsageAt had the same defect and stayed silent, being nullable and primitive. InitOrderContractTest now fails the build on any property declared after its class's init. WAITING IS NOW VISIBLE The binary is launched BEFORE the tab is built (start() only dispatches, so `claude` boots while JCEF creates its browser), and a boot screen holds the tab until the process is up. Three states, not two: running, starting, and neither — that last one is a failed launch and must clear the screen, or a missing binary leaves the tab covered forever. Context and cost no longer wait on a clock. The poll timer's initial delay equals its interval, so the first reading was a minute late; worse, that tick landed while the binary was still starting and returned early, costing a second minute. Ready, tab-open and both turn edges now poll directly, and the timer retires at the end of a turn instead of round-tripping forever on an idle session. Plan limits were pushed to the dashboard but not to the composer, so the same number appeared instantly in one place and "later" in the other. Reasoning tokens, context and output now settle at 0 rather than being omitted until non-zero: an absent figure is indistinguishable from one that failed to load. FAILURES THAT WERE HARD TO READ - The CLI wraps failed tool results in <tool_use_error>...</tool_use_error> (verified in claude 2.1.222, which carries the same text unwrapped in a sibling field). Rendered verbatim it put raw markup in a native GUI. Stripped only when it encloses the whole payload; is_error already conveys the failure. - A failed tool card stayed collapsed, so the entire message was "the header is red", and its output scrolled sideways rather than wrapping, hiding the half that says what to do instead. It now opens itself once, and error text wraps. - ToolSearch was missing from SensitiveGuard.AGENT_TOOLS, so on a session that defers tools, the call that loads every other tool's schema landed in the untrusted branch. Added, with AskUserQuestion, Mcp and FileRead/Edit/Write. Entries are only ever added here: this is a trust allowlist, not an inventory, and a missing first-party name is the 4.4.0 hard-DENY incident. - A markdown link whose href is a path did nothing when clicked: the host handled https:// and jb://open and dropped the rest without a sound. Both paths now go through one gate (LinkResolver.isOpenable). The scheme test requires two or more characters so a Windows drive letter stays a path. - Copy on a message copied and gave no feedback, which reads as a broken button; the shared flash helper is now exported rather than reimplemented, and the .copied class finally has a CSS rule — it had never had one. Suite: 692 -> 694 JVM, 67 -> 84 frontend. verifyPlugin Compatible on IC-251, IC-252, IU-253, IU-261 and IU-262.
Two CI failures, one of them ours. OURS — the `Static analysis` job died with an IOException from DriftLiveCheck. Kover instruments and aggregates EVERY `Test` task in the project, and `checkDrift` is registered as one, so it had silently become a dependency of `koverVerify`. That task is on-demand by design: it downloads the latest SDK and probes a `claude` binary installed on the machine. A runner has no such binary. The task's own documentation already said "NOT wired into `check`" — it just was not true of the coverage graph, and nothing checked that it was. It passed locally for the one reason that makes this class of bug expensive: the maintainer's machine has the binary, so "on-demand" and "wired in" looked identical until CI ran it. Verified with `koverVerify --dry-run`, which listed `:checkDrift` in the graph before and does not after. NOT OURS — the `JVM tests` job failed resolving com.jetbrains.intellij.platform:test-framework with a 502 Bad Gateway from cache-redirector.jetbrains.com. Nothing to fix; it needs a re-run. Also bumps the IntelliJ Platform Gradle Plugin 2.16.0 -> 2.18.1, which the build had been warning about on every run. Stated plainly: the plugin bump is NOT verified locally. A 2.16 -> 2.18 jump touches the whole build, and CI is the verification here rather than a claim made in advance.
Reverts the 2.16.0 -> 2.18.1 half of the previous commit. The `checkDrift`
coverage-graph fix in that commit stands; only the plugin version goes back.
On 2.18.1 the headless suite never runs. `ChatSessionManagerHeadlessTest` hangs
before its first assertion, and a thread dump puts the EDT here:
BasePlatformTestCase.setUp
-> LightPlatformTestCase.doSetup
-> IndexingTestUtil.waitUntilIndexesAreReady (267s and counting)
That is the platform's own fixture waiting for indexing that never completes.
Our code is not on the stack at all — the bump changes which platform
test-framework is resolved, so this is a fixture-level regression rather than
something fixable from this side.
Measured both ways rather than inferred: the same test on 2.16.0 finishes in 19
seconds, and the full gate is green at 694 tests.
The version is now pinned with the reason written AT the setting, because the
build prints an "outdated" warning on every single run and the next person to
see it will otherwise do exactly what I did. Re-attempt it as its own change
with the headless suite as the acceptance test, not as a drive-by on a release
branch.
I shipped that bump unverified and said so in the message; this is what it cost.
The one-line lesson is the boring one: a build-tooling bump is a change like any
other and does not get to skip the suite because it looks like configuration.
…5-standards Feature/release 5 standards
Merging `develop` into `main` now publishes. The tag is DERIVED from the version in build.gradle.kts rather than supplied alongside it, so the two can no longer disagree — the mismatch the old flow guarded against with a comparison simply cannot occur. A merge that does not bump the version publishes nothing: the workflow finds the tag present, logs a notice and stops. It does not fail. `main` legitimately receives merges that are not releases, and a red run on each of those is an alarm people learn to ignore. Published tags stay immutable, which is the correction ADR 0001 records after v4.3.2 and v4.4.1 were each force-re-cut three times. Pushing a tag by hand still works, as the escape hatch for re-cutting after a failed publish without an empty commit on main. A tag this workflow creates does not re-trigger it, so there is no loop. The tag is cut AFTER the approval and the publish, not in the guard. Created earlier it would name a version that was never published when a build fails or an approval is declined — and since tags here are immutable, that would block the next attempt. Cutting it last makes it mean "this was published", which is the only claim it can honestly make once the version, not the tag, is the input. WHAT THIS COSTS, STATED RATHER THAN GLOSSED The tag is signed by the CI key, not the maintainer's YubiKey — which cannot sign inside a runner, and whose non-exportability is exactly what makes it worth trusting. The chain still ends in hardware because the CI key is certified by it. But no signature on a release now asserts that a person authorised it. That claim moves entirely to the two gates around publication: `main` accepts only reviewed pull requests, and publishing requires an approval from a named reviewer on a protected environment. BRANCHING.md and SECURITY.md said the opposite and are corrected — a verification instruction that overstates what it proves is worse than none, because someone acts on it. ADR 0001 still describes the tag-triggered flow and needs superseding.
STRICTER — the policy is now a gate rather than a promise CLAUDE.md says "never ship a deprecated or scheduled-for-removal API — treat it as a blocker, not a warning". Nothing enforced it. The verifier's default failure level is COMPATIBILITY_PROBLEMS + INTERNAL_API_USAGES + OVERRIDE_ONLY_API_USAGES, so a deprecated usage was reported and the job went green anyway. A rule that lives only in prose is not a rule; DEPRECATED_API_USAGES is what makes the sentence true. Verified green today, so it lands with no debt to forgive. EXPERIMENTAL_API_USAGES is deliberately excluded, and that is a decision rather than an oversight: DiffTabCleanup uses ProjectCloseListener.projectClosingBeforeSave knowingly, because it is the only hook that runs before the workspace is written. An experimental API is acceptable with a reason. A deprecated one is not, because it has an announced removal and this plugin must keep working across 251 → 262. LEANER — four measured duplications, none of them a weakened gate - Every commit ran the pipeline TWICE. `push` on topic branches and `pull_request` both fire, and the concurrency group keyed on `github.ref` differs between them (refs/heads/x vs refs/pull/N/merge), so neither cancelled the other. Keyed on the commit now: same SHA, same group, duplicate cancelled. - Superseded runs were only cancelled for pull requests. Three pushes in a row left three full pipelines racing, each spending ten minutes downloading IDEs for a commit that had already been replaced. - The JVM suite ran twice. `koverVerify` depends on `:test`, so putting it in the `Static analysis` job re-ran the entire suite on a second runner with a cold cache. Coverage is a property OF a test run and now shares its job. - The plugin was built twice. `verifyPlugin` already produces the distributable; `Build plugin` built its own. That was not only wasteful but subtly wrong — the bytes being asserted were never the bytes that were verified. It now downloads the verified artifact. Job DISPLAY NAMES are unchanged, deliberately: a ruleset references a required check by its name, so renaming one does not fail the gate — it silently stops applying it. No rulesets need reapplying. Also gives drift.yml the concurrency group it was missing (queue, do not cancel: a half-written drift report is worse than a late one). Caught while writing this: the SHA I pinned actions/download-artifact to was invented. Verified against the API and corrected. Pinning by SHA protects nothing if the SHA is made up.
Nine open Dependabot PRs, each firing the full pipeline — verifier included, ten
minutes and 1.25 GB of IDE downloads apiece — for dependency bumps that are
reviewed in seconds.
Five of them were SECURITY updates (undici, ip-address, fast-uri, hono, postcss:
exactly the `npm audit` findings). Two things about that stream were not
understood when this file was written, and both are documented at the setting now:
- security updates ignore `open-pull-requests-limit` entirely, which is how five
arrived under a limit of three;
- the existing group did not cover them, because `applies-to` defaults to
version-updates.
So each ecosystem now has an explicit `applies-to: security-updates` group. These
are transitive devDependencies that are never distributed — `npm audit --omit=dev`
reports 0, and the artifact contains zero node_modules entries — so reviewing them
one at a time bought nothing.
github-actions and gradle had no grouping at all. Actions are pinned by full
commit SHA, so a bump is a one-line change per action; grouping them costs no
review fidelity. Gradle groups minor and patch only: a MAJOR keeps its own PR
deliberately, because that is the ecosystem where a bump can hang the headless
suite — the 2.16 -> 2.18 platform-plugin attempt did exactly that, and it deserved
its own run and its own decision.
Syntax verified against GitHub's Dependabot options reference rather than written
from memory, after inventing an action SHA earlier today.
I edited build.gradle.kts and ran verifyPlugin and `help` against it, but not spotlessCheck — so `Static analysis` failed on spotlessKotlinGradleCheck for a purely mechanical reason. The formatter is a gate like any other and running a subset of the gate is the same as not running it.
…ipelines ci: release from main, enforce zero deprecations, remove duplicate work
Bumps [org.junit:junit-bom](https://github.com/junit-team/junit-framework) from 5.11.4 to 6.1.2. - [Release notes](https://github.com/junit-team/junit-framework/releases) - [Commits](junit-team/junit-framework@r5.11.4...r6.1.2) --- updated-dependencies: - dependency-name: org.junit:junit-bom dependency-version: 6.1.2 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps the actions group with 1 update: [actions/download-artifact](https://github.com/actions/download-artifact). Updates `actions/download-artifact` from 7.0.0 to 8.0.1 - [Release notes](https://github.com/actions/download-artifact/releases) - [Commits](actions/download-artifact@37930b1...3e5f45b) --- updated-dependencies: - dependency-name: actions/download-artifact dependency-version: 8.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: actions ... Signed-off-by: dependabot[bot] <support@github.com>
…b_actions/actions-1fa769d870 build(deps): bump actions/download-artifact from 7.0.0 to 8.0.1 in the actions group
Three changes, all aimed at the same thing: the pipeline was spending its time on work nobody needed. DEVELOP NO LONGER REQUIRES BRANCHES TO BE UP TO DATE `strict_required_status_checks_policy` meant every merge into develop invalidated every other open PR, forcing each to update and re-run the entire suite. With two grouped Dependabot PRs that is annoying; with five it makes a release impractical, and the cost is paid on exactly the changes that least deserve scrutiny. What the setting guards against is real — two PRs that are green apart can break together — so it is not simply dropped. It is dropped where a second net exists: CI runs on every push to develop, so a semantic conflict is caught there, on develop, before anything reaches main. `main` KEEPS the strict policy: one merge per release, and it is the merge that publishes. THE VERIFIER RUNS WHERE THE ANSWER MATTERS ~10 minutes and 1.25 GB of IDE downloads, previously on every push to every topic branch. Now on pull requests and on the protected branches. It remains a required check, so nothing merges without it. The loss is early detection mid-branch, which is a genuine cost rather than free savings. TOPIC BRANCHES WRITE THEIR OWN CACHE Read-only was blanket-applied to everything but develop and main, so a topic branch restored the shared cache and saved nothing — every push re-downloaded what the previous one had already fetched. Read-only now applies to pull requests from forks only. Verified against GitHub's cache documentation rather than assumed: "Workflow runs cannot restore caches created for child branches or sibling branches", and a cache created on a pull request is written to the merge ref and "can only be restored by re-runs of the pull request". A topic branch cannot reach what develop reads back; the isolation is the platform's, not ours. NB the ruleset change needs ./scripts/apply-rulesets.sh to take effect.
~10 minutes and 1.25 GB of IDE downloads per run, previously paid on every iteration of a branch nobody was about to merge. It now runs only on pushes to develop and main — not on topic branches, not on pull requests. This REQUIRED dropping "Plugin verifier" and "Build plugin" from the required checks on both rulesets, and that is not a detail: a required check whose job never runs is never reported, so the pull request would wait forever. The two changes have to move together or the branch becomes unmergeable. What still covers a release: release.yml re-runs the FULL gate — verifier included — on the exact tree being published, behind the environment approval, so nothing reaches the Marketplace unverified. What is genuinely lost is catching a binary incompatibility at pull-request time rather than after the merge into develop. A real regression in feedback latency, accepted deliberately. Dependabot also moves to monthly, and majors are no longer bot-proposed in any ecosystem: they are where a bump actually breaks something and where CI is not enough on its own — the platform-plugin 2.16 -> 2.18 attempt passed CI and hung the headless suite locally, inside the platform's own test fixture. Security updates are unaffected by either change.
Corrects the previous commit, which moved the verifier in the wrong direction. It ran it only on PUSHES to develop and main and dropped it from main's required checks — so a pull request from develop into main, the merge that publishes, would not have run it at all. That is precisely the door that has to be guarded. The policy now matches the intent: topic branches and pull requests into develop run the fast checks (compile, our own JVM and frontend suites, static analysis, dependency audit) and iterate quickly. Any pull request targeting main also runs the plugin verifier and the artifact assertions, both required again on main alongside the two CodeQL analyses. The verifier is the only thing that catches a BINARY incompatibility across the declared 251 -> 262 range — compiling against 252 proves nothing about 262, which is exactly how the 4.4.1 /login regression shipped. Skipping it on a branch is a latency trade; skipping it on the way to a release would not be.
Branches now run JVM tests and frontend tests and nothing else. Static analysis
and the dependency audit join the plugin verifier behind the develop -> main
door, where the exhaustive gate belongs.
Required checks match, because they have to: a check whose job does not run is
never reported and would block the pull request forever. develop requires the two
suites; main requires all eight.
The costs, named rather than discovered later:
- A formatting, detekt, ESLint or coverage failure now lands ON develop and is
fixed by a follow-up commit, instead of being caught in the pull request. That
happened today with spotlessKotlinGradleCheck, and the PR is what caught it.
- The dependency audit no longer runs on the Dependabot PR that proposes a bump.
It runs once that bump is on develop, and again before it can reach main, so
nothing ships un-audited — the finding just arrives one merge later.
What this buys is the thing that was actually hurting: a branch iterates in about
three minutes instead of thirteen, and nothing is promoted to main without the
full gate, plus release.yml re-running all of it on the exact published tree.
…e/org.junit-junit-bom-6.1.2 build(deps): bump org.junit:junit-bom from 5.11.4 to 6.1.2
…ipelines Feature/update pipelines
A branch with an open pull request already fires `pull_request` on every push to
it (the `synchronize` event), so keeping a `push` trigger meant two complete
pipelines per commit for identical information. This removes the duplication at
its source instead of relying on the concurrency group to cancel one in time.
The result is the intended shape: a PR into develop runs the two required suites
once, a PR from develop into main runs all eight.
Two things this gives up, recorded because each removes something the current
setup was leaning on:
- There is no longer a CI run on the push a merge into develop creates. That
run was the stated justification for dropping the up-to-date requirement on
develop: two pull requests that are green apart can break together, and
develop's own run was what would have caught it. It is now caught at the pull
request into main, where the full gate runs — later, but still before
anything is published.
- A branch with no open pull request gets no checks at all, and a pull request
from a fork is the only path that would ever exercise them for an outside
contributor.
release.yml is untouched: it carries its own `push: branches: [main]` trigger and
still fires on the merge that publishes.
Points the five toolchain jobs at ghcr.io/serialexperimentslainnnn/cc-ci and drops
the setup-java and setup-node steps, which provisioned inside a container that
already has both. `Build plugin` keeps no container: it only unzips an artifact.
GRADLE_USER_HOME is set per job and MUST match the value in the Dockerfile. If the
two diverge nothing fails — the run simply re-downloads everything the image
already holds, and the image appears to have bought nothing. That silence is the
reason it is stated at the setting rather than assumed.
NOT VERIFIED against a real image: at the time of writing it has not been built or
pushed. Two things have to be true before this can merge, and both fail in ways
that look like something else:
- the package must be public (or linked to this repo), or every job dies on a
401 that reads like a wrong image name;
- the warmed caches must actually be in the image — `docker run --rm IMAGE
sh -c 'ls /opt/gradle-home/caches'` answers it in seconds.
Also still open: whether gradle/actions/setup-gradle should stay. It restores its
own cache over GRADLE_USER_HOME, so it now layers on top of the baked one. That
may be a useful increment or redundant work; it needs measuring, not guessing.
The package stays private. Each container job authenticates with the GITHUB_TOKEN the run already has, so there is no new secret to create, store or rotate, and the credential expires with the job. `packages: read` is granted per job rather than at the top level, keeping the default token read-only on everything else. Without it the pull fails with a 401 that reads like a wrong image name rather than a permission problem — which is exactly the kind of error that gets debugged in the wrong place. One prerequisite this does NOT remove: the package must be linked to this repository, or the token has no grant on it. That is done once, from the package settings, and it is what makes "same account" mean "same permissions" here.
The cleanup step wipes /warmup, and `npm ci` had installed node_modules inside it — so the image built the frontend dependencies and deleted them moments later. The warm-up looked like it worked and bought nothing: CI would re-download the whole tree on every run, silently, because nothing fails when a cache is missing. Setting npm_config_cache moves the reusable part to /opt/npm-cache, which the cleanup does not touch. node_modules stays disposable, and that is correct independently of this bug: it must match the package-lock.json of the commit CI checks out, not the one that happened to be current when the image was cut. Found by a question about what that `rm -rf` actually deletes, which is a better review than reading the line I had just written myself.
The image was already pulled by most jobs; the remaining ones provisioned their own JDK, Node and Gradle cache and so ran on a toolchain nothing else had used. Now `Build plugin`, the protocol-drift check and the release gate use it too, which is the point of having built it. `gradle/actions/setup-gradle` is removed everywhere rather than set to read-only, because it was not doing the job it appeared to be doing. The warm GRADLE_USER_HOME measures 31 GB — 23 GB of extracted IDE transforms under caches/9.5.1 and 7.7 GB of downloaded IDE artifacts under modules-2 — and an Actions cache entry is capped at 10 GB per repository. It could only ever have stored a fraction, evicted it, and re-downloaded the rest next run. The image has no such ceiling. The trade is explicit and worth stating: refreshing what CI has cached is now a deliberate rebuild-and-push, not something that drifts between runs. Also removes the verifier's `Free disk space` step. Inside a container those paths are the IMAGE's, not the runner's, so it had been freeing nothing while looking like this job's safety margin. The margin now comes from the IDEs being baked: nothing is downloaded or extracted at verify time. Recorded because it is the failure everyone hits once: the private package must be granted Read access to this repository in its own settings. The `packages: read` permission widens what the token may ASK for; it does not authorise it against a package the repo was never linked to, and without the link the pull fails with a bare `denied` that reads like a wrong image name. Not containerised, deliberately: `publish`, which holds the Marketplace token and the signing key and is not a test, and CodeQL, which is weekly and is where a container breaks quietly. Not verified: that the image builds with no network at all. The one attempt failed on uid mapping, which says nothing about CI, where the container runs as root.
…code-for-jetbrains into feature/update-pipelines
The CI image was 38.1 GB, 29.1 GB of it the IDEs `verifyPlugin` downloads. Every job in ci.yml pulls its own copy on its own runner, so `Initialize containers` measured 5m37s on a job whose actual work is an 8-second vitest run, and 38 GB on a runner with ~25-30 GB free on the root volume was also flirting with `No space left on device`. The verifier is the only consumer of those IDEs and it runs only on a pull request from develop into main, so it now downloads them when it runs. The set also moves — the verifier resolves from the EAP/RC channels — so the baked copies stopped matching on JetBrains' release schedule, not ours. The warm-up itself was not warming anything. `dependencies --configuration compileClasspath` resolves dependency metadata; it never triggers the artifact transform that EXTRACTS the IntelliJ Platform, which is where the GB are. Measured: it leaves caches/*/transforms at 179 MB with no extracted IDE in it. What warmed the platform was the `verifyPlugin` step, so removing that step alone would have shipped a cold image that re-resolved the platform in every job — and the `> /dev/null 2>&1 || true` meant nothing would have said so. It is now `testClasses`, unredirected and without `|| true`, so a failed warm-up fails the image build. Verified with the network disabled inside the image: `testClasses` compiles in 33s and the 84 frontend tests pass from the baked npm cache. 38.1 GB -> 8.07 GB, 3.61 GB compressed. Also in this change: - .dockerignore: the tree reaching `COPY` goes from 2.09 GB to 3.2 MB (node_modules, build/, .git). `node_modules` is now removed in the same layer `npm ci` creates it, since a later `rm` hides space rather than reclaiming it. - python3 is installed explicitly. It was arriving transitively through dnf-plugins-core, which was itself unused — the Adoptium repo file is written with printf, not dnf config-manager, and curl is already in the base image — and bin/fake-claude, the stand-in the integration tests drive a real ClaudeSession against, is a `#!/usr/bin/env python3` script. Dropping the unused package without naming python3 would have broken the integration suite. - One dnf transaction with tsflags=nodocs, git-core instead of git, /usr/share/locale removed. - codeql.yml: the matrix is split so java-kotlin runs in the image and inherits the warm cache, dropping a setup-java that handed it a different JDK from the rest of the pipeline, while javascript-typescript stays on a bare runner where it is faster. Both display names are unchanged: they are required checks in .github/rulesets/main.json, and a renamed job stops applying its gate silently. - No `:latest` anywhere; the image is `:base`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ipelines Feature/update pipelines
The release ran as: build, publish to the Marketplace, then stamp a tag on what had already gone out. It now runs as the sequence the process actually describes — merge to main, cut and sign the tag, open the GitHub Release, then build and publish FROM that tag. Building from the tag rather than from `main` is the part that changes behaviour: `main` is a moving ref, so a merge landing between the `guard` job and the `publish` job was silently included in a release named after a different tree. `publish` now checks out `refs/tags/vX.Y.Z`. Cutting the tag first used to be unsafe, and the comment saying so was right at the time: a tag could exist for a version that was never published, and published tags are immutable here. What makes it safe now is that the whole job sits behind the `marketplace` environment, so nothing — including the tag — happens before a human approves. The residual case is a publish that fails after the tag exists; the recovery is re-running the job on that tag, which is why the asset upload carries `--clobber`. This is deliberately NOT two workflows chained by the tag push. A tag pushed with the GITHUB_TOKEN does not create a workflow run (the recursion guard), so chaining would need a PAT, a GitHub App or a deploy key — a long-lived write credential — to buy an ordering one run already achieves. `buildPlugin signPlugin publishPlugin` stays a single Gradle invocation. `publishPlugin` uploads the signed archive only if `signPlugin.didWork` and falls back to the UNSIGNED one otherwise, so splitting it to fit the new ordering is exactly how an unsigned plugin ships unnoticed. The GitHub Release is created as a draft and undrafted once the artifacts are attached: created final and empty, its download links would 404 for the length of the build. Out of band but part of the same path: the `marketplace` environment's deployment branch policy allowed only `tag: v*.*.*`, and the run's ref on the primary path is `refs/heads/main` — the deployment would have been rejected before even asking for the approval. `main` was added to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
codeql.yml already triggers on every pull request into develop, so both jobs were running there and nobody was obliged to read the result. Requiring them changes only that. Affordable in a way the other main-only checks are not: no extra run is created. And a SAST finding is the class of defect worth catching before the merge rather than at the release door, where it arrives mixed in with everything else that landed on the branch since. The contexts are the jobs' DISPLAY names. Renaming a job in codeql.yml does not fail this gate — it silently stops applying it, which is why the names are duplicated in a comment beside them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
apply-rulesets.sh deleted the key named exactly `_comment`, so the moment a second annotation was needed in one object the natural name — `_comment_codeql` — sailed through the filter and reached the API, which rejected the whole ruleset with a bare 422 naming no property. Observed, not hypothetical: it is how the CodeQL required check failed to apply. The failure reads as "the ruleset is wrong" rather than "a comment leaked into the payload", which is the expensive part. Now every key with the `_comment` prefix is stripped, so annotating a block twice is safe. The key added in the previous commit is renamed back to the convention as well. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A release is a claim that develop is a finished state. An open pull request from Claude or from Dependabot contradicts it: the change was meant to be in this release and sits one click away from being in it. Merging past it does not lose the work — it ships a version whose CHANGELOG was written as if the work had landed. For Dependabot it also means releasing with a known dependency update unmerged, the one class of pending change an advisory gets written about. A status check rather than a ruleset entry because it cannot be a ruleset entry: rulesets speak of checks, signatures and approvals, and have no vocabulary for "no other pull request exists". main.json requires this job by DISPLAY name, like every other gate. The author match is anchored, not a substring, so a human whose username contains "claude" is not caught by a release gate. It covers both renderings, since which one appears depends on how each integration is installed: `app/<slug>` for an app, `<name>[bot]` for a bot user. The cost is stated in the workflow rather than left to be discovered: Dependabot's resting state is "has something open", so draining that queue becomes a release step. That is the intended trade, and if the gate starts being routinely in the way the answer is to merge Dependabot more often, not to widen the filter. NB the check cannot be required until this job exists on develop — a required check that never reports blocks the pull request forever. Merge first, run scripts/apply-rulesets.sh second. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
5 of 6 tasks
One image served every job, and an image is a cost paid PER JOB: each one pulls its own
copy onto its own runner. So `Frontend tests` — whose work is an 8-second vitest run —
spent 1m05s pulling a JDK, a Gradle distribution and 3.4 GB of extracted IntelliJ
Platform it never opened.
There are now two, named for what they carry and pinned by version, not a floating tag:
node-test:v1.0.0 462 MB. Node, npm, warm npm cache. For `Frontend tests` and
`Dependency audit`.
jvm-test:v1.0.0 8.08 GB. The above plus the JDK, Gradle and the extracted platform.
For every job that runs Gradle: JVM tests, Static analysis, Plugin
verifier, CodeQL (java-kotlin), drift, and the release gate.
jvm-test is built FROM node-test, so it is not a second copy — the registry stores the
shared layers once — and it carries Node deliberately: `Static analysis`, `drift` and the
release gate each run Gradle AND npm in one job. Splitting those would add a whole extra
image pull, which is the cost this change exists to remove.
Two jobs now pull NOTHING. `Build plugin` downloads an artifact and runs `unzip`, `grep`
and `ls` — it does not build anything despite the name, and it was pulling GB to do it.
The bot-PR gate only calls `gh`. Both run on a bare runner.
Why the pull cannot simply be cached, since it is the obvious first idea: a `container:`
job pulls in `Initialize containers`, which runs BEFORE the job's first step, so there is
no point at which an `actions/cache` step could run first — and every job starts on a
fresh runner with no shared layer cache. Restoring a tarball instead is slower, not
faster: it moves the same bytes from a store further away than ghcr and capped at 10 GB
per repository. The only lever is how much each job downloads, which is what this does.
The images also run `dnf upgrade --refresh` before installing. That makes the build
non-reproducible, which is acceptable here precisely because the tag is explicit: what CI
runs is frozen at v1.0.0, and bumping it is the deliberate act that moves it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Since `CodeQL (java-kotlin)` moved into a container, every run logs:
ERROR: ld.so: object '.../${LIB}_${PLATFORM}_trace.so' from LD_PRELOAD
cannot be preloaded (cannot open shared object file): ignored
`$LIB` and `$PLATFORM` are glibc dynamic string tokens that ld.so expands at load time, so one
variable covers several ABIs. Measured locally rather than assumed: Fedora 44's loader expands
them to `lib64_x86_64` and loads the file happily; Ubuntu 24.04's does not resolve that form at
all. CodeQL is built and tested on Ubuntu runners, so the shipped filename matches Ubuntu's
expansion and Fedora asks for a name the bundle does not contain.
The step symlinks the name Fedora asks for onto the 64-bit tracer that is actually shipped, and
prints the directory listing first — that `ls` is the evidence the layout still matches, and it is
deliberately not guarded with `|| true` so a future CodeQL release that moves these files fails
loudly instead of quietly reverting to the current behaviour.
Whether the message was ever more than noise is NOT established, and this commit does not claim it
was: github/codeql-action#1113 records the same line as harmless with the real failure elsewhere.
It is removed because a permanent ERROR in a security gate's log trains you to skim past the one
that matters — and this gate has just become a required check on develop. The same run will settle
it: if extraction was actually being skipped, the build step's behaviour changes with the symlink
in place.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous attempt globbed the tool cache:
ls -d /__t/CodeQL/*/x64/codeql/tools/linux64
and that matched TWO directories — the runner image ships one CodeQL bundle and the action had
downloaded another, so 2.26.1 and 2.26.2 sat side by side. The variable held two newline-separated
paths and every command after it failed:
ls: cannot access '.../2.26.1/...'$'\n''.../2.26.2/...*_trace.so': No such file or directory
Picking one by sort order would have been a nicer-looking guess. LD_PRELOAD already names the exact
file the loader will be asked for, so the whole inference disappears: the step now reads it, takes
its dirname, and substitutes the tokens the way Fedora's ld.so resolves them.
The substitution is on the value read from the ENVIRONMENT, where `${LIB}` and `${PLATFORM}` are
literal characters that a shell assignment does not re-expand. Verified with a literal env value
rather than assumed — the first test of it was wrong (it built the string with double quotes, so
bash expanded both tokens to empty before the substitution ever ran) and looked like a real failure.
It also fails loudly if LD_PRELOAD is unset, since that means tracing was never initialised and this
step is patching a problem that no longer exists in the shape it was written for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`Initialize CodeQL` runs inside the container and writes to $GITHUB_ENV:
LD_PRELOAD=/__t/CodeQL/<v>/x64/codeql/tools/linux64/${LIB}_${PLATFORM}_trace.so
`/__t` is the name the tool cache has INSIDE the container; on the host the same directory is
/opt/hostedtoolcache — which is precisely what this job exported back when it ran on a bare runner,
and why the message never appeared there. But $GITHUB_ENV is consumed by the runner process, which
lives on the HOST, so its helpers start with an LD_PRELOAD naming a path that does not exist from
where they stand, and ld.so logs "cannot be preloaded ... ignored" on every step.
The fix makes one string valid from both namespaces: symlink /opt/hostedtoolcache to /__t inside the
container, and rewrite the variable to use it. Verified inside the real jvm-test image before being
written here — the rewritten path resolves and the real tracer loads through it.
Three earlier hypotheses were wrong and are recorded in the workflow so nobody re-runs them:
- Not a missing Fedora package. The bundle ships the full matrix (lib/lib64/lib32/x86_64-linux-gnu
times x86_64/haswell/i686/xeon_phi) and lib64_x86_64_trace.so is present in the container, 0755.
- Not a glibc difference. Fedora 44's loader expands the tokens and loads the real tracer fine in
this exact image: AT_PLATFORM x86_64, every dependency satisfied.
- Not a broken database. The tracing that matters happens inside the container, where the path was
always valid — which is why the scan succeeded throughout.
This supersedes the symlink patch from df99ffe and 284233e, which the evidence showed was a no-op:
the file it created already existed.
Residual, stated rather than hidden: the ERROR still appears once at the start of the rewriting step
itself, since the new value cannot apply before the step that sets it. The steps that run the
compiler get the corrected value.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ipelines ci: release ordering, CodeQL gate and a release-readiness check
The recovery path the workflow documents did not work. If the publish failed after the tag was cut, the intended fix was to re-run the job — but `fetch-depth: 0` fetches tags, so a bare `git tag -s` exited non-zero on the second pass, the job died before reaching `gh release upload --clobber`, and the release was stuck needing a human to delete a tag this repository treats as immutable. Which is the exact situation immutability exists to prevent. Re-running the whole WORKFLOW cannot substitute: `guard` would see the tag on the remote and correctly report the version as already released, skipping verify and publish entirely. So the job re-run is the only path, and it has to survive an existing tag. It now detects the ref, verifies the signature that is already on it, and skips re-cutting. Also corrects a justification that was simply false. The header claimed the tag is checked out because `main` is a moving ref that a later merge could slip into the release. `actions/checkout` defaults to `github.sha`, so every job was already pinned to the triggering commit — building from the tag is a provenance statement, not a race fix. Both claims were flagged in the review of the commit that introduced them; this is that follow-up. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
4 of 5 tasks
…ag-idempotent fix(release): make the tag step idempotent for job re-runs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 5.0.0
Summary
The release door for 5.0.0, the standards-compliance major.
developis 38 commits ahead ofmain;the last published tag is
v4.4.1andbuild.gradle.ktsalready declares5.0.0. Merging this is whatmakes a
v5.0.0tag reachable frommain, which is the preconditionrelease.yml'sguardjob assertsbefore any credential comes into scope.
This PR carries no new code of its own. Its purpose is to run the full gate —
Plugin verifier,Static analysisandDependency auditonly open on a pull request intomain, so this is the first timethe 38 commits are judged by all of them together.
Related issue
n/a
Type of change
Risk and rollback
Risk: this is a major, and it changes packaging and process rather than the plugin's runtime
behaviour. The user-facing surface that is genuinely new is the plan-limits panel; the rest is CI, docs,
static analysis and the dependency-scope correction. The concrete risk is the release machinery itself —
release.ymlpublishes to the Marketplace behind a human approval on themarketplaceenvironment, andthis is the first release to run through it end to end.
Rollback: reverting the merge does not roll anything back for a user who already updated — a published
plugin only moves forward. If 5.0.0 is bad in the wild, the exit is a 5.0.1 patch, not a revert. Marketplace
also allows delisting a version, which stops new installs but does not touch existing ones. Nothing in this
release migrates persisted settings, the transcript format, or the permission surface, so a user who runs it
and then downgrades manually keeps a readable
claude-code.xml.Checklist
developbranch (ormainonly for hotfixes).Deliberately not. This is the
develop->mainrelease door described inADR 0001, not a feature PR.
./gradlew testandnpm testpass locally — 694 JVM tests (0 failures, 2 Windows-only skips)and 84 frontend tests, plus
koverVerify,detekt,spotlessCheck, ESLint, Prettier andnpm audit --omit=dev(0 vulnerabilities).buildPluginclean.verifyPluginpasses locally / is Compatible across 251 -> 263.*.Not run locally for this PR — it is the
Plugin verifierjob on this very pull request, which isthe first place it opens. Judge it there rather than on my word.
failureLevelincludesDEPRECATED_API_USAGES, so the verifier job enforces this rather than a reviewer)../gradlew checkDriftgreen. Not re-run for this PR. The baseline was last verified atclaude2.1.222 / SDK 0.3.222 during the 5.0.0 work;drift.ymlruns weekly and files an issue.THIRD-PARTY-NOTICES.md, which now ships inside the artifact.CHANGELOG.md— reproduced below.CONTRIBUTING.md,CLAUDE.mdanddocs/adr/.How was this tested?
./gradlew test) and frontend tests (npm test)../gradlew runIde).Stated plainly rather than ticked: no live in-IDE manual pass was run for this PR specifically. For a
major that touches the permission surface's documentation and ships a new dashboard card, that is the gap a
reviewer should weigh — the automated gate does not cover it.
Notes for reviewers
The CHANGELOG entry below predates the last commit on
develop.04957a8("ci: drop the verifier IDEsfrom the image and fix the warm-up") landed after the 5.0.0 section was written, and is not reflected in it.
It is CI-only and reaches no user, but the entry is not a complete account of what is on this branch. Worth
folding in before tagging if the CHANGELOG is meant to be read as the record of the release. Summary of it:
the CI image was 38.1 GB because it baked the verifier's IDEs, costing every job a 5m37s container init;
they are gone (38.1 GB -> 8.07 GB, container init now measured at 1m05s), and the image's warm-up was found
to have never warmed anything —
dependencies --configuration compileClasspathresolves metadata but nevertriggers the transform that extracts the IntelliJ Platform, and
> /dev/null 2>&1 || truehid it.What this PR is really for is the three jobs that have never run on these commits:
Plugin verifier(the only thing that catches a binary incompatibility across 251 -> 263.*, and the asymmetry that let the
4.4.1
/loginregression ship),Static analysisandDependency audit. If any of them is red, that isnew information and not a flake.
[5.0.0] — 2026-08-05
The standards-compliance major. Not a feature release: the repository was taken through the standards
catalogue domain by domain, and the major number reflects that the code changed to comply, not only the
documentation. The break this major records is one of process and packaging, and it is stated plainly rather
than hidden in a patch.
It did not stay purely that, and saying so is cheaper than letting a reader discover it: the release also
carries the plan-limits panel and a run of user-facing fixes (below). Nothing is removed or behaves
differently on purpose — but a release note claiming "no user-facing change" while shipping a new dashboard
card would be the kind of small untruth that makes the rest of the document unusable as evidence.
Security
@anthropic-ai/claude-agent-sdksat independenciesalthough it is protocol reference material — kept so the Kotlin layer can be diffed against the binary's real surface — and is not executed or packaged. The published artifact contains jars and inlined web assets and zeronode_modulesentries, which anyone can confirm withunzip -l build/distributions/*.zip | grep -c node_modules. The consequence of the wrong declaration was seven permanentnpm auditfindings (three high) against code no user ever receives: an alarm backlog that cannot be acted on, which is worse than no alarm because it trains you to ignore the one that matters. Moved todevDependencies, sonpm audit --omit=dev— the distributed scope — now reports zero.SECURITY.mdstates the triage boundary explicitly, with the command to verify it rather than a request to trust it.SensitiveGuardwas strong and undocumented: nothing said what it defends against, which makes coverage unarguable and restarts every bypass discussion from first principles. The ADR states the trust model (the user trusted; theclaudebinary trusted as software but untrusted as a channel; everything it relays — model output, tool inputs, MCP traffic, file contents, fetched pages — untrusted) and runs STRIDE over the three real surfaces. On indirect prompt injection it records the position deliberately: detection is not attempted, because content-level detection is unsolved and a control built on it would be a liability. Injection is assumed to succeed, and the defence sits where success does not pay — the guard judges the tool call and never the reasoning behind it, so a perfectly-injected model still has to ask to read the key, and still gets the same answer. Non-goals are listed as explicitly as goals..gitignorecovered build output and nothing else, so the working tree was one wrong answer away from a committed private key:scripts/bootstrap-ci.shasks where to save a generated JetBrains signing key, and answering.dropsprivate.peminto the repository. It now leads with a secrets section —*.pem,*.key,*.p12,*.jks,chain.crt,passphrase,private.asc,*.token,.npmrc,.netrc— with a single negation fordocs/ci-signing-key.asc, the one key file that must be committed since without the public half nobody can verify a release. Verified by creating each of those files and confirminggit check-ignoreblocks it while the public key stays committable. Secrets come first in the file for a reason: a build artifact committed by accident is noise, whereas a private key committed by accident is burned — forks, clones, forge caches and CI logs mean rewriting history does not un-leak it, and the key has to be rotated regardless. The file itself remains untracked by design (it ignores itself): a published.gitignoreis a public inventory of a maintainer's local directories and tooling, which is reconnaissance for no benefit to anyone installing the plugin.SECURITY.md's supported-versions table still said2.x.Added
.gitlab-ci.ymlhad been asserting for months that GitHub Actions was "capped (billing)". That was false — the repository is public, and Actions on standard hosted runners is free and unmetered for public repositories; the account's Actions permissions were verified enabled. A false constraint written into a config file gets believed for years, which is precisely what happened.ci.ymlnow runs the full gate ondevelop,mainand everyfeature/**,bugfix/**andhotfix/**branch — not only on the PR, because a bar you meet only at PR time is a bar you discover late.codeql.ymladds SAST over Kotlin and JavaScript.release.ymlpublishes to the Marketplace only when three things hold at once: avX.Y.Ztag; the tagged commit reachable frommain, asserted before any credential is in scope; and a human approval on themarketplaceenvironment, where the four credentials are scoped and exist for no other job. The middle gate is the load-bearing one — without it, anyone who can push a tag can publish from any code, and the review the approval assumes becomes optional.drift.ymlrunscheckDriftweekly and files an issue rather than committing: whether a new protocol message should be modelled or ignored is a judgement call, and a bot that answers it would bless a gap silently. Every action is pinned by full commit SHA (a tag is mutable, and the action runs with this repository's token), with Dependabot proposing the bumps so the pinning stays free. Build provenance is attested and deliberately not overtrusted — a compromised runner can sign a build that genuinely happened on it..asctherefore needs a software key in a secret, and that weakening is bounded rather than waved through: the secret is scoped to the approval-gatedmarketplaceenvironment (no job reachable from a bare tag push can see it), the key expires after a year so an unnoticed leak stops mattering on its own, and its user ID says out loud that it is a CI key. That last point is the actual mitigation — if the two signatures were indistinguishable, a leaked CI key would impersonate a person. The two claims are now documented as distinct: the tag signature says a person authorised this release, the artifact signature says this workflow produced these bytes, andSECURITY.mdtells users to check both. Generated byscripts/gen-ci-signing-key.sh, which works in a throwaway keyring and never touches the maintainer's. The CI key is certified by the hardware key, which is what makes the arrangement defensible rather than merely documented: without it a user is asked to trust a fingerprint printed in a file inside the very repository an attacker who could swap the key would control — a tautology, not a trust anchor. With it the chain terminates in hardware, and there is a revocation lever nobody holding the leaked key can undo.scripts/bootstrap-ci.shperforms the whole one-time setup, anddocs/CI_SETUP.mddocuments each step for when it has to be done by hand..github/rulesets/*.json, applied byscripts/apply-rulesets.sh). Bothmainanddeveloprequire a pull request, an up-to-date branch, signed commits, and every CI check. Required approvals are zero, which reads like a hole and is the opposite: GitHub does not let an author approve their own pull request, so on a single-maintainer repository "require 1 approval" with no bypass actors means nothing can ever be merged — not by push, not by PR, not by admin. We established that empirically, by locking the repository and having to unlock it. The gate that remains is the mechanical one, which is also the one that cannot be talked out of. Raise it to 1 when a second maintainer exists; the rulesets carry that instruction inline. No bypass actors, including admins — the previous documentation preserved an admin bypass for a structural blocker that never existed, and a bypass is by construction used at the worst possible moment on the least-reviewed change..gitlab-ci.ymlis removed rather than retained: two pipelines that can each publish is one publisher too many.role="status" aria-live="polite"region declared in the staticshell.html— created lazily it would never announce its first message, which is the classic way to ship a silent live region — plusCC.announcewith duplicate suppression, so a screen-reader user is told when a turn starts, finishes, or is blocked on a permission card. The transcript streams without ever moving focus, so without this the turn simply stalls in silence. Also a:focus-visiblebaseline covering every element whose outline the stylesheet suppresses (the find bar's input had no replacement at all), honoured underforced-colorsrather than overridden. Ten frontend tests pin the structural guarantees; they do not certify conformance, which still requires a keyboard and screen-reader pass by a person.THIRD-PARTY-NOTICES.md,LICENSEandLICENSES/*are packaged underMETA-INF/. The plugin redistributesmarked,DOMPurifyandhighlight.js, and a permissive licence's notice obligation binds on redistribution: a notices file that exists only in the repository does not discharge it for someone who installs the zip. DOMPurify is dualApache-2.0 OR MPL-2.0, so the choice is recorded rather than left implicit.AGENTS.md— the operational runbook for agentic development (commands, gates, boundaries), complementingCLAUDE.md, which stays the architecture.docs/adr/— three decision records: 0001 release process, 0002 threat model, 0003 i18n deferred with the triggers that reopen it.commitlintand a versioned.githooks/commit-msg(enable withgit config core.hooksPath .githooks). The hook self-tests and degrades to advisory if its own toolchain fails, specifically so it can never become a reason to reach for--no-verify.Changed
v4.3.2andv4.4.1were each force-re-cut three times after being pushed. A tag is the identity of a shipped artifact; moving one means two people can hold different trees, different zips and different checksums while both believe they have the same version — which defeats the single thing a signature is for. A mistake found after tagging is now fixed by the next patch version. The already-moved tags are left alone, because re-cutting them to "fix" history would repeat the exact mistake.LoginCoordinatorextracted fromClaudeSession(1965 → 1826 lines). The OAuth sign-in is a subsystem in its own right — the TTY-less--printsession cannot host an interactive login, so it happens outside the session entirely through three ordered paths — and it now owns its own state. Mechanical, no behaviour change, full suite green across it. The two further extractions that were considered (SessionRestorer,RewindCoordinator) were deliberately not made:restoreis 23 lines that touch six pieces of session state, and rewind is one of six identically-shaped control-request delegates. Both would have bought indirection rather than cohesion, and saying so is the point of recording it.package.jsondeclared"license": "ISC"on a GPL-3.0-only repository and lacked"private": true— i.e. it was publishable to npm under the wrong licence. Corrected.<vendor email>attribute is optional and has been dropped fromplugin.xml; vulnerability reports now go through GitHub private security advisories rather than an inbox. That is the better channel on its own merits and not only a privacy measure: the report lands in a private thread attached to the repository, the discussion and fix stay linked to it, and a CVE can be requested from the same advisory — whereas an address in a public file is scraped far more often than it is used by a reporter.claude2.1.222 / SDK 0.3.222;./gradlew checkDriftgreen, protocol surface unchanged.Internal
src/test/frontend/helpers/load.js) now extracts the shell DOM from the realshell.htmlinstead of a hand-copied approximation. The copy had already drifted — it lacked#a11y-status— which is the worst failure mode a harness has: it does not fail loudly, it quietly tests something that is not the product.Static analysis, formatting and coverage — installed, then acted on
config/detekt/baseline.xmlholds exactly those two, both aboutClaudeSession, both explained inside the file — it is a record of a decision, not a drawer. The distinction matters: a 203-entry baseline is a promise to nobody, a 2-entry one is a claim somebody has to defend in review.ClaudeSession.onEventwas a singlewhenover 47 event types — 244 lines, cyclomatic complexity 111 — the one function where every protocol concern in the plugin met.ClaudeEventnow declares seven sealed sub-interfaces (Stream,Conversation,Control,Task,Notice,SessionSignal,HookTelemetry) and dispatch picks the group, then the variant. The grouping is expressed in the type on purpose: a sealed hierarchy keeps the compiler checking exhaustiveness at both levels, so a new protocol event that nobody handles is a compile error rather than a silently dropped frame — which is the propertycheckDriftexists to protect, and was not up for trade against a complexity threshold. The groups are semantic, not cosmetic: they differ in what the host owes the binary (aControlframe must be answered or the binary hangs; aNoticeis fire-and-forget).JcefBridge.MsgandJcefChatPanel.onBridgeMessage(complexity 46) got the same treatment, with the message groups mirroring the bridge's parsers one-for-one.whenchains were dictionaries written as control flow, and are now data:ProtocolParser.parseSystemhad 25 arms of which 21 were the same expression with two names substituted (complexity 29 → aMap), likewise the top-level frame decoder, andEditorContextProvider.langForExtension(26 arms → a lookup table). Adding a protocol subtype is now one line, and the shared fallback wiring is written once instead of 21 times where a mistyped argument would have been invisible.koverVerify), because risk here is not evenly spread:permission/decides whether the agent may read your SSH key,ui/paints a browser. Thresholds sit slightly below what each package measures, so they catch regression instead of inviting test-padding.ui/,context/,process/,actions/andutil/are excluded with the reason stated rather than gated at a token value — gating them at 20% would dress the same fact up as a passing check. Policy, measured numbers and the known gaps are indocs/RELEASE_CHECKLIST.md§Coverage policy.build.gradle.ktsclaimed the figure was "documented indocs/RELEASE_CHECKLIST.md"; that file had never mentioned coverage, and the real number was 53.3%. A number nobody measured, pointing at a rule nobody wrote.no-eval,no-implied-evalandno-new-funcare errors because the page runs under a hash-pinned CSP with no'unsafe-eval': without the gate, code Chromium will silently refuse in a user's IDE can still reachmain. Vendoredmarked/DOMPurify/highlight.jsare excluded — a finding in them is not ours to fix, and fixing it would fork a dependency.Static analysisjob (detekt,spotlessCheck,koverVerify,npm run lint,npm run format:check) is now a required check on both protected branches. Everything above is only worth having if breaking it fails a merge.max-line-lengthandfunction-namingare detekt's, because only detekt can scope an exception to the test tree. Running both meant the stricter-but-blinder tool decided, which is how you end up reformatting single-line NDJSON protocol fixtures to satisfy a tool that cannot be told they are fixtures.Fixed — defects the tooling surfaced
TokenFormat.trimDecimalused the default-locale"%.1f", so on a comma-decimal machine (Spanish, German, French…) a count rendered as1,2kinside otherwise-English UI — and worse, the trailing-.0test stopped matching, so a flat 1000 tokens displayed as1,0kinstead of1k. The same bug inJcefTheme.rgbawas not cosmetic at all: it emittedrgba(217, 119, 87, 0,140)— four components instead of three — so the browser discarded the declaration and the--accent-soft/--link-softwashes (text selection, the code-block Copy hover, the "View diff" hover, blockquote backgrounds) never rendered on those machines. Also fixed in the context-usage percentage and the colour-to-hex helper. All now pinLocale.ROOT.ChainDiffVirtualFileover amock:///URL); the platform persists every open editor tab by URL without filtering by file system, so on the next start each one resolved to nothing. One workspace here had accumulated 13 such entries — all namedClaude · SKILL.md, since the tab title is the file name and a skills repository has oneSKILL.mdper directory — producing 26WARN EditorsSplitters - No file existslines on every single launch.DiffTabCleanupnow closes them onprojectClosingBeforeSave, the one hook that runs before the state is written (projectClosingwould be one step too late), and a wiring test pins theplugin.xmlregistration against the shipped descriptor — the failure mode being silence, not a stack trace.CloseAllDiffsActionmoved to a background update thread. It reads oneCopyOnWriteArraySet's size; keeping it on the EDT put it in the queue behind everything the IDE does at startup.InterruptActiondeliberately stays on the EDT and now says so in the code: it readsContentManagerImpl.mySelection, anArrayListmutated on the EDT with no synchronisation and no threading assertion, so moving it would trade a cosmetic log line for a rareIndexOutOfBoundsException.obj.hasOwnProperty(k)in both DOM-building helpers (breaks if the object carries its ownhasOwnProperty— and those helpers build DOM from host-supplied data), an emptycatchin the Vibe Mode theme restore that silently left the theme half-reverted, and two dead functions (isAgentTool,esc) nobody called.sniffMediaTypeno longer confuses any RIFF container for WEBP. Rewritten around named signatures (complexity 23 → 4), it now checks the four-byte form type that actually identifies the format, not just theRIFFheader that WAV and AVI share.Fixed — a tab-killing regression, and the silences it hid
JcefChatPanel.pendingUntilReadywas declared below theinitblock that uses it. Kotlin runs property initializers andinitblocks in declaration order, so the list was stillnullwhileinitran and the constructor threwNullPointerException— taking the whole tab with it, on new chats and on startup restore alike.lastUsage/lastUsageAthad the identical defect and stayed silent, because a nullable reference and a primitive read asnull/0instead of throwing: the loud version of this bug was the lucky one. The compiler does not catch it — it flags a direct reference in an initializer, but here the read happens inside a function called frominit, which it cannot see through — soInitOrderContractTestscans the sources and fails the build on any class-body property declared after its owninit.start()only dispatches, soclaudeboots while JCEF creates its browser), and a boot screen holds the tab until the process is up. Three states, not two:running,starting, and neither — that last one is a launch that failed (missing binary, declined trust prompt, refused remote-mount project) and it must bring the screen down, or the tab stays covered forever with no way to reach the notification explaining why. The screen is declared visible in the static shell, since at page load the process genuinely is not up yet.javax.swing.Timer's initial delay equals its interval, so the first poll came a fullQUOTA_POLL_MSafter the panel attached — and that tick landed while the binary was still launching and returned early on the not-running guard, costing a second interval. Process-ready, tab-open and both turn edges now poll directly, and the timer retires at the end of a turn: context and cost cannot move while a session sits idle, so polling forever was a round-trip through the binary, per tab, for two numbers that provably had not changed.get_usagereply refreshed the dashboard bars but not the composer's dots, so the same number appeared immediately in one surface and "a while later" in the other, whenever some unrelated state change happened to re-push. Both are pushed together now. Opening the dashboard also refreshes them, whichrequestUsage's own contract had claimed and the code had never done.0instead of omitted. An item that only appears once it is non-zero is indistinguishable from one that failed to load — which is exactly how a fresh tab read: a lone "Idle" and no figures. Cost stays gated, because a currency amount of zero is noise rather than an ambiguity to resolve.<tool_use_error>wrapper reached the transcript verbatim.claude2.1.222 wraps a failed tool result'scontentin that tag pair and carries the same message unwrapped in a sibling field — framing for the model, not text for a human. Rendered as-is it put raw markup in a native GUI, the "never mirror raw CLI output" antipattern this plugin exists to avoid. Stripped only when it encloses the whole payload, so output that legitimately mentions the tag survives;is_erroralready conveys the failure structurally, and is what reddens the card.ToolSearchwas missing from theSensitiveGuardtrust allowlist, along withAskUserQuestion,McpandFileRead/FileEdit/FileWrite.ToolSearchis the one that mattered: it loads the schema of every deferred tool, so on a session that defers them, the call that unlocks all the others was the one landing in the third-party branch. Entries are only ever added to that list — it is a trust allowlist, not an inventory, and a missing first-party name is precisely the 4.4.0 hard-DENY incident. Found by diffing it against a live session's real tool inventory rather than against the SDK's type names, which are not the runtime registry (the SDK calls themFileRead/FileEdit/FileWrite; the tools areRead/Edit/Write).https://andjb://openand dropped everything else without a sound — so[BACKLOG](docs/BACKLOG.md)was inert while bare paths written in prose worked, making the more deliberate link the one that failed. Both routes now go through a single authorising gate (LinkResolver.isOpenable). The scheme test requires two or more characters before the colon, so a Windows drive (C:\src\main.kt) stays a path rather than being mistaken for a URI scheme..copiedclass had been applied by the JS since 4.0.4 and had no CSS rule at all; it now has one.