Release 5.0.0 - #34
Closed
serialexperimentslainnnn wants to merge 38 commits into
Closed
serialexperimentslainnnn wants to merge 38 commits into
serialexperimentslainnnn wants to merge 38 commits into
Conversation
Audited the repository against git-workflow-standards. Adds the governance the standard requires and that this repo lacked, minus anything CI-bound (no CI gate exists yet - a local lab is planned). - commitlint.config.mjs + a VERSIONED .githooks/commit-msg gate, so the rule travels with the repo instead of living in one laptop. The hook self-tests against a known-good message first and degrades to a warning if the toolchain is broken: a hook that fails closed on its own bugs gets bypassed with --no-verify, and then protects nothing. - .gitattributes: this plugin ships and is tested on Windows, and publishes signed zips. Without a declared policy a clone with autocrlf=true rewrites every file and invalidates artifact signatures. Vendored frontend bundles marked -diff so a bump is reviewable. - docs/adr/0001: records the two justified deviations (GitFlow for installable versioned software; GPG-on-YubiKey for real revocation and because the same key signs the distributed artifacts) and corrects a real violation - v4.3.2 and v4.4.1 were each re-cut and force-pushed three times. Published tags are now immutable; a post-tag mistake is fixed by the next patch version. CHANGELOG generation is deliberately still open: the tooling silently skips non-Conventional commits, so the message gate has to land first. Refs: git-workflow-standards §3.2, §4.1, §5.2, §5.3
The published zip redistributes third-party code - marked 12.0.0 (MIT), DOMPurify 3.0.11, highlight.js 11.9.0 (BSD-3-Clause) vendored into the plugin jar, and kotlinx.serialization 1.7.3 (Apache-2.0) as its own jars. MIT, BSD-3-Clause and Apache-2.0 all require the copyright notice and licence text to be preserved ON REDISTRIBUTION, and a file sitting in the Git repository does not accompany the binary a user installs from the Marketplace. This was unmet. processResources now packages THIRD-PARTY-NOTICES.md, LICENSE and the three licence texts into META-INF/ of the plugin jar - verified present in the built artifact, not just in the source tree. Generated at build time from one root-level source of truth rather than a checked-in copy, so the notices cannot drift from the files they describe. Every licence was verified by reading the LICENSE of the exact shipped version, never the manifest or a badge. That found a real one: DOMPurify 3.0.11 is dual "Apache-2.0 OR MPL-2.0". An OR expression is a choice the redistributor must make and record; leaving it unstated is an unmade decision. Apache-2.0 is selected, with the reasoning recorded in the notices file. kotlinx.serialization ships no META-INF/LICENSE in its jars, so its licence was read from the project source instead of inferred. Also corrects the documented JDK path across CLAUDE.md, README, CONTRIBUTING and docs/: ~/.local/jdks/jdk-21.0.11+10 no longer exists on this machine (it is now ~/.jdks/jbr-21.0.11), and gradlew failed with "JAVA_HOME is set to an invalid directory" rather than anything a build log filter would obviously catch. Refs: opensource-licensing-standards §3.5, §3.6, §5.4, §8
Audited the JCEF web UI against WCAG 2.2 AA. This is a conformance requirement, not a nice-to-have: the plugin is distributed to consumers in the EU, where Directive (EU) 2019/882 has applied since 28 June 2025 (in Spain, Ley 11/2023). 4.1.3 Status Messages (AA) - the transcript streams without ever moving focus, so a screen-reader user got NO signal that Claude started, finished, or was blocked on them: the turn simply stopped, silently. Adds a polite live region declared in the static shell (it has to exist in the DOM before text is written into it, or the first change is never announced) plus CC.announce. Only turn EDGES are announced, never per token - a region updated on every delta talks over itself and gets switched off, which is worse than silence. The most important case is a pending permission card: it appears without taking focus, so it was previously undetectable. 2.4.7 Focus Visible / 1.4.11 Non-text Contrast (AA) - the stylesheet suppressed the default outline in five places. Four had some replacement; .find-input had NONE, so focus there was invisible. Adds a :focus-visible baseline covering the controls built as <span role="button"> (the code-block Copy control, menu items, tool cards), which get nothing for free because they are not native buttons, and honours forced-colors instead of overriding the system's focus colour. prefers-reduced-motion was already correct (a universal reset covering all eight keyframe animations) and is now pinned by a test so it cannot decay into a hand-picked subset as animations are added. Refactor: the frontend test harness hand-copied the shell DOM and had already drifted - shell.html gained the live region, the harness did not, so tests exercised a DOM the product does not have. It now extracts the body from the real shell.html. A mount point added to the shell reaches tests automatically; one removed breaks the tests that relied on it. Scope stated honestly in the test file: automated checks catch ~40-57% of real barriers and none of the semantic judgements. These pin the structural guarantees that regress silently. They do not certify conformance - that still needs keyboard-only and screen-reader passes by a person. Refs: accessibility-standards §3.2, §3.4, §3.8, §4.1
`@anthropic-ai/claude-agent-sdk` sat in `dependencies` while CLAUDE.md has always said it is "protocol reference only, not distributed" — and it is: the published artifact is jars and inlined web assets, with zero `node_modules` entries (verifiable with `unzip -l`). The mismatch produced seven permanent npm-audit findings (three high) against code no user ever receives, which is worse than useless: a backlog of alarms that cannot be acted on trains you to ignore the one that matters. Moving it to `devDependencies` makes the declaration match reality, so `npm audit --omit=dev` — the distributed scope — now reports zero. The full tree still reports the same seven; they are build-time tooling and are handled as maintenance. SECURITY.md states that triage boundary explicitly, with the command to verify it rather than trust it. `checkDrift` was the constraint on this change: it reads the SDK out of `node_modules` and runs `npm update`. Both survive, because `npm install` installs devDependencies — only `--omit=dev` would break it, and nothing uses it. Verified green, and it reported real drift while doing so, so the baseline moves with it: SDK 0.3.220 → 0.3.222, binary 2.1.220 → 2.1.222, protocol surface unchanged. Also corrects the package manifest itself: it declared `"license": "ISC"` on a GPL-3.0-only repository and lacked `"private": true`, i.e. it was publishable to npm with the wrong licence. And the supported-versions table still said 2.x.
The control was strong and undocumented: nothing stated what it is defending against, which makes it unreviewable. Coverage cannot be argued, and every discussion of a proposed bypass restarts from first principles. ADR 0002 states the trust model — user trusted, `claude` binary trusted as software but untrusted as a channel, everything it relays untrusted — and runs STRIDE over the three surfaces that actually exist: the binary as a child process, third-party MCP servers, and model-returned content. The third is the one worth having written down. We do not attempt to detect prompt injection; content-level detection is unsolved and a control built on it would be a liability. The document says so explicitly and records where the defence actually sits instead: the guard judges the tool call and never the reasoning behind it, so an injection that succeeds completely still has to ask to read the key and still gets the same answer. Also records why MCP servers are denied rather than asked about, and why caller trust is an allowlist — names are attacker-supplied, and 4.4.0 is the cautionary tale of that list going stale (it failed safe, which is why the allowlist was the right shape, but it still failed). Non-goals are as load-bearing as the goals, so they are listed: we do not defend against the user, against a compromised host, or against every obfuscation. That last one draws the line triage keeps blurring — failing to *recognise* a path is a pattern gap; enforcement of a match, once made, is absolute. Claims verified against the code before writing, not from memory: `binaryPermissionMode` rewriting acceptEdits/bypassPermissions to `default`, and the CSP's `default-src 'none'` / `connect-src 'none'` / hash-pinned scripts. SECURITY.md links here so a reporter can tell a finding from a stated position before writing the email.
Until now the entire quality bar for ~13k lines of Kotlin and ~3.6k lines of shipped JavaScript rested on review. That is the thing the standards say to mechanise: if formatting is discussed in a review, a formatter is missing. Adds and configures: - detekt 1.23.8 (config/detekt/detekt.yml). Every non-default setting carries its reasoning AT the setting, not in a message nobody will find again. Three rules were re-calibrated rather than suppressed: TooManyFunctions counts only the public surface (the default penalised extracting helpers, which is the exact fix the complexity rules demand); LongParameterList ignores defaulted parameters (one absent from the call site is not part of the problem the rule exists for); MagicNumber skips the colour palette (a 0xRRGGBB inside a property named DIFF_ADDED_BG is that name's definition). - Spotless 8.9.0 / ktlint 1.8.0. max-line-length and function-naming are disabled here and left to detekt: both tools enforced them, only detekt can scope an exception to the test tree, and running both meant the stricter-but-blinder one decided. - ESLint + Prettier over the JCEF frontend, which ships inside the plugin jar and had never passed through any tool. no-eval / no-implied-eval / no-new-func are errors because the page runs under a hash-pinned CSP with no 'unsafe-eval': without this gate, code Chromium will silently refuse in a user's IDE can still reach main. Vendored marked/DOMPurify/highlight.js are excluded — a finding there is not ours to fix. - koverVerify, gated per package. Risk is not evenly spread: permission/ decides whether the agent may read an SSH key, ui/ paints a browser. Thresholds sit slightly BELOW what each package measures today, so they catch regression rather than invite test-padding. ui/, context/, process/, actions/ and util/ are excluded with the reason stated rather than gated at a token value. config/detekt/baseline.xml holds TWO entries, both about ClaudeSession, both explained inside the file. It started this work at 203; the other 201 were fixed, not frozen. A `Static analysis` job runs all of it and is a required check on main and develop, so none of the above depends on anyone remembering to run it.
Three strands that could not be separated into their own commits: the formatter rewrote the same files the refactor touched, hunk for hunk, so splitting would have meant manual surgery across ~100 files for commits that would not build in isolation. Recorded here rather than pretended otherwise — it also means there is no pure formatting commit to add to .git-blame-ignore-revs. REFACTOR — the dispatch tables, split in two levels ClaudeSession.onEvent was one `when` over 47 event types: 244 lines, cyclomatic complexity 111, the single function where every protocol concern in the plugin met. ClaudeEvent now declares seven sealed sub-interfaces and dispatch picks the group, then the variant. The grouping lives in the TYPE on purpose: a sealed hierarchy keeps the compiler checking exhaustiveness at BOTH levels, so a new protocol event nobody handles is a compile error rather than a silently dropped frame — the property checkDrift exists to protect, and not up for trade against a complexity threshold. The groups are semantic: they differ in what the host OWES the binary (a Control frame must be answered or the binary hangs; a Notice is fire-and-forget). JcefBridge.Msg and JcefChatPanel.onBridgeMessage got the same treatment. Several `when` chains were dictionaries written as control flow and are now data (ProtocolParser.parseSystem: 25 arms, 21 identical bar two names). FIXED — defects the tooling and live probing surfaced - Locale-dependent formatting. TokenFormat rendered "1,2k" on a comma-decimal machine, and worse, the trailing-.0 test then stopped matching so a flat 1000 showed as "1,0k". The same bug in JcefTheme.rgba emitted four rgba components instead of three, so the browser DISCARDED the declaration and the accent/link washes never rendered. Now pinned to Locale.ROOT. - Diff tabs were persisted into the workspace and could never be restored: they are mock:// in-memory previews, and the platform persists open tabs by URL without filtering by file system. One workspace had 13 such entries, all named "Claude · SKILL.md", producing 26 warnings on every launch. DiffTabCleanup now drops them on projectClosingBeforeSave — the one hook that runs BEFORE the state is written. - Every chat animation was dead. The accessibility pass had added a prefers-reduced-motion block, and inside JCEF that query reports `reduce` regardless of the desktop (measured: it matched while GNOME had animations enabled), so it flattened everything. The tell was Vibe Mode's rainbow, which survived because it is a setInterval, not a keyframe. Motion is now reduced only when the user asks, via a plugin setting, off by default. - Requests needing a live binary were dropped. The panel is built BEFORE session.start() runs, so requestMcp/requestVersion/requestUsage all found isRunning() == false and did nothing; each looked like its own bug. They now queue until ready. - sniffMediaType treated any RIFF container as WEBP; it now checks the form type that actually identifies the format (WAV and AVI share the RIFF header). FEAT — plan limits, in both surfaces get_usage is a control request the plugin has known about since 4.0.1 and never sent; it returns every rate-limit window at once plus the extra-credit balance. The dashboard shows one bar per window with its reset countdown; the composer readout shows a dot per window. One severity scale for both: blue under 65%, amber under 85%, red above — status hues, never the product accent, because a quota bar that is coral at 10% and coral at 95% has said nothing. Crossing 65% and 85% announces once per threshold per window (a warning that repeats every refresh is one users learn to ignore), and 85% also raises an IDE notification. ClaudeSession.rateLimits is now per-window: the single field could only ever hold whichever event arrived last, so a user on a weekly limit saw their five-hour bar and concluded they had room. Also adds CC.diagnostics, which reports what the embedded browser actually resolves to idea.log. The UI is a browser nobody can open devtools on, and without it a CSS rule that silently does not apply is indistinguishable from a backend that never sent the state — which cost three wrong diagnoses before it existed. Suite: 682 -> 689 JVM, 54 -> 67 frontend.
Documents what the release actually did, including the parts that did not go to plan — a release note that only lists wins is one nobody can act on. - CHANGELOG gains two sections: the static-analysis work (203 findings fixed rather than frozen, down to 2 recorded exceptions) and the defects it exposed. - docs/RELEASE_CHECKLIST.md gains a Coverage policy with the MEASURED per-package numbers, what is excluded and why, and the Kover 0.9.2 limitation behind the current floor+aggregate shape. It also corrects a claim that never had a basis: build.gradle.kts cited a ">=90% target documented in RELEASE_CHECKLIST.md", that document had never mentioned coverage, and the real figure was 53%. - docs/BACKLOG.md is new. Every entry was PROBED against the real binary rather than inferred from the SDK types, including the two that are deliberately not worth doing, so nobody investigates them a third time. - AGENTS.md, CI_SETUP.md, BRANCHING.md and the ADRs updated for the GitHub Actions pipeline that replaced the deleted GitLab one. Known gaps, stated rather than left to be discovered: the risk-ranked human review is not done; the reduced-motion decision generalises from a measurement on one machine; and Fable's weekly window cannot be shown at all, because get_usage returns every per-model window as null (measured, not assumed).
Squash and rebase merging are now disabled at the repository level, leaving
merge commit as the only option. BRANCHING.md said the opposite for `develop`
("squash or merge"), which was worse than saying nothing: it invited the exact
click that destroys signatures.
Both disabled methods REWRITE commits — new SHAs, new committer information —
which invalidates the author's hardware-backed signatures and replaces them with
GitHub's web-flow key. The "signed commits required" rule would still pass, and
that is the trap: the commits are signed, just no longer by the author, which is
the whole property the rule exists to provide.
The cost is recorded too: the merge node itself is GitHub-signed, because the
alternative is a local merge and a direct push, which the pull-request
requirement blocks. Relaxing that to save one commit's provenance would be a far
worse trade. Release provenance therefore rests on the signed tag, not on the
merge node.
Also corrects the status-check lists, which had not been updated when the
`Static analysis` gate was added.
CodeQL flagged two high-severity alerts on the same line of the frontend test harness: "incomplete multi-character sanitization" and "bad HTML filtering regexp". Both were correct: the pattern used to strip script tags does not match a closing tag written with a trailing space, so something that reads like a sanitiser silently is not one. The exposure was low. This reads OUR OWN shell.html off disk inside a test, and the reason script tags are stripped is load control — the harness injects the app modules itself, in a controlled order — rather than sanitisation. That is a reason to rank the finding, not a reason to keep it: a regex that looks like HTML sanitisation is a pattern someone copies to a place where it does matter. Fixed at the root rather than by widening the pattern. jsdom is already a dependency of this harness, so the file is parsed and the script elements are removed as nodes. All 67 frontend tests still pass, which is the check that the harness still builds the same DOM as the product. Recorded for the risk review: the two CodeQL contexts the rulesets require (java-kotlin, javascript-typescript) both PASSED — they report "the analysis ran". The aggregate CodeQL check, which reports "something was found", is required by neither branch. A high-severity finding can therefore merge with no gate objecting.
A regression introduced by this branch made the plugin unusable, and fixing it surfaced a run of smaller defects that all shared one shape: something failed or was still loading, and the UI said nothing at all. THE REGRESSION JcefChatPanel.pendingUntilReady was declared below the `init` block that uses it. Kotlin runs property initializers and init blocks in declaration order, so the list was still null while init ran and the constructor threw. It took the whole tab with it: no chat could be opened or restored. lastUsage/lastUsageAt had the same defect and stayed silent, being nullable and primitive. InitOrderContractTest now fails the build on any property declared after its class's init. WAITING IS NOW VISIBLE The binary is launched BEFORE the tab is built (start() only dispatches, so `claude` boots while JCEF creates its browser), and a boot screen holds the tab until the process is up. Three states, not two: running, starting, and neither — that last one is a failed launch and must clear the screen, or a missing binary leaves the tab covered forever. Context and cost no longer wait on a clock. The poll timer's initial delay equals its interval, so the first reading was a minute late; worse, that tick landed while the binary was still starting and returned early, costing a second minute. Ready, tab-open and both turn edges now poll directly, and the timer retires at the end of a turn instead of round-tripping forever on an idle session. Plan limits were pushed to the dashboard but not to the composer, so the same number appeared instantly in one place and "later" in the other. Reasoning tokens, context and output now settle at 0 rather than being omitted until non-zero: an absent figure is indistinguishable from one that failed to load. FAILURES THAT WERE HARD TO READ - The CLI wraps failed tool results in <tool_use_error>...</tool_use_error> (verified in claude 2.1.222, which carries the same text unwrapped in a sibling field). Rendered verbatim it put raw markup in a native GUI. Stripped only when it encloses the whole payload; is_error already conveys the failure. - A failed tool card stayed collapsed, so the entire message was "the header is red", and its output scrolled sideways rather than wrapping, hiding the half that says what to do instead. It now opens itself once, and error text wraps. - ToolSearch was missing from SensitiveGuard.AGENT_TOOLS, so on a session that defers tools, the call that loads every other tool's schema landed in the untrusted branch. Added, with AskUserQuestion, Mcp and FileRead/Edit/Write. Entries are only ever added here: this is a trust allowlist, not an inventory, and a missing first-party name is the 4.4.0 hard-DENY incident. - A markdown link whose href is a path did nothing when clicked: the host handled https:// and jb://open and dropped the rest without a sound. Both paths now go through one gate (LinkResolver.isOpenable). The scheme test requires two or more characters so a Windows drive letter stays a path. - Copy on a message copied and gave no feedback, which reads as a broken button; the shared flash helper is now exported rather than reimplemented, and the .copied class finally has a CSS rule — it had never had one. Suite: 692 -> 694 JVM, 67 -> 84 frontend. verifyPlugin Compatible on IC-251, IC-252, IU-253, IU-261 and IU-262.
Two CI failures, one of them ours. OURS — the `Static analysis` job died with an IOException from DriftLiveCheck. Kover instruments and aggregates EVERY `Test` task in the project, and `checkDrift` is registered as one, so it had silently become a dependency of `koverVerify`. That task is on-demand by design: it downloads the latest SDK and probes a `claude` binary installed on the machine. A runner has no such binary. The task's own documentation already said "NOT wired into `check`" — it just was not true of the coverage graph, and nothing checked that it was. It passed locally for the one reason that makes this class of bug expensive: the maintainer's machine has the binary, so "on-demand" and "wired in" looked identical until CI ran it. Verified with `koverVerify --dry-run`, which listed `:checkDrift` in the graph before and does not after. NOT OURS — the `JVM tests` job failed resolving com.jetbrains.intellij.platform:test-framework with a 502 Bad Gateway from cache-redirector.jetbrains.com. Nothing to fix; it needs a re-run. Also bumps the IntelliJ Platform Gradle Plugin 2.16.0 -> 2.18.1, which the build had been warning about on every run. Stated plainly: the plugin bump is NOT verified locally. A 2.16 -> 2.18 jump touches the whole build, and CI is the verification here rather than a claim made in advance.
Reverts the 2.16.0 -> 2.18.1 half of the previous commit. The `checkDrift`
coverage-graph fix in that commit stands; only the plugin version goes back.
On 2.18.1 the headless suite never runs. `ChatSessionManagerHeadlessTest` hangs
before its first assertion, and a thread dump puts the EDT here:
BasePlatformTestCase.setUp
-> LightPlatformTestCase.doSetup
-> IndexingTestUtil.waitUntilIndexesAreReady (267s and counting)
That is the platform's own fixture waiting for indexing that never completes.
Our code is not on the stack at all — the bump changes which platform
test-framework is resolved, so this is a fixture-level regression rather than
something fixable from this side.
Measured both ways rather than inferred: the same test on 2.16.0 finishes in 19
seconds, and the full gate is green at 694 tests.
The version is now pinned with the reason written AT the setting, because the
build prints an "outdated" warning on every single run and the next person to
see it will otherwise do exactly what I did. Re-attempt it as its own change
with the headless suite as the acceptance test, not as a drive-by on a release
branch.
I shipped that bump unverified and said so in the message; this is what it cost.
The one-line lesson is the boring one: a build-tooling bump is a change like any
other and does not get to skip the suite because it looks like configuration.
…5-standards Feature/release 5 standards
Merging `develop` into `main` now publishes. The tag is DERIVED from the version in build.gradle.kts rather than supplied alongside it, so the two can no longer disagree — the mismatch the old flow guarded against with a comparison simply cannot occur. A merge that does not bump the version publishes nothing: the workflow finds the tag present, logs a notice and stops. It does not fail. `main` legitimately receives merges that are not releases, and a red run on each of those is an alarm people learn to ignore. Published tags stay immutable, which is the correction ADR 0001 records after v4.3.2 and v4.4.1 were each force-re-cut three times. Pushing a tag by hand still works, as the escape hatch for re-cutting after a failed publish without an empty commit on main. A tag this workflow creates does not re-trigger it, so there is no loop. The tag is cut AFTER the approval and the publish, not in the guard. Created earlier it would name a version that was never published when a build fails or an approval is declined — and since tags here are immutable, that would block the next attempt. Cutting it last makes it mean "this was published", which is the only claim it can honestly make once the version, not the tag, is the input. WHAT THIS COSTS, STATED RATHER THAN GLOSSED The tag is signed by the CI key, not the maintainer's YubiKey — which cannot sign inside a runner, and whose non-exportability is exactly what makes it worth trusting. The chain still ends in hardware because the CI key is certified by it. But no signature on a release now asserts that a person authorised it. That claim moves entirely to the two gates around publication: `main` accepts only reviewed pull requests, and publishing requires an approval from a named reviewer on a protected environment. BRANCHING.md and SECURITY.md said the opposite and are corrected — a verification instruction that overstates what it proves is worse than none, because someone acts on it. ADR 0001 still describes the tag-triggered flow and needs superseding.
STRICTER — the policy is now a gate rather than a promise CLAUDE.md says "never ship a deprecated or scheduled-for-removal API — treat it as a blocker, not a warning". Nothing enforced it. The verifier's default failure level is COMPATIBILITY_PROBLEMS + INTERNAL_API_USAGES + OVERRIDE_ONLY_API_USAGES, so a deprecated usage was reported and the job went green anyway. A rule that lives only in prose is not a rule; DEPRECATED_API_USAGES is what makes the sentence true. Verified green today, so it lands with no debt to forgive. EXPERIMENTAL_API_USAGES is deliberately excluded, and that is a decision rather than an oversight: DiffTabCleanup uses ProjectCloseListener.projectClosingBeforeSave knowingly, because it is the only hook that runs before the workspace is written. An experimental API is acceptable with a reason. A deprecated one is not, because it has an announced removal and this plugin must keep working across 251 → 262. LEANER — four measured duplications, none of them a weakened gate - Every commit ran the pipeline TWICE. `push` on topic branches and `pull_request` both fire, and the concurrency group keyed on `github.ref` differs between them (refs/heads/x vs refs/pull/N/merge), so neither cancelled the other. Keyed on the commit now: same SHA, same group, duplicate cancelled. - Superseded runs were only cancelled for pull requests. Three pushes in a row left three full pipelines racing, each spending ten minutes downloading IDEs for a commit that had already been replaced. - The JVM suite ran twice. `koverVerify` depends on `:test`, so putting it in the `Static analysis` job re-ran the entire suite on a second runner with a cold cache. Coverage is a property OF a test run and now shares its job. - The plugin was built twice. `verifyPlugin` already produces the distributable; `Build plugin` built its own. That was not only wasteful but subtly wrong — the bytes being asserted were never the bytes that were verified. It now downloads the verified artifact. Job DISPLAY NAMES are unchanged, deliberately: a ruleset references a required check by its name, so renaming one does not fail the gate — it silently stops applying it. No rulesets need reapplying. Also gives drift.yml the concurrency group it was missing (queue, do not cancel: a half-written drift report is worse than a late one). Caught while writing this: the SHA I pinned actions/download-artifact to was invented. Verified against the API and corrected. Pinning by SHA protects nothing if the SHA is made up.
Nine open Dependabot PRs, each firing the full pipeline — verifier included, ten
minutes and 1.25 GB of IDE downloads apiece — for dependency bumps that are
reviewed in seconds.
Five of them were SECURITY updates (undici, ip-address, fast-uri, hono, postcss:
exactly the `npm audit` findings). Two things about that stream were not
understood when this file was written, and both are documented at the setting now:
- security updates ignore `open-pull-requests-limit` entirely, which is how five
arrived under a limit of three;
- the existing group did not cover them, because `applies-to` defaults to
version-updates.
So each ecosystem now has an explicit `applies-to: security-updates` group. These
are transitive devDependencies that are never distributed — `npm audit --omit=dev`
reports 0, and the artifact contains zero node_modules entries — so reviewing them
one at a time bought nothing.
github-actions and gradle had no grouping at all. Actions are pinned by full
commit SHA, so a bump is a one-line change per action; grouping them costs no
review fidelity. Gradle groups minor and patch only: a MAJOR keeps its own PR
deliberately, because that is the ecosystem where a bump can hang the headless
suite — the 2.16 -> 2.18 platform-plugin attempt did exactly that, and it deserved
its own run and its own decision.
Syntax verified against GitHub's Dependabot options reference rather than written
from memory, after inventing an action SHA earlier today.
I edited build.gradle.kts and ran verifyPlugin and `help` against it, but not spotlessCheck — so `Static analysis` failed on spotlessKotlinGradleCheck for a purely mechanical reason. The formatter is a gate like any other and running a subset of the gate is the same as not running it.
…ipelines ci: release from main, enforce zero deprecations, remove duplicate work
Bumps [org.junit:junit-bom](https://github.com/junit-team/junit-framework) from 5.11.4 to 6.1.2. - [Release notes](https://github.com/junit-team/junit-framework/releases) - [Commits](junit-team/junit-framework@r5.11.4...r6.1.2) --- updated-dependencies: - dependency-name: org.junit:junit-bom dependency-version: 6.1.2 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps the actions group with 1 update: [actions/download-artifact](https://github.com/actions/download-artifact). Updates `actions/download-artifact` from 7.0.0 to 8.0.1 - [Release notes](https://github.com/actions/download-artifact/releases) - [Commits](actions/download-artifact@37930b1...3e5f45b) --- updated-dependencies: - dependency-name: actions/download-artifact dependency-version: 8.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: actions ... Signed-off-by: dependabot[bot] <support@github.com>
…b_actions/actions-1fa769d870 build(deps): bump actions/download-artifact from 7.0.0 to 8.0.1 in the actions group
Three changes, all aimed at the same thing: the pipeline was spending its time on work nobody needed. DEVELOP NO LONGER REQUIRES BRANCHES TO BE UP TO DATE `strict_required_status_checks_policy` meant every merge into develop invalidated every other open PR, forcing each to update and re-run the entire suite. With two grouped Dependabot PRs that is annoying; with five it makes a release impractical, and the cost is paid on exactly the changes that least deserve scrutiny. What the setting guards against is real — two PRs that are green apart can break together — so it is not simply dropped. It is dropped where a second net exists: CI runs on every push to develop, so a semantic conflict is caught there, on develop, before anything reaches main. `main` KEEPS the strict policy: one merge per release, and it is the merge that publishes. THE VERIFIER RUNS WHERE THE ANSWER MATTERS ~10 minutes and 1.25 GB of IDE downloads, previously on every push to every topic branch. Now on pull requests and on the protected branches. It remains a required check, so nothing merges without it. The loss is early detection mid-branch, which is a genuine cost rather than free savings. TOPIC BRANCHES WRITE THEIR OWN CACHE Read-only was blanket-applied to everything but develop and main, so a topic branch restored the shared cache and saved nothing — every push re-downloaded what the previous one had already fetched. Read-only now applies to pull requests from forks only. Verified against GitHub's cache documentation rather than assumed: "Workflow runs cannot restore caches created for child branches or sibling branches", and a cache created on a pull request is written to the merge ref and "can only be restored by re-runs of the pull request". A topic branch cannot reach what develop reads back; the isolation is the platform's, not ours. NB the ruleset change needs ./scripts/apply-rulesets.sh to take effect.
~10 minutes and 1.25 GB of IDE downloads per run, previously paid on every iteration of a branch nobody was about to merge. It now runs only on pushes to develop and main — not on topic branches, not on pull requests. This REQUIRED dropping "Plugin verifier" and "Build plugin" from the required checks on both rulesets, and that is not a detail: a required check whose job never runs is never reported, so the pull request would wait forever. The two changes have to move together or the branch becomes unmergeable. What still covers a release: release.yml re-runs the FULL gate — verifier included — on the exact tree being published, behind the environment approval, so nothing reaches the Marketplace unverified. What is genuinely lost is catching a binary incompatibility at pull-request time rather than after the merge into develop. A real regression in feedback latency, accepted deliberately. Dependabot also moves to monthly, and majors are no longer bot-proposed in any ecosystem: they are where a bump actually breaks something and where CI is not enough on its own — the platform-plugin 2.16 -> 2.18 attempt passed CI and hung the headless suite locally, inside the platform's own test fixture. Security updates are unaffected by either change.
Corrects the previous commit, which moved the verifier in the wrong direction. It ran it only on PUSHES to develop and main and dropped it from main's required checks — so a pull request from develop into main, the merge that publishes, would not have run it at all. That is precisely the door that has to be guarded. The policy now matches the intent: topic branches and pull requests into develop run the fast checks (compile, our own JVM and frontend suites, static analysis, dependency audit) and iterate quickly. Any pull request targeting main also runs the plugin verifier and the artifact assertions, both required again on main alongside the two CodeQL analyses. The verifier is the only thing that catches a BINARY incompatibility across the declared 251 -> 262 range — compiling against 252 proves nothing about 262, which is exactly how the 4.4.1 /login regression shipped. Skipping it on a branch is a latency trade; skipping it on the way to a release would not be.
Branches now run JVM tests and frontend tests and nothing else. Static analysis
and the dependency audit join the plugin verifier behind the develop -> main
door, where the exhaustive gate belongs.
Required checks match, because they have to: a check whose job does not run is
never reported and would block the pull request forever. develop requires the two
suites; main requires all eight.
The costs, named rather than discovered later:
- A formatting, detekt, ESLint or coverage failure now lands ON develop and is
fixed by a follow-up commit, instead of being caught in the pull request. That
happened today with spotlessKotlinGradleCheck, and the PR is what caught it.
- The dependency audit no longer runs on the Dependabot PR that proposes a bump.
It runs once that bump is on develop, and again before it can reach main, so
nothing ships un-audited — the finding just arrives one merge later.
What this buys is the thing that was actually hurting: a branch iterates in about
three minutes instead of thirteen, and nothing is promoted to main without the
full gate, plus release.yml re-running all of it on the exact published tree.
…e/org.junit-junit-bom-6.1.2 build(deps): bump org.junit:junit-bom from 5.11.4 to 6.1.2
…ipelines Feature/update pipelines
A branch with an open pull request already fires `pull_request` on every push to
it (the `synchronize` event), so keeping a `push` trigger meant two complete
pipelines per commit for identical information. This removes the duplication at
its source instead of relying on the concurrency group to cancel one in time.
The result is the intended shape: a PR into develop runs the two required suites
once, a PR from develop into main runs all eight.
Two things this gives up, recorded because each removes something the current
setup was leaning on:
- There is no longer a CI run on the push a merge into develop creates. That
run was the stated justification for dropping the up-to-date requirement on
develop: two pull requests that are green apart can break together, and
develop's own run was what would have caught it. It is now caught at the pull
request into main, where the full gate runs — later, but still before
anything is published.
- A branch with no open pull request gets no checks at all, and a pull request
from a fork is the only path that would ever exercise them for an outside
contributor.
release.yml is untouched: it carries its own `push: branches: [main]` trigger and
still fires on the merge that publishes.
Points the five toolchain jobs at ghcr.io/serialexperimentslainnnn/cc-ci and drops
the setup-java and setup-node steps, which provisioned inside a container that
already has both. `Build plugin` keeps no container: it only unzips an artifact.
GRADLE_USER_HOME is set per job and MUST match the value in the Dockerfile. If the
two diverge nothing fails — the run simply re-downloads everything the image
already holds, and the image appears to have bought nothing. That silence is the
reason it is stated at the setting rather than assumed.
NOT VERIFIED against a real image: at the time of writing it has not been built or
pushed. Two things have to be true before this can merge, and both fail in ways
that look like something else:
- the package must be public (or linked to this repo), or every job dies on a
401 that reads like a wrong image name;
- the warmed caches must actually be in the image — `docker run --rm IMAGE
sh -c 'ls /opt/gradle-home/caches'` answers it in seconds.
Also still open: whether gradle/actions/setup-gradle should stay. It restores its
own cache over GRADLE_USER_HOME, so it now layers on top of the baked one. That
may be a useful increment or redundant work; it needs measuring, not guessing.
The package stays private. Each container job authenticates with the GITHUB_TOKEN the run already has, so there is no new secret to create, store or rotate, and the credential expires with the job. `packages: read` is granted per job rather than at the top level, keeping the default token read-only on everything else. Without it the pull fails with a 401 that reads like a wrong image name rather than a permission problem — which is exactly the kind of error that gets debugged in the wrong place. One prerequisite this does NOT remove: the package must be linked to this repository, or the token has no grant on it. That is done once, from the package settings, and it is what makes "same account" mean "same permissions" here.
The cleanup step wipes /warmup, and `npm ci` had installed node_modules inside it — so the image built the frontend dependencies and deleted them moments later. The warm-up looked like it worked and bought nothing: CI would re-download the whole tree on every run, silently, because nothing fails when a cache is missing. Setting npm_config_cache moves the reusable part to /opt/npm-cache, which the cleanup does not touch. node_modules stays disposable, and that is correct independently of this bug: it must match the package-lock.json of the commit CI checks out, not the one that happened to be current when the image was cut. Found by a question about what that `rm -rf` actually deletes, which is a better review than reading the line I had just written myself.
The image was already pulled by most jobs; the remaining ones provisioned their own JDK, Node and Gradle cache and so ran on a toolchain nothing else had used. Now `Build plugin`, the protocol-drift check and the release gate use it too, which is the point of having built it. `gradle/actions/setup-gradle` is removed everywhere rather than set to read-only, because it was not doing the job it appeared to be doing. The warm GRADLE_USER_HOME measures 31 GB — 23 GB of extracted IDE transforms under caches/9.5.1 and 7.7 GB of downloaded IDE artifacts under modules-2 — and an Actions cache entry is capped at 10 GB per repository. It could only ever have stored a fraction, evicted it, and re-downloaded the rest next run. The image has no such ceiling. The trade is explicit and worth stating: refreshing what CI has cached is now a deliberate rebuild-and-push, not something that drifts between runs. Also removes the verifier's `Free disk space` step. Inside a container those paths are the IMAGE's, not the runner's, so it had been freeing nothing while looking like this job's safety margin. The margin now comes from the IDEs being baked: nothing is downloaded or extracted at verify time. Recorded because it is the failure everyone hits once: the private package must be granted Read access to this repository in its own settings. The `packages: read` permission widens what the token may ASK for; it does not authorise it against a package the repo was never linked to, and without the link the pull fails with a bare `denied` that reads like a wrong image name. Not containerised, deliberately: `publish`, which holds the Marketplace token and the signing key and is not a test, and CodeQL, which is weekly and is where a container breaks quietly. Not verified: that the image builds with no network at all. The one attempt failed on uid mapping, which says nothing about CI, where the container runs as root.
…code-for-jetbrains into feature/update-pipelines
The CI image was 38.1 GB, 29.1 GB of it the IDEs `verifyPlugin` downloads. Every job in ci.yml pulls its own copy on its own runner, so `Initialize containers` measured 5m37s on a job whose actual work is an 8-second vitest run, and 38 GB on a runner with ~25-30 GB free on the root volume was also flirting with `No space left on device`. The verifier is the only consumer of those IDEs and it runs only on a pull request from develop into main, so it now downloads them when it runs. The set also moves — the verifier resolves from the EAP/RC channels — so the baked copies stopped matching on JetBrains' release schedule, not ours. The warm-up itself was not warming anything. `dependencies --configuration compileClasspath` resolves dependency metadata; it never triggers the artifact transform that EXTRACTS the IntelliJ Platform, which is where the GB are. Measured: it leaves caches/*/transforms at 179 MB with no extracted IDE in it. What warmed the platform was the `verifyPlugin` step, so removing that step alone would have shipped a cold image that re-resolved the platform in every job — and the `> /dev/null 2>&1 || true` meant nothing would have said so. It is now `testClasses`, unredirected and without `|| true`, so a failed warm-up fails the image build. Verified with the network disabled inside the image: `testClasses` compiles in 33s and the 84 frontend tests pass from the baked npm cache. 38.1 GB -> 8.07 GB, 3.61 GB compressed. Also in this change: - .dockerignore: the tree reaching `COPY` goes from 2.09 GB to 3.2 MB (node_modules, build/, .git). `node_modules` is now removed in the same layer `npm ci` creates it, since a later `rm` hides space rather than reclaiming it. - python3 is installed explicitly. It was arriving transitively through dnf-plugins-core, which was itself unused — the Adoptium repo file is written with printf, not dnf config-manager, and curl is already in the base image — and bin/fake-claude, the stand-in the integration tests drive a real ClaudeSession against, is a `#!/usr/bin/env python3` script. Dropping the unused package without naming python3 would have broken the integration suite. - One dnf transaction with tsflags=nodocs, git-core instead of git, /usr/share/locale removed. - codeql.yml: the matrix is split so java-kotlin runs in the image and inherits the warm cache, dropping a setup-java that handed it a different JDK from the rest of the pipeline, while javascript-typescript stays on a bare runner where it is faster. Both display names are unchanged: they are required checks in .github/rulesets/main.json, and a renamed job stops applying its gate silently. - No `:latest` anywhere; the image is `:base`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ipelines Feature/update pipelines
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull request
Summary
What does this PR change and why? One short paragraph is fine.
Related issue
Closes #
Type of change
Risk and rollback
Risk: what breaks if this is wrong, and for whom? (
noneis a validanswer for docs-only changes — say so rather than leaving it blank.)
Rollback: how is this undone once released? Reverting the commit is not a
rollback for a published plugin — a user on the bad version stays there until
they update. If the change touches persisted settings, the transcript format,
or the permission surface, say what happens to a user who already ran it.
Checklist
developbranch (ormainonly for hotfixes).commit-msghook enforces it —install once with
git config core.hooksPath .githooks)../gradlew test verifyPlugin buildPluginpasses locally.verifyPluginis Compatible across the declared range (251 → 263.*)and reports no new internal-API usage (
@ApiStatus.Internal).The CDN download is unreliable here; use
-PlocalIdePath=<dir>[,<dir>…]with locally-extracted IDEs.src/test/kotlin/…forKotlin,
src/test/frontend/…(npm test) for anything undersrc/main/resources/jcef/../gradlew checkDriftis green and the baseline inscripts/drift-baseline.propertiesmatches what was verified.recorded in
THIRD-PARTY-NOTICES.mdif itships in the artifact.
CHANGELOG.mdand
RELEASE_NOTES.mdunderUnreleased.paths in the diff or commit messages.
CONTRIBUTING.md, thearchitectural contract in
CLAUDE.md, and the recordeddecisions in
docs/adr/.How was this tested?
./gradlew test) and frontend tests (npm test)./gradlew runIde) — describe the scenarios youexercised.
visible on every control touched. Automated checks catch roughly half of
real accessibility barriers and none of the judgement calls, so this one
is not delegable to a tool.
Notes for reviewers
Anything tricky, follow-up work, or open questions.