Skip to content

Release 5.0.0 - #35

Merged
serialexperimentslainnnn merged 49 commits into
mainfrom
develop
Aug 6, 2026
Merged

serialexperimentslainnnn merged 49 commits into
mainfrom
develop

Conversation

@serialexperimentslainnnn

Copy link
Copy Markdown
Owner

Release 5.0.0

Summary

The release door for 5.0.0, the standards-compliance major. develop is 38 commits ahead of main;
the last published tag is v4.4.1 and build.gradle.kts already declares 5.0.0. Merging this is what
makes a v5.0.0 tag reachable from main, which is the precondition release.yml's guard job asserts
before any credential comes into scope.

This PR carries no new code of its own. Its purpose is to run the full gate — Plugin verifier,
Static analysis and Dependency audit only open on a pull request into main, so this is the first time
the 38 commits are judged by all of them together.

Related issue

n/a

Type of change

  • New feature
  • Docs / build / CI
  • Security fix
  • Bug fix
  • Refactor (no behavioural change)

Risk and rollback

Risk: this is a major, and it changes packaging and process rather than the plugin's runtime
behaviour. The user-facing surface that is genuinely new is the plan-limits panel; the rest is CI, docs,
static analysis and the dependency-scope correction. The concrete risk is the release machinery itself —
release.yml publishes to the Marketplace behind a human approval on the marketplace environment, and
this is the first release to run through it end to end.

Rollback: reverting the merge does not roll anything back for a user who already updated — a published
plugin only moves forward. If 5.0.0 is bad in the wild, the exit is a 5.0.1 patch, not a revert. Marketplace
also allows delisting a version, which stops new installs but does not touch existing ones. Nothing in this
release migrates persisted settings, the transcript format, or the permission surface, so a user who runs it
and then downgrades manually keeps a readable claude-code.xml.

Checklist

  • PR targets the develop branch (or main only for hotfixes).
    Deliberately not. This is the develop -> main release door described in
    ADR 0001, not a feature PR.
  • Commits follow Conventional Commits.
  • ./gradlew test and npm test pass locally — 694 JVM tests (0 failures, 2 Windows-only skips)
    and 84 frontend tests, plus koverVerify, detekt, spotlessCheck, ESLint, Prettier and
    npm audit --omit=dev (0 vulnerabilities). buildPlugin clean.
  • verifyPlugin passes locally / is Compatible across 251 -> 263.*.
    Not run locally for this PR — it is the Plugin verifier job on this very pull request, which is
    the first place it opens. Judge it there rather than on my word.
  • No new deprecated or scheduled-for-removal IntelliJ Platform APIs (failureLevel includes
    DEPRECATED_API_USAGES, so the verifier job enforces this rather than a reviewer).
  • Tests added or updated for the new behaviour.
  • ./gradlew checkDrift green. Not re-run for this PR. The baseline was last verified at
    claude 2.1.222 / SDK 0.3.222 during the 5.0.0 work; drift.yml runs weekly and files an issue.
  • New dependencies recorded in THIRD-PARTY-NOTICES.md, which now ships inside the artifact.
  • User-visible changes documented in CHANGELOG.md — reproduced below.
  • No secrets, tokens, transcripts or personal absolute paths in the diff or commit messages.
  • Follows CONTRIBUTING.md, CLAUDE.md and docs/adr/.

How was this tested?

  • Unit tests (./gradlew test) and frontend tests (npm test).
  • Manual sandbox (./gradlew runIde).
  • Smoke test on a real IDE install.

Stated plainly rather than ticked: no live in-IDE manual pass was run for this PR specifically. For a
major that touches the permission surface's documentation and ships a new dashboard card, that is the gap a
reviewer should weigh — the automated gate does not cover it.

Notes for reviewers

The CHANGELOG entry below predates the last commit on develop. 04957a8 ("ci: drop the verifier IDEs
from the image and fix the warm-up") landed after the 5.0.0 section was written, and is not reflected in it.
It is CI-only and reaches no user, but the entry is not a complete account of what is on this branch. Worth
folding in before tagging if the CHANGELOG is meant to be read as the record of the release. Summary of it:
the CI image was 38.1 GB because it baked the verifier's IDEs, costing every job a 5m37s container init;
they are gone (38.1 GB -> 8.07 GB, container init now measured at 1m05s), and the image's warm-up was found
to have never warmed anything — dependencies --configuration compileClasspath resolves metadata but never
triggers the transform that extracts the IntelliJ Platform, and > /dev/null 2>&1 || true hid it.

What this PR is really for is the three jobs that have never run on these commits: Plugin verifier
(the only thing that catches a binary incompatibility across 251 -> 263.*, and the asymmetry that let the
4.4.1 /login regression ship), Static analysis and Dependency audit. If any of them is red, that is
new information and not a flake.


[5.0.0] — 2026-08-05

The standards-compliance major. Not a feature release: the repository was taken through the standards
catalogue domain by domain, and the major number reflects that the code changed to comply, not only the
documentation. The break this major records is one of process and packaging, and it is stated plainly rather
than hidden in a patch.

It did not stay purely that, and saying so is cheaper than letting a reader discover it: the release also
carries the plan-limits panel and a run of user-facing fixes (below). Nothing is removed or behaves
differently on purpose — but a release note claiming "no user-facing change" while shipping a new dashboard
card would be the kind of small untruth that makes the rest of the document unusable as evidence.

Security

  • The protocol SDK was declared as a runtime dependency while never being one. @anthropic-ai/claude-agent-sdk sat in dependencies although it is protocol reference material — kept so the Kotlin layer can be diffed against the binary's real surface — and is not executed or packaged. The published artifact contains jars and inlined web assets and zero node_modules entries, which anyone can confirm with unzip -l build/distributions/*.zip | grep -c node_modules. The consequence of the wrong declaration was seven permanent npm audit findings (three high) against code no user ever receives: an alarm backlog that cannot be acted on, which is worse than no alarm because it trains you to ignore the one that matters. Moved to devDependencies, so npm audit --omit=dev — the distributed scope — now reports zero. SECURITY.md states the triage boundary explicitly, with the command to verify it rather than a request to trust it.
  • Written threat model (ADR 0002). SensitiveGuard was strong and undocumented: nothing said what it defends against, which makes coverage unarguable and restarts every bypass discussion from first principles. The ADR states the trust model (the user trusted; the claude binary trusted as software but untrusted as a channel; everything it relays — model output, tool inputs, MCP traffic, file contents, fetched pages — untrusted) and runs STRIDE over the three real surfaces. On indirect prompt injection it records the position deliberately: detection is not attempted, because content-level detection is unsolved and a control built on it would be a liability. Injection is assumed to succeed, and the defence sits where success does not pay — the guard judges the tool call and never the reasoning behind it, so a perfectly-injected model still has to ask to read the key, and still gets the same answer. Non-goals are listed as explicitly as goals.
  • The ignore rules had no protection for key material. .gitignore covered build output and nothing else, so the working tree was one wrong answer away from a committed private key: scripts/bootstrap-ci.sh asks where to save a generated JetBrains signing key, and answering . drops private.pem into the repository. It now leads with a secrets section — *.pem, *.key, *.p12, *.jks, chain.crt, passphrase, private.asc, *.token, .npmrc, .netrc — with a single negation for docs/ci-signing-key.asc, the one key file that must be committed since without the public half nobody can verify a release. Verified by creating each of those files and confirming git check-ignore blocks it while the public key stays committable. Secrets come first in the file for a reason: a build artifact committed by accident is noise, whereas a private key committed by accident is burned — forks, clones, forge caches and CI logs mean rewriting history does not un-leak it, and the key has to be rotated regardless. The file itself remains untracked by design (it ignores itself): a published .gitignore is a public inventory of a maintainer's local directories and tooling, which is reconnaissance for no benefit to anyone installing the plugin.
  • SECURITY.md's supported-versions table still said 2.x.

Added

  • CI/CD on GitHub Actions, with publication gated three independent ways. The repository had no working pipeline at all: the workflows had been deleted, and a comment in .gitlab-ci.yml had been asserting for months that GitHub Actions was "capped (billing)". That was false — the repository is public, and Actions on standard hosted runners is free and unmetered for public repositories; the account's Actions permissions were verified enabled. A false constraint written into a config file gets believed for years, which is precisely what happened. ci.yml now runs the full gate on develop, main and every feature/**, bugfix/** and hotfix/** branch — not only on the PR, because a bar you meet only at PR time is a bar you discover late. codeql.yml adds SAST over Kotlin and JavaScript. release.yml publishes to the Marketplace only when three things hold at once: a vX.Y.Z tag; the tagged commit reachable from main, asserted before any credential is in scope; and a human approval on the marketplace environment, where the four credentials are scoped and exist for no other job. The middle gate is the load-bearing one — without it, anyone who can push a tag can publish from any code, and the review the approval assumes becomes optional. drift.yml runs checkDrift weekly and files an issue rather than committing: whether a new protocol message should be modelled or ignored is a judgement call, and a bot that answers it would bless a gap silently. Every action is pinned by full commit SHA (a tag is mutable, and the action runs with this repository's token), with Dependabot proposing the bumps so the pinning stays free. Build provenance is attested and deliberately not overtrusted — a compromised runner can sign a build that genuinely happened on it.
  • Release artifacts are signed in the pipeline, by a key that is deliberately not the maintainer's. The maintainer key is hardware-backed and non-exportable — which is what makes it worth trusting, and also why it cannot sign inside a runner. Automating the .asc therefore needs a software key in a secret, and that weakening is bounded rather than waved through: the secret is scoped to the approval-gated marketplace environment (no job reachable from a bare tag push can see it), the key expires after a year so an unnoticed leak stops mattering on its own, and its user ID says out loud that it is a CI key. That last point is the actual mitigation — if the two signatures were indistinguishable, a leaked CI key would impersonate a person. The two claims are now documented as distinct: the tag signature says a person authorised this release, the artifact signature says this workflow produced these bytes, and SECURITY.md tells users to check both. Generated by scripts/gen-ci-signing-key.sh, which works in a throwaway keyring and never touches the maintainer's. The CI key is certified by the hardware key, which is what makes the arrangement defensible rather than merely documented: without it a user is asked to trust a fingerprint printed in a file inside the very repository an attacker who could swap the key would control — a tautology, not a trust anchor. With it the chain terminates in hardware, and there is a revocation lever nobody holding the leaked key can undo. scripts/bootstrap-ci.sh performs the whole one-time setup, and docs/CI_SETUP.md documents each step for when it has to be done by hand.
  • Branch protection as versioned code (.github/rulesets/*.json, applied by scripts/apply-rulesets.sh). Both main and develop require a pull request, an up-to-date branch, signed commits, and every CI check. Required approvals are zero, which reads like a hole and is the opposite: GitHub does not let an author approve their own pull request, so on a single-maintainer repository "require 1 approval" with no bypass actors means nothing can ever be merged — not by push, not by PR, not by admin. We established that empirically, by locking the repository and having to unlock it. The gate that remains is the mechanical one, which is also the one that cannot be talked out of. Raise it to 1 when a second maintainer exists; the rulesets carry that instruction inline. No bypass actors, including admins — the previous documentation preserved an admin bypass for a structural blocker that never existed, and a bypass is by construction used at the worst possible moment on the least-reviewed change. .gitlab-ci.yml is removed rather than retained: two pipelines that can each publish is one publisher too many.
  • Accessibility conformance work (WCAG 2.2 AA; the EU Accessibility Act has applied since 28 June 2025). A role="status" aria-live="polite" region declared in the static shell.html — created lazily it would never announce its first message, which is the classic way to ship a silent live region — plus CC.announce with duplicate suppression, so a screen-reader user is told when a turn starts, finishes, or is blocked on a permission card. The transcript streams without ever moving focus, so without this the turn simply stalls in silence. Also a :focus-visible baseline covering every element whose outline the stylesheet suppresses (the find bar's input had no replacement at all), honoured under forced-colors rather than overridden. Ten frontend tests pin the structural guarantees; they do not certify conformance, which still requires a keyboard and screen-reader pass by a person.
  • Third-party attribution ships inside the artifact — THIRD-PARTY-NOTICES.md, LICENSE and LICENSES/* are packaged under META-INF/. The plugin redistributes marked, DOMPurify and highlight.js, and a permissive licence's notice obligation binds on redistribution: a notices file that exists only in the repository does not discharge it for someone who installs the zip. DOMPurify is dual Apache-2.0 OR MPL-2.0, so the choice is recorded rather than left implicit.
  • AGENTS.md — the operational runbook for agentic development (commands, gates, boundaries), complementing CLAUDE.md, which stays the architecture.
  • docs/adr/ — three decision records: 0001 release process, 0002 threat model, 0003 i18n deferred with the triggers that reopen it.
  • Conventional Commits enforcement via commitlint and a versioned .githooks/commit-msg (enable with git config core.hooksPath .githooks). The hook self-tests and degrades to advisory if its own toolchain fails, specifically so it can never become a reason to reach for --no-verify.

Changed

  • Published tags are now immutable, recorded in ADR 0001 §3 as a correction of a real violation: v4.3.2 and v4.4.1 were each force-re-cut three times after being pushed. A tag is the identity of a shipped artifact; moving one means two people can hold different trees, different zips and different checksums while both believe they have the same version — which defeats the single thing a signature is for. A mistake found after tagging is now fixed by the next patch version. The already-moved tags are left alone, because re-cutting them to "fix" history would repeat the exact mistake.
  • LoginCoordinator extracted from ClaudeSession (1965 → 1826 lines). The OAuth sign-in is a subsystem in its own right — the TTY-less --print session cannot host an interactive login, so it happens outside the session entirely through three ordered paths — and it now owns its own state. Mechanical, no behaviour change, full suite green across it. The two further extractions that were considered (SessionRestorer, RewindCoordinator) were deliberately not made: restore is 23 lines that touch six pieces of session state, and rewind is one of six identically-shaped control-request delegates. Both would have bought indirection rather than cohesion, and saying so is the point of recording it.
  • package.json declared "license": "ISC" on a GPL-3.0-only repository and lacked "private": true — i.e. it was publishable to npm under the wrong licence. Corrected.
  • No contact email is published anywhere in the project. The <vendor email> attribute is optional and has been dropped from plugin.xml; vulnerability reports now go through GitHub private security advisories rather than an inbox. That is the better channel on its own merits and not only a privacy measure: the report lands in a private thread attached to the repository, the discussion and fix stay linked to it, and a CVE can be requested from the same advisory — whereas an address in a public file is scraped far more often than it is used by a reporter.
  • Protocol baseline re-verified and advanced to claude 2.1.222 / SDK 0.3.222; ./gradlew checkDrift green, protocol surface unchanged.
  • The pull-request template now asks for risk and rollback — and for a published plugin, reverting a commit is not a rollback: a user on the bad version stays there until they update.

Internal

  • The frontend test harness (src/test/frontend/helpers/load.js) now extracts the shell DOM from the real shell.html instead of a hand-copied approximation. The copy had already drifted — it lacked #a11y-status — which is the worst failure mode a harness has: it does not fail loudly, it quietly tests something that is not the product.
  • Frontend suite: 44 → 54 tests. JVM suite 677 → 682.

Static analysis, formatting and coverage — installed, then acted on

  • detekt and Spotless/ktlint added, and the 203 findings they raised were fixed rather than frozen. Until now the entire quality bar for 13k lines of Kotlin rested on review, which is precisely what the standards say to mechanise. The first run produced 492 findings; tuning the rules with the reasoning written at each setting brought it to 203, and those were then worked down to 2. config/detekt/baseline.xml holds exactly those two, both about ClaudeSession, both explained inside the file — it is a record of a decision, not a drawer. The distinction matters: a 203-entry baseline is a promise to nobody, a 2-entry one is a claim somebody has to defend in review.
  • The dispatch tables were split in two levels, keeping compile-time exhaustiveness. ClaudeSession.onEvent was a single when over 47 event types — 244 lines, cyclomatic complexity 111 — the one function where every protocol concern in the plugin met. ClaudeEvent now declares seven sealed sub-interfaces (Stream, Conversation, Control, Task, Notice, SessionSignal, HookTelemetry) and dispatch picks the group, then the variant. The grouping is expressed in the type on purpose: a sealed hierarchy keeps the compiler checking exhaustiveness at both levels, so a new protocol event that nobody handles is a compile error rather than a silently dropped frame — which is the property checkDrift exists to protect, and was not up for trade against a complexity threshold. The groups are semantic, not cosmetic: they differ in what the host owes the binary (a Control frame must be answered or the binary hangs; a Notice is fire-and-forget). JcefBridge.Msg and JcefChatPanel.onBridgeMessage (complexity 46) got the same treatment, with the message groups mirroring the bridge's parsers one-for-one.
  • Several when chains were dictionaries written as control flow, and are now data: ProtocolParser.parseSystem had 25 arms of which 21 were the same expression with two names substituted (complexity 29 → a Map), likewise the top-level frame decoder, and EditorContextProvider.langForExtension (26 arms → a lookup table). Adding a protocol subtype is now one line, and the shared fallback wiring is written once instead of 21 times where a mistyped argument would have been invisible.
  • Coverage is gated per package (koverVerify), because risk here is not evenly spread: permission/ decides whether the agent may read your SSH key, ui/ paints a browser. Thresholds sit slightly below what each package measures, so they catch regression instead of inviting test-padding. ui/, context/, process/, actions/ and util/ are excluded with the reason stated rather than gated at a token value — gating them at 20% would dress the same fact up as a passing check. Policy, measured numbers and the known gaps are in docs/RELEASE_CHECKLIST.md §Coverage policy.
  • A "≥90% coverage target" was cited in the build for a requirement that did not exist. build.gradle.kts claimed the figure was "documented in docs/RELEASE_CHECKLIST.md"; that file had never mentioned coverage, and the real number was 53.3%. A number nobody measured, pointing at a rule nobody wrote.
  • ESLint and Prettier now cover the shipped JCEF frontend — ~3.6k lines of JavaScript that ride inside the plugin jar and had never passed through any tool. no-eval, no-implied-eval and no-new-func are errors because the page runs under a hash-pinned CSP with no 'unsafe-eval': without the gate, code Chromium will silently refuse in a user's IDE can still reach main. Vendored marked/DOMPurify/highlight.js are excluded — a finding in them is not ours to fix, and fixing it would fork a dependency.
  • A Static analysis job (detekt, spotlessCheck, koverVerify, npm run lint, npm run format:check) is now a required check on both protected branches. Everything above is only worth having if breaking it fails a merge.
  • Two rules that both tools enforced were given a single owner each: max-line-length and function-naming are detekt's, because only detekt can scope an exception to the test tree. Running both meant the stricter-but-blinder tool decided, which is how you end up reformatting single-line NDJSON protocol fixtures to satisfy a tool that cannot be told they are fixtures.

Fixed — defects the tooling surfaced

  • Token counts and CSS alpha values were formatted with the machine's locale. TokenFormat.trimDecimal used the default-locale "%.1f", so on a comma-decimal machine (Spanish, German, French…) a count rendered as 1,2k inside otherwise-English UI — and worse, the trailing-.0 test stopped matching, so a flat 1000 tokens displayed as 1,0k instead of 1k. The same bug in JcefTheme.rgba was not cosmetic at all: it emitted rgba(217, 119, 87, 0,140) — four components instead of three — so the browser discarded the declaration and the --accent-soft/--link-soft washes (text selection, the code-block Copy hover, the "View diff" hover, blockquote backgrounds) never rendered on those machines. Also fixed in the context-usage percentage and the colour-to-hex helper. All now pin Locale.ROOT.
  • Diff tabs were being persisted into the workspace and could never be restored. Our diffs are in-memory previews (ChainDiffVirtualFile over a mock:/// URL); the platform persists every open editor tab by URL without filtering by file system, so on the next start each one resolved to nothing. One workspace here had accumulated 13 such entries — all named Claude · SKILL.md, since the tab title is the file name and a skills repository has one SKILL.md per directory — producing 26 WARN EditorsSplitters - No file exists lines on every single launch. DiffTabCleanup now closes them on projectClosingBeforeSave, the one hook that runs before the state is written (projectClosing would be one step too late), and a wiring test pins the plugin.xml registration against the shipped descriptor — the failure mode being silence, not a stack trace.
  • CloseAllDiffsAction moved to a background update thread. It reads one CopyOnWriteArraySet's size; keeping it on the EDT put it in the queue behind everything the IDE does at startup. InterruptAction deliberately stays on the EDT and now says so in the code: it reads ContentManagerImpl.mySelection, an ArrayList mutated on the EDT with no synchronisation and no threading assertion, so moving it would trade a cosmetic log line for a rare IndexOutOfBoundsException.
  • Four defects in the shipped frontend, all found by its first lint run: obj.hasOwnProperty(k) in both DOM-building helpers (breaks if the object carries its own hasOwnProperty — and those helpers build DOM from host-supplied data), an empty catch in the Vibe Mode theme restore that silently left the theme half-reverted, and two dead functions (isAgentTool, esc) nobody called.
  • sniffMediaType no longer confuses any RIFF container for WEBP. Rewritten around named signatures (complexity 23 → 4), it now checks the four-byte form type that actually identifies the format, not just the RIFF header that WAV and AVI share.

Fixed — a tab-killing regression, and the silences it hid

  • No chat could be opened or restored (regression, introduced on this branch). JcefChatPanel.pendingUntilReady was declared below the init block that uses it. Kotlin runs property initializers and init blocks in declaration order, so the list was still null while init ran and the constructor threw NullPointerException — taking the whole tab with it, on new chats and on startup restore alike. lastUsage/lastUsageAt had the identical defect and stayed silent, because a nullable reference and a primitive read as null/0 instead of throwing: the loud version of this bug was the lucky one. The compiler does not catch it — it flags a direct reference in an initializer, but here the read happens inside a function called from init, which it cannot see through — so InitOrderContractTest scans the sources and fails the build on any class-body property declared after its own init.
  • Nothing said the agent was still starting. The binary is now launched before the tab is built (start() only dispatches, so claude boots while JCEF creates its browser), and a boot screen holds the tab until the process is up. Three states, not two: running, starting, and neither — that last one is a launch that failed (missing binary, declined trust prompt, refused remote-mount project) and it must bring the screen down, or the tab stays covered forever with no way to reach the notification explaining why. The screen is declared visible in the static shell, since at page load the process genuinely is not up yet.
  • Context and cost were a minute late, twice over. A javax.swing.Timer's initial delay equals its interval, so the first poll came a full QUOTA_POLL_MS after the panel attached — and that tick landed while the binary was still launching and returned early on the not-running guard, costing a second interval. Process-ready, tab-open and both turn edges now poll directly, and the timer retires at the end of a turn: context and cost cannot move while a session sits idle, so polling forever was a round-trip through the binary, per tab, for two numbers that provably had not changed.
  • The plan-limit figures disagreed with themselves. A get_usage reply refreshed the dashboard bars but not the composer's dots, so the same number appeared immediately in one surface and "a while later" in the other, whenever some unrelated state change happened to re-push. Both are pushed together now. Opening the dashboard also refreshes them, which requestUsage's own contract had claimed and the code had never done.
  • Reasoning tokens, context and output are rendered at 0 instead of omitted. An item that only appears once it is non-zero is indistinguishable from one that failed to load — which is exactly how a fresh tab read: a lone "Idle" and no figures. Cost stays gated, because a currency amount of zero is noise rather than an ambiguity to resolve.
  • The CLI's <tool_use_error> wrapper reached the transcript verbatim. claude 2.1.222 wraps a failed tool result's content in that tag pair and carries the same message unwrapped in a sibling field — framing for the model, not text for a human. Rendered as-is it put raw markup in a native GUI, the "never mirror raw CLI output" antipattern this plugin exists to avoid. Stripped only when it encloses the whole payload, so output that legitimately mentions the tag survives; is_error already conveys the failure structurally, and is what reddens the card.
  • A failed tool card hid its own error. Tool output lives behind the card's collapse, so for a failure the entire message was "the header is red", and the text scrolled sideways rather than wrapping — hiding the actionable half at the end of the line. A failed card now opens itself once (tracked on the node, so it never fights a user who deliberately collapsed it) and its error text wraps. Healthy output still scrolls: wrapping code or a log corrupts its alignment.
  • ToolSearch was missing from the SensitiveGuard trust allowlist, along with AskUserQuestion, Mcp and FileRead/FileEdit/FileWrite. ToolSearch is the one that mattered: it loads the schema of every deferred tool, so on a session that defers them, the call that unlocks all the others was the one landing in the third-party branch. Entries are only ever added to that list — it is a trust allowlist, not an inventory, and a missing first-party name is precisely the 4.4.0 hard-DENY incident. Found by diffing it against a live session's real tool inventory rather than against the SDK's type names, which are not the runtime registry (the SDK calls them FileRead/FileEdit/FileWrite; the tools are Read/Edit/Write).
  • A Markdown link whose href is a path did nothing when clicked. The host handled https:// and jb://open and dropped everything else without a sound — so [BACKLOG](docs/BACKLOG.md) was inert while bare paths written in prose worked, making the more deliberate link the one that failed. Both routes now go through a single authorising gate (LinkResolver.isOpenable). The scheme test requires two or more characters before the colon, so a Windows drive (C:\src\main.kt) stays a path rather than being mistaken for a URI scheme.
  • Copy on a message copied and said nothing. Message-level buttons carry their own click handler and never reached the delegated code-block path that flashes "Copied", which reads as a broken button — and was reported as one. The flash helper is now exported and shared rather than reimplemented, so wording and duration cannot drift. The .copied class had been applied by the JS since 4.0.4 and had no CSS rule at all; it now has one.

serialexperimentslainnnn and others added 30 commits August 5, 2026 16:43
Audited the repository against git-workflow-standards. Adds the governance
the standard requires and that this repo lacked, minus anything CI-bound
(no CI gate exists yet - a local lab is planned).

- commitlint.config.mjs + a VERSIONED .githooks/commit-msg gate, so the
  rule travels with the repo instead of living in one laptop. The hook
  self-tests against a known-good message first and degrades to a warning
  if the toolchain is broken: a hook that fails closed on its own bugs
  gets bypassed with --no-verify, and then protects nothing.
- .gitattributes: this plugin ships and is tested on Windows, and
  publishes signed zips. Without a declared policy a clone with
  autocrlf=true rewrites every file and invalidates artifact signatures.
  Vendored frontend bundles marked -diff so a bump is reviewable.
- docs/adr/0001: records the two justified deviations (GitFlow for
  installable versioned software; GPG-on-YubiKey for real revocation and
  because the same key signs the distributed artifacts) and corrects a
  real violation - v4.3.2 and v4.4.1 were each re-cut and force-pushed
  three times. Published tags are now immutable; a post-tag mistake is
  fixed by the next patch version.

CHANGELOG generation is deliberately still open: the tooling silently
skips non-Conventional commits, so the message gate has to land first.

Refs: git-workflow-standards §3.2, §4.1, §5.2, §5.3
The published zip redistributes third-party code - marked 12.0.0 (MIT),
DOMPurify 3.0.11, highlight.js 11.9.0 (BSD-3-Clause) vendored into the
plugin jar, and kotlinx.serialization 1.7.3 (Apache-2.0) as its own jars.
MIT, BSD-3-Clause and Apache-2.0 all require the copyright notice and
licence text to be preserved ON REDISTRIBUTION, and a file sitting in the
Git repository does not accompany the binary a user installs from the
Marketplace. This was unmet.

processResources now packages THIRD-PARTY-NOTICES.md, LICENSE and the
three licence texts into META-INF/ of the plugin jar - verified present in
the built artifact, not just in the source tree. Generated at build time
from one root-level source of truth rather than a checked-in copy, so the
notices cannot drift from the files they describe.

Every licence was verified by reading the LICENSE of the exact shipped
version, never the manifest or a badge. That found a real one: DOMPurify
3.0.11 is dual "Apache-2.0 OR MPL-2.0". An OR expression is a choice the
redistributor must make and record; leaving it unstated is an unmade
decision. Apache-2.0 is selected, with the reasoning recorded in the
notices file. kotlinx.serialization ships no META-INF/LICENSE in its jars,
so its licence was read from the project source instead of inferred.

Also corrects the documented JDK path across CLAUDE.md, README,
CONTRIBUTING and docs/: ~/.local/jdks/jdk-21.0.11+10 no longer exists on
this machine (it is now ~/.jdks/jbr-21.0.11), and gradlew failed with
"JAVA_HOME is set to an invalid directory" rather than anything a build
log filter would obviously catch.

Refs: opensource-licensing-standards §3.5, §3.6, §5.4, §8
Audited the JCEF web UI against WCAG 2.2 AA. This is a conformance
requirement, not a nice-to-have: the plugin is distributed to consumers in
the EU, where Directive (EU) 2019/882 has applied since 28 June 2025 (in
Spain, Ley 11/2023).

4.1.3 Status Messages (AA) - the transcript streams without ever moving
focus, so a screen-reader user got NO signal that Claude started, finished,
or was blocked on them: the turn simply stopped, silently. Adds a polite
live region declared in the static shell (it has to exist in the DOM before
text is written into it, or the first change is never announced) plus
CC.announce. Only turn EDGES are announced, never per token - a region
updated on every delta talks over itself and gets switched off, which is
worse than silence. The most important case is a pending permission card:
it appears without taking focus, so it was previously undetectable.

2.4.7 Focus Visible / 1.4.11 Non-text Contrast (AA) - the stylesheet
suppressed the default outline in five places. Four had some replacement;
.find-input had NONE, so focus there was invisible. Adds a :focus-visible
baseline covering the controls built as <span role="button"> (the code-block
Copy control, menu items, tool cards), which get nothing for free because
they are not native buttons, and honours forced-colors instead of overriding
the system's focus colour.

prefers-reduced-motion was already correct (a universal reset covering all
eight keyframe animations) and is now pinned by a test so it cannot decay
into a hand-picked subset as animations are added.

Refactor: the frontend test harness hand-copied the shell DOM and had
already drifted - shell.html gained the live region, the harness did not, so
tests exercised a DOM the product does not have. It now extracts the body
from the real shell.html. A mount point added to the shell reaches tests
automatically; one removed breaks the tests that relied on it.

Scope stated honestly in the test file: automated checks catch ~40-57% of
real barriers and none of the semantic judgements. These pin the structural
guarantees that regress silently. They do not certify conformance - that
still needs keyboard-only and screen-reader passes by a person.

Refs: accessibility-standards §3.2, §3.4, §3.8, §4.1
`@anthropic-ai/claude-agent-sdk` sat in `dependencies` while CLAUDE.md has
always said it is "protocol reference only, not distributed" — and it is: the
published artifact is jars and inlined web assets, with zero `node_modules`
entries (verifiable with `unzip -l`). The mismatch produced seven permanent
npm-audit findings (three high) against code no user ever receives, which is
worse than useless: a backlog of alarms that cannot be acted on trains you to
ignore the one that matters.

Moving it to `devDependencies` makes the declaration match reality, so
`npm audit --omit=dev` — the distributed scope — now reports zero. The full
tree still reports the same seven; they are build-time tooling and are handled
as maintenance. SECURITY.md states that triage boundary explicitly, with the
command to verify it rather than trust it.

`checkDrift` was the constraint on this change: it reads the SDK out of
`node_modules` and runs `npm update`. Both survive, because `npm install`
installs devDependencies — only `--omit=dev` would break it, and nothing uses
it. Verified green, and it reported real drift while doing so, so the baseline
moves with it: SDK 0.3.220 → 0.3.222, binary 2.1.220 → 2.1.222, protocol
surface unchanged.

Also corrects the package manifest itself: it declared `"license": "ISC"` on a
GPL-3.0-only repository and lacked `"private": true`, i.e. it was publishable
to npm with the wrong licence. And the supported-versions table still said 2.x.
The control was strong and undocumented: nothing stated what it is defending
against, which makes it unreviewable. Coverage cannot be argued, and every
discussion of a proposed bypass restarts from first principles.

ADR 0002 states the trust model — user trusted, `claude` binary trusted as
software but untrusted as a channel, everything it relays untrusted — and runs
STRIDE over the three surfaces that actually exist: the binary as a child
process, third-party MCP servers, and model-returned content.

The third is the one worth having written down. We do not attempt to detect
prompt injection; content-level detection is unsolved and a control built on it
would be a liability. The document says so explicitly and records where the
defence actually sits instead: the guard judges the tool call and never the
reasoning behind it, so an injection that succeeds completely still has to ask
to read the key and still gets the same answer. Also records why MCP servers
are denied rather than asked about, and why caller trust is an allowlist —
names are attacker-supplied, and 4.4.0 is the cautionary tale of that list
going stale (it failed safe, which is why the allowlist was the right shape,
but it still failed).

Non-goals are as load-bearing as the goals, so they are listed: we do not
defend against the user, against a compromised host, or against every
obfuscation. That last one draws the line triage keeps blurring — failing to
*recognise* a path is a pattern gap; enforcement of a match, once made, is
absolute.

Claims verified against the code before writing, not from memory:
`binaryPermissionMode` rewriting acceptEdits/bypassPermissions to `default`,
and the CSP's `default-src 'none'` / `connect-src 'none'` / hash-pinned scripts.
SECURITY.md links here so a reporter can tell a finding from a stated position
before writing the email.
Until now the entire quality bar for ~13k lines of Kotlin and ~3.6k lines of
shipped JavaScript rested on review. That is the thing the standards say to
mechanise: if formatting is discussed in a review, a formatter is missing.

Adds and configures:

- detekt 1.23.8 (config/detekt/detekt.yml). Every non-default setting carries
  its reasoning AT the setting, not in a message nobody will find again. Three
  rules were re-calibrated rather than suppressed: TooManyFunctions counts only
  the public surface (the default penalised extracting helpers, which is the
  exact fix the complexity rules demand); LongParameterList ignores defaulted
  parameters (one absent from the call site is not part of the problem the rule
  exists for); MagicNumber skips the colour palette (a 0xRRGGBB inside a
  property named DIFF_ADDED_BG is that name's definition).
- Spotless 8.9.0 / ktlint 1.8.0. max-line-length and function-naming are
  disabled here and left to detekt: both tools enforced them, only detekt can
  scope an exception to the test tree, and running both meant the
  stricter-but-blinder one decided.
- ESLint + Prettier over the JCEF frontend, which ships inside the plugin jar
  and had never passed through any tool. no-eval / no-implied-eval / no-new-func
  are errors because the page runs under a hash-pinned CSP with no
  'unsafe-eval': without this gate, code Chromium will silently refuse in a
  user's IDE can still reach main. Vendored marked/DOMPurify/highlight.js are
  excluded — a finding there is not ours to fix.
- koverVerify, gated per package. Risk is not evenly spread: permission/ decides
  whether the agent may read an SSH key, ui/ paints a browser. Thresholds sit
  slightly BELOW what each package measures today, so they catch regression
  rather than invite test-padding. ui/, context/, process/, actions/ and util/
  are excluded with the reason stated rather than gated at a token value.

config/detekt/baseline.xml holds TWO entries, both about ClaudeSession, both
explained inside the file. It started this work at 203; the other 201 were
fixed, not frozen.

A `Static analysis` job runs all of it and is a required check on main and
develop, so none of the above depends on anyone remembering to run it.
Three strands that could not be separated into their own commits: the formatter
rewrote the same files the refactor touched, hunk for hunk, so splitting would
have meant manual surgery across ~100 files for commits that would not build in
isolation. Recorded here rather than pretended otherwise — it also means there
is no pure formatting commit to add to .git-blame-ignore-revs.

REFACTOR — the dispatch tables, split in two levels

ClaudeSession.onEvent was one `when` over 47 event types: 244 lines, cyclomatic
complexity 111, the single function where every protocol concern in the plugin
met. ClaudeEvent now declares seven sealed sub-interfaces and dispatch picks the
group, then the variant. The grouping lives in the TYPE on purpose: a sealed
hierarchy keeps the compiler checking exhaustiveness at BOTH levels, so a new
protocol event nobody handles is a compile error rather than a silently dropped
frame — the property checkDrift exists to protect, and not up for trade against
a complexity threshold. The groups are semantic: they differ in what the host
OWES the binary (a Control frame must be answered or the binary hangs; a Notice
is fire-and-forget). JcefBridge.Msg and JcefChatPanel.onBridgeMessage got the
same treatment. Several `when` chains were dictionaries written as control flow
and are now data (ProtocolParser.parseSystem: 25 arms, 21 identical bar two
names).

FIXED — defects the tooling and live probing surfaced

- Locale-dependent formatting. TokenFormat rendered "1,2k" on a comma-decimal
  machine, and worse, the trailing-.0 test then stopped matching so a flat 1000
  showed as "1,0k". The same bug in JcefTheme.rgba emitted four rgba components
  instead of three, so the browser DISCARDED the declaration and the accent/link
  washes never rendered. Now pinned to Locale.ROOT.
- Diff tabs were persisted into the workspace and could never be restored: they
  are mock:// in-memory previews, and the platform persists open tabs by URL
  without filtering by file system. One workspace had 13 such entries, all named
  "Claude · SKILL.md", producing 26 warnings on every launch. DiffTabCleanup now
  drops them on projectClosingBeforeSave — the one hook that runs BEFORE the
  state is written.
- Every chat animation was dead. The accessibility pass had added a
  prefers-reduced-motion block, and inside JCEF that query reports `reduce`
  regardless of the desktop (measured: it matched while GNOME had animations
  enabled), so it flattened everything. The tell was Vibe Mode's rainbow, which
  survived because it is a setInterval, not a keyframe. Motion is now reduced
  only when the user asks, via a plugin setting, off by default.
- Requests needing a live binary were dropped. The panel is built BEFORE
  session.start() runs, so requestMcp/requestVersion/requestUsage all found
  isRunning() == false and did nothing; each looked like its own bug. They now
  queue until ready.
- sniffMediaType treated any RIFF container as WEBP; it now checks the form type
  that actually identifies the format (WAV and AVI share the RIFF header).

FEAT — plan limits, in both surfaces

get_usage is a control request the plugin has known about since 4.0.1 and never
sent; it returns every rate-limit window at once plus the extra-credit balance.
The dashboard shows one bar per window with its reset countdown; the composer
readout shows a dot per window. One severity scale for both: blue under 65%,
amber under 85%, red above — status hues, never the product accent, because a
quota bar that is coral at 10% and coral at 95% has said nothing. Crossing 65%
and 85% announces once per threshold per window (a warning that repeats every
refresh is one users learn to ignore), and 85% also raises an IDE notification.

ClaudeSession.rateLimits is now per-window: the single field could only ever
hold whichever event arrived last, so a user on a weekly limit saw their
five-hour bar and concluded they had room.

Also adds CC.diagnostics, which reports what the embedded browser actually
resolves to idea.log. The UI is a browser nobody can open devtools on, and
without it a CSS rule that silently does not apply is indistinguishable from a
backend that never sent the state — which cost three wrong diagnoses before it
existed.

Suite: 682 -> 689 JVM, 54 -> 67 frontend.
Documents what the release actually did, including the parts that did not go to
plan — a release note that only lists wins is one nobody can act on.

- CHANGELOG gains two sections: the static-analysis work (203 findings fixed
  rather than frozen, down to 2 recorded exceptions) and the defects it exposed.
- docs/RELEASE_CHECKLIST.md gains a Coverage policy with the MEASURED
  per-package numbers, what is excluded and why, and the Kover 0.9.2 limitation
  behind the current floor+aggregate shape. It also corrects a claim that never
  had a basis: build.gradle.kts cited a ">=90% target documented in
  RELEASE_CHECKLIST.md", that document had never mentioned coverage, and the
  real figure was 53%.
- docs/BACKLOG.md is new. Every entry was PROBED against the real binary rather
  than inferred from the SDK types, including the two that are deliberately not
  worth doing, so nobody investigates them a third time.
- AGENTS.md, CI_SETUP.md, BRANCHING.md and the ADRs updated for the GitHub
  Actions pipeline that replaced the deleted GitLab one.

Known gaps, stated rather than left to be discovered: the risk-ranked human
review is not done; the reduced-motion decision generalises from a measurement
on one machine; and Fable's weekly window cannot be shown at all, because
get_usage returns every per-model window as null (measured, not assumed).
Squash and rebase merging are now disabled at the repository level, leaving
merge commit as the only option. BRANCHING.md said the opposite for `develop`
("squash or merge"), which was worse than saying nothing: it invited the exact
click that destroys signatures.

Both disabled methods REWRITE commits — new SHAs, new committer information —
which invalidates the author's hardware-backed signatures and replaces them with
GitHub's web-flow key. The "signed commits required" rule would still pass, and
that is the trap: the commits are signed, just no longer by the author, which is
the whole property the rule exists to provide.

The cost is recorded too: the merge node itself is GitHub-signed, because the
alternative is a local merge and a direct push, which the pull-request
requirement blocks. Relaxing that to save one commit's provenance would be a far
worse trade. Release provenance therefore rests on the signed tag, not on the
merge node.

Also corrects the status-check lists, which had not been updated when the
`Static analysis` gate was added.
CodeQL flagged two high-severity alerts on the same line of the frontend test
harness: "incomplete multi-character sanitization" and "bad HTML filtering
regexp". Both were correct: the pattern used to strip script tags does not match
a closing tag written with a trailing space, so something that reads like a
sanitiser silently is not one.

The exposure was low. This reads OUR OWN shell.html off disk inside a test, and
the reason script tags are stripped is load control — the harness injects the
app modules itself, in a controlled order — rather than sanitisation. That is a
reason to rank the finding, not a reason to keep it: a regex that looks like
HTML sanitisation is a pattern someone copies to a place where it does matter.

Fixed at the root rather than by widening the pattern. jsdom is already a
dependency of this harness, so the file is parsed and the script elements are
removed as nodes. All 67 frontend tests still pass, which is the check that the
harness still builds the same DOM as the product.

Recorded for the risk review: the two CodeQL contexts the rulesets require
(java-kotlin, javascript-typescript) both PASSED — they report "the analysis
ran". The aggregate CodeQL check, which reports "something was found", is
required by neither branch. A high-severity finding can therefore merge with no
gate objecting.
A regression introduced by this branch made the plugin unusable, and fixing it
surfaced a run of smaller defects that all shared one shape: something failed or
was still loading, and the UI said nothing at all.

THE REGRESSION

JcefChatPanel.pendingUntilReady was declared below the `init` block that uses it.
Kotlin runs property initializers and init blocks in declaration order, so the
list was still null while init ran and the constructor threw. It took the whole
tab with it: no chat could be opened or restored. lastUsage/lastUsageAt had the
same defect and stayed silent, being nullable and primitive. InitOrderContractTest
now fails the build on any property declared after its class's init.

WAITING IS NOW VISIBLE

The binary is launched BEFORE the tab is built (start() only dispatches, so
`claude` boots while JCEF creates its browser), and a boot screen holds the tab
until the process is up. Three states, not two: running, starting, and neither —
that last one is a failed launch and must clear the screen, or a missing binary
leaves the tab covered forever.

Context and cost no longer wait on a clock. The poll timer's initial delay equals
its interval, so the first reading was a minute late; worse, that tick landed
while the binary was still starting and returned early, costing a second minute.
Ready, tab-open and both turn edges now poll directly, and the timer retires at
the end of a turn instead of round-tripping forever on an idle session.

Plan limits were pushed to the dashboard but not to the composer, so the same
number appeared instantly in one place and "later" in the other. Reasoning tokens,
context and output now settle at 0 rather than being omitted until non-zero: an
absent figure is indistinguishable from one that failed to load.

FAILURES THAT WERE HARD TO READ

- The CLI wraps failed tool results in <tool_use_error>...</tool_use_error>
  (verified in claude 2.1.222, which carries the same text unwrapped in a sibling
  field). Rendered verbatim it put raw markup in a native GUI. Stripped only when
  it encloses the whole payload; is_error already conveys the failure.
- A failed tool card stayed collapsed, so the entire message was "the header is
  red", and its output scrolled sideways rather than wrapping, hiding the half
  that says what to do instead. It now opens itself once, and error text wraps.
- ToolSearch was missing from SensitiveGuard.AGENT_TOOLS, so on a session that
  defers tools, the call that loads every other tool's schema landed in the
  untrusted branch. Added, with AskUserQuestion, Mcp and FileRead/Edit/Write.
  Entries are only ever added here: this is a trust allowlist, not an inventory,
  and a missing first-party name is the 4.4.0 hard-DENY incident.
- A markdown link whose href is a path did nothing when clicked: the host handled
  https:// and jb://open and dropped the rest without a sound. Both paths now go
  through one gate (LinkResolver.isOpenable). The scheme test requires two or more
  characters so a Windows drive letter stays a path.
- Copy on a message copied and gave no feedback, which reads as a broken button;
  the shared flash helper is now exported rather than reimplemented, and the
  .copied class finally has a CSS rule — it had never had one.

Suite: 692 -> 694 JVM, 67 -> 84 frontend. verifyPlugin Compatible on IC-251,
IC-252, IU-253, IU-261 and IU-262.
Two CI failures, one of them ours.

OURS — the `Static analysis` job died with an IOException from DriftLiveCheck.
Kover instruments and aggregates EVERY `Test` task in the project, and
`checkDrift` is registered as one, so it had silently become a dependency of
`koverVerify`. That task is on-demand by design: it downloads the latest SDK and
probes a `claude` binary installed on the machine. A runner has no such binary.

The task's own documentation already said "NOT wired into `check`" — it just was
not true of the coverage graph, and nothing checked that it was. It passed
locally for the one reason that makes this class of bug expensive: the
maintainer's machine has the binary, so "on-demand" and "wired in" looked
identical until CI ran it. Verified with `koverVerify --dry-run`, which listed
`:checkDrift` in the graph before and does not after.

NOT OURS — the `JVM tests` job failed resolving
com.jetbrains.intellij.platform:test-framework with a 502 Bad Gateway from
cache-redirector.jetbrains.com. Nothing to fix; it needs a re-run.

Also bumps the IntelliJ Platform Gradle Plugin 2.16.0 -> 2.18.1, which the build
had been warning about on every run.

Stated plainly: the plugin bump is NOT verified locally. A 2.16 -> 2.18 jump
touches the whole build, and CI is the verification here rather than a claim
made in advance.
Reverts the 2.16.0 -> 2.18.1 half of the previous commit. The `checkDrift`
coverage-graph fix in that commit stands; only the plugin version goes back.

On 2.18.1 the headless suite never runs. `ChatSessionManagerHeadlessTest` hangs
before its first assertion, and a thread dump puts the EDT here:

  BasePlatformTestCase.setUp
    -> LightPlatformTestCase.doSetup
      -> IndexingTestUtil.waitUntilIndexesAreReady   (267s and counting)

That is the platform's own fixture waiting for indexing that never completes.
Our code is not on the stack at all — the bump changes which platform
test-framework is resolved, so this is a fixture-level regression rather than
something fixable from this side.

Measured both ways rather than inferred: the same test on 2.16.0 finishes in 19
seconds, and the full gate is green at 694 tests.

The version is now pinned with the reason written AT the setting, because the
build prints an "outdated" warning on every single run and the next person to
see it will otherwise do exactly what I did. Re-attempt it as its own change
with the headless suite as the acceptance test, not as a drive-by on a release
branch.

I shipped that bump unverified and said so in the message; this is what it cost.
The one-line lesson is the boring one: a build-tooling bump is a change like any
other and does not get to skip the suite because it looks like configuration.
Merging `develop` into `main` now publishes. The tag is DERIVED from the version
in build.gradle.kts rather than supplied alongside it, so the two can no longer
disagree — the mismatch the old flow guarded against with a comparison simply
cannot occur.

A merge that does not bump the version publishes nothing: the workflow finds the
tag present, logs a notice and stops. It does not fail. `main` legitimately
receives merges that are not releases, and a red run on each of those is an alarm
people learn to ignore. Published tags stay immutable, which is the correction
ADR 0001 records after v4.3.2 and v4.4.1 were each force-re-cut three times.

Pushing a tag by hand still works, as the escape hatch for re-cutting after a
failed publish without an empty commit on main. A tag this workflow creates does
not re-trigger it, so there is no loop.

The tag is cut AFTER the approval and the publish, not in the guard. Created
earlier it would name a version that was never published when a build fails or an
approval is declined — and since tags here are immutable, that would block the
next attempt. Cutting it last makes it mean "this was published", which is the
only claim it can honestly make once the version, not the tag, is the input.

WHAT THIS COSTS, STATED RATHER THAN GLOSSED

The tag is signed by the CI key, not the maintainer's YubiKey — which cannot sign
inside a runner, and whose non-exportability is exactly what makes it worth
trusting. The chain still ends in hardware because the CI key is certified by it.

But no signature on a release now asserts that a person authorised it. That claim
moves entirely to the two gates around publication: `main` accepts only reviewed
pull requests, and publishing requires an approval from a named reviewer on a
protected environment. BRANCHING.md and SECURITY.md said the opposite and are
corrected — a verification instruction that overstates what it proves is worse
than none, because someone acts on it.

ADR 0001 still describes the tag-triggered flow and needs superseding.
STRICTER — the policy is now a gate rather than a promise

CLAUDE.md says "never ship a deprecated or scheduled-for-removal API — treat it
as a blocker, not a warning". Nothing enforced it. The verifier's default failure
level is COMPATIBILITY_PROBLEMS + INTERNAL_API_USAGES + OVERRIDE_ONLY_API_USAGES,
so a deprecated usage was reported and the job went green anyway. A rule that
lives only in prose is not a rule; DEPRECATED_API_USAGES is what makes the
sentence true. Verified green today, so it lands with no debt to forgive.

EXPERIMENTAL_API_USAGES is deliberately excluded, and that is a decision rather
than an oversight: DiffTabCleanup uses ProjectCloseListener.projectClosingBeforeSave
knowingly, because it is the only hook that runs before the workspace is written.
An experimental API is acceptable with a reason. A deprecated one is not, because
it has an announced removal and this plugin must keep working across 251 → 262.

LEANER — four measured duplications, none of them a weakened gate

- Every commit ran the pipeline TWICE. `push` on topic branches and
  `pull_request` both fire, and the concurrency group keyed on `github.ref`
  differs between them (refs/heads/x vs refs/pull/N/merge), so neither cancelled
  the other. Keyed on the commit now: same SHA, same group, duplicate cancelled.
- Superseded runs were only cancelled for pull requests. Three pushes in a row
  left three full pipelines racing, each spending ten minutes downloading IDEs
  for a commit that had already been replaced.
- The JVM suite ran twice. `koverVerify` depends on `:test`, so putting it in the
  `Static analysis` job re-ran the entire suite on a second runner with a cold
  cache. Coverage is a property OF a test run and now shares its job.
- The plugin was built twice. `verifyPlugin` already produces the distributable;
  `Build plugin` built its own. That was not only wasteful but subtly wrong — the
  bytes being asserted were never the bytes that were verified. It now downloads
  the verified artifact.

Job DISPLAY NAMES are unchanged, deliberately: a ruleset references a required
check by its name, so renaming one does not fail the gate — it silently stops
applying it. No rulesets need reapplying.

Also gives drift.yml the concurrency group it was missing (queue, do not cancel:
a half-written drift report is worse than a late one).

Caught while writing this: the SHA I pinned actions/download-artifact to was
invented. Verified against the API and corrected. Pinning by SHA protects nothing
if the SHA is made up.
Nine open Dependabot PRs, each firing the full pipeline — verifier included, ten
minutes and 1.25 GB of IDE downloads apiece — for dependency bumps that are
reviewed in seconds.

Five of them were SECURITY updates (undici, ip-address, fast-uri, hono, postcss:
exactly the `npm audit` findings). Two things about that stream were not
understood when this file was written, and both are documented at the setting now:

  - security updates ignore `open-pull-requests-limit` entirely, which is how five
    arrived under a limit of three;
  - the existing group did not cover them, because `applies-to` defaults to
    version-updates.

So each ecosystem now has an explicit `applies-to: security-updates` group. These
are transitive devDependencies that are never distributed — `npm audit --omit=dev`
reports 0, and the artifact contains zero node_modules entries — so reviewing them
one at a time bought nothing.

github-actions and gradle had no grouping at all. Actions are pinned by full
commit SHA, so a bump is a one-line change per action; grouping them costs no
review fidelity. Gradle groups minor and patch only: a MAJOR keeps its own PR
deliberately, because that is the ecosystem where a bump can hang the headless
suite — the 2.16 -> 2.18 platform-plugin attempt did exactly that, and it deserved
its own run and its own decision.

Syntax verified against GitHub's Dependabot options reference rather than written
from memory, after inventing an action SHA earlier today.
I edited build.gradle.kts and ran verifyPlugin and `help` against it, but not
spotlessCheck — so `Static analysis` failed on spotlessKotlinGradleCheck for a
purely mechanical reason. The formatter is a gate like any other and running a
subset of the gate is the same as not running it.
…ipelines

ci: release from main, enforce zero deprecations, remove duplicate work
Bumps [org.junit:junit-bom](https://github.com/junit-team/junit-framework) from 5.11.4 to 6.1.2.
- [Release notes](https://github.com/junit-team/junit-framework/releases)
- [Commits](junit-team/junit-framework@r5.11.4...r6.1.2)

---
updated-dependencies:
- dependency-name: org.junit:junit-bom
  dependency-version: 6.1.2
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Bumps the actions group with 1 update: [actions/download-artifact](https://github.com/actions/download-artifact).


Updates `actions/download-artifact` from 7.0.0 to 8.0.1
- [Release notes](https://github.com/actions/download-artifact/releases)
- [Commits](actions/download-artifact@37930b1...3e5f45b)

---
updated-dependencies:
- dependency-name: actions/download-artifact
  dependency-version: 8.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
  dependency-group: actions
...

Signed-off-by: dependabot[bot] <support@github.com>
…b_actions/actions-1fa769d870

build(deps): bump actions/download-artifact from 7.0.0 to 8.0.1 in the actions group
Three changes, all aimed at the same thing: the pipeline was spending its time on
work nobody needed.

DEVELOP NO LONGER REQUIRES BRANCHES TO BE UP TO DATE

`strict_required_status_checks_policy` meant every merge into develop invalidated
every other open PR, forcing each to update and re-run the entire suite. With two
grouped Dependabot PRs that is annoying; with five it makes a release impractical,
and the cost is paid on exactly the changes that least deserve scrutiny.

What the setting guards against is real — two PRs that are green apart can break
together — so it is not simply dropped. It is dropped where a second net exists:
CI runs on every push to develop, so a semantic conflict is caught there, on
develop, before anything reaches main. `main` KEEPS the strict policy: one merge
per release, and it is the merge that publishes.

THE VERIFIER RUNS WHERE THE ANSWER MATTERS

~10 minutes and 1.25 GB of IDE downloads, previously on every push to every topic
branch. Now on pull requests and on the protected branches. It remains a required
check, so nothing merges without it. The loss is early detection mid-branch, which
is a genuine cost rather than free savings.

TOPIC BRANCHES WRITE THEIR OWN CACHE

Read-only was blanket-applied to everything but develop and main, so a topic
branch restored the shared cache and saved nothing — every push re-downloaded what
the previous one had already fetched. Read-only now applies to pull requests from
forks only.

Verified against GitHub's cache documentation rather than assumed: "Workflow runs
cannot restore caches created for child branches or sibling branches", and a cache
created on a pull request is written to the merge ref and "can only be restored by
re-runs of the pull request". A topic branch cannot reach what develop reads back;
the isolation is the platform's, not ours.

NB the ruleset change needs ./scripts/apply-rulesets.sh to take effect.
~10 minutes and 1.25 GB of IDE downloads per run, previously paid on every
iteration of a branch nobody was about to merge. It now runs only on pushes to
develop and main — not on topic branches, not on pull requests.

This REQUIRED dropping "Plugin verifier" and "Build plugin" from the required
checks on both rulesets, and that is not a detail: a required check whose job
never runs is never reported, so the pull request would wait forever. The two
changes have to move together or the branch becomes unmergeable.

What still covers a release: release.yml re-runs the FULL gate — verifier
included — on the exact tree being published, behind the environment approval, so
nothing reaches the Marketplace unverified. What is genuinely lost is catching a
binary incompatibility at pull-request time rather than after the merge into
develop. A real regression in feedback latency, accepted deliberately.

Dependabot also moves to monthly, and majors are no longer bot-proposed in any
ecosystem: they are where a bump actually breaks something and where CI is not
enough on its own — the platform-plugin 2.16 -> 2.18 attempt passed CI and hung
the headless suite locally, inside the platform's own test fixture. Security
updates are unaffected by either change.
Corrects the previous commit, which moved the verifier in the wrong direction. It
ran it only on PUSHES to develop and main and dropped it from main's required
checks — so a pull request from develop into main, the merge that publishes,
would not have run it at all. That is precisely the door that has to be guarded.

The policy now matches the intent: topic branches and pull requests into develop
run the fast checks (compile, our own JVM and frontend suites, static analysis,
dependency audit) and iterate quickly. Any pull request targeting main also runs
the plugin verifier and the artifact assertions, both required again on main
alongside the two CodeQL analyses.

The verifier is the only thing that catches a BINARY incompatibility across the
declared 251 -> 262 range — compiling against 252 proves nothing about 262, which
is exactly how the 4.4.1 /login regression shipped. Skipping it on a branch is a
latency trade; skipping it on the way to a release would not be.
Branches now run JVM tests and frontend tests and nothing else. Static analysis
and the dependency audit join the plugin verifier behind the develop -> main
door, where the exhaustive gate belongs.

Required checks match, because they have to: a check whose job does not run is
never reported and would block the pull request forever. develop requires the two
suites; main requires all eight.

The costs, named rather than discovered later:

  - A formatting, detekt, ESLint or coverage failure now lands ON develop and is
    fixed by a follow-up commit, instead of being caught in the pull request. That
    happened today with spotlessKotlinGradleCheck, and the PR is what caught it.
  - The dependency audit no longer runs on the Dependabot PR that proposes a bump.
    It runs once that bump is on develop, and again before it can reach main, so
    nothing ships un-audited — the finding just arrives one merge later.

What this buys is the thing that was actually hurting: a branch iterates in about
three minutes instead of thirteen, and nothing is promoted to main without the
full gate, plus release.yml re-running all of it on the exact published tree.
…e/org.junit-junit-bom-6.1.2

build(deps): bump org.junit:junit-bom from 5.11.4 to 6.1.2
serialexperimentslainnnn and others added 12 commits August 6, 2026 02:27
A branch with an open pull request already fires `pull_request` on every push to
it (the `synchronize` event), so keeping a `push` trigger meant two complete
pipelines per commit for identical information. This removes the duplication at
its source instead of relying on the concurrency group to cancel one in time.

The result is the intended shape: a PR into develop runs the two required suites
once, a PR from develop into main runs all eight.

Two things this gives up, recorded because each removes something the current
setup was leaning on:

  - There is no longer a CI run on the push a merge into develop creates. That
    run was the stated justification for dropping the up-to-date requirement on
    develop: two pull requests that are green apart can break together, and
    develop's own run was what would have caught it. It is now caught at the pull
    request into main, where the full gate runs — later, but still before
    anything is published.
  - A branch with no open pull request gets no checks at all, and a pull request
    from a fork is the only path that would ever exercise them for an outside
    contributor.

release.yml is untouched: it carries its own `push: branches: [main]` trigger and
still fires on the merge that publishes.
Points the five toolchain jobs at ghcr.io/serialexperimentslainnnn/cc-ci and drops
the setup-java and setup-node steps, which provisioned inside a container that
already has both. `Build plugin` keeps no container: it only unzips an artifact.

GRADLE_USER_HOME is set per job and MUST match the value in the Dockerfile. If the
two diverge nothing fails — the run simply re-downloads everything the image
already holds, and the image appears to have bought nothing. That silence is the
reason it is stated at the setting rather than assumed.

NOT VERIFIED against a real image: at the time of writing it has not been built or
pushed. Two things have to be true before this can merge, and both fail in ways
that look like something else:

  - the package must be public (or linked to this repo), or every job dies on a
    401 that reads like a wrong image name;
  - the warmed caches must actually be in the image — `docker run --rm IMAGE
    sh -c 'ls /opt/gradle-home/caches'` answers it in seconds.

Also still open: whether gradle/actions/setup-gradle should stay. It restores its
own cache over GRADLE_USER_HOME, so it now layers on top of the baked one. That
may be a useful increment or redundant work; it needs measuring, not guessing.
The package stays private. Each container job authenticates with the GITHUB_TOKEN
the run already has, so there is no new secret to create, store or rotate, and the
credential expires with the job.

`packages: read` is granted per job rather than at the top level, keeping the
default token read-only on everything else. Without it the pull fails with a 401
that reads like a wrong image name rather than a permission problem — which is
exactly the kind of error that gets debugged in the wrong place.

One prerequisite this does NOT remove: the package must be linked to this
repository, or the token has no grant on it. That is done once, from the package
settings, and it is what makes "same account" mean "same permissions" here.
The cleanup step wipes /warmup, and `npm ci` had installed node_modules inside it
— so the image built the frontend dependencies and deleted them moments later.
The warm-up looked like it worked and bought nothing: CI would re-download the
whole tree on every run, silently, because nothing fails when a cache is missing.

Setting npm_config_cache moves the reusable part to /opt/npm-cache, which the
cleanup does not touch. node_modules stays disposable, and that is correct
independently of this bug: it must match the package-lock.json of the commit CI
checks out, not the one that happened to be current when the image was cut.

Found by a question about what that `rm -rf` actually deletes, which is a better
review than reading the line I had just written myself.
The image was already pulled by most jobs; the remaining ones provisioned
their own JDK, Node and Gradle cache and so ran on a toolchain nothing else
had used. Now `Build plugin`, the protocol-drift check and the release gate
use it too, which is the point of having built it.

`gradle/actions/setup-gradle` is removed everywhere rather than set to
read-only, because it was not doing the job it appeared to be doing. The warm
GRADLE_USER_HOME measures 31 GB — 23 GB of extracted IDE transforms under
caches/9.5.1 and 7.7 GB of downloaded IDE artifacts under modules-2 — and an
Actions cache entry is capped at 10 GB per repository. It could only ever have
stored a fraction, evicted it, and re-downloaded the rest next run. The image
has no such ceiling. The trade is explicit and worth stating: refreshing what
CI has cached is now a deliberate rebuild-and-push, not something that drifts
between runs.

Also removes the verifier's `Free disk space` step. Inside a container those
paths are the IMAGE's, not the runner's, so it had been freeing nothing while
looking like this job's safety margin. The margin now comes from the IDEs
being baked: nothing is downloaded or extracted at verify time.

Recorded because it is the failure everyone hits once: the private package
must be granted Read access to this repository in its own settings. The
`packages: read` permission widens what the token may ASK for; it does not
authorise it against a package the repo was never linked to, and without the
link the pull fails with a bare `denied` that reads like a wrong image name.

Not containerised, deliberately: `publish`, which holds the Marketplace token
and the signing key and is not a test, and CodeQL, which is weekly and is
where a container breaks quietly.

Not verified: that the image builds with no network at all. The one attempt
failed on uid mapping, which says nothing about CI, where the container runs
as root.
…code-for-jetbrains into feature/update-pipelines
The CI image was 38.1 GB, 29.1 GB of it the IDEs `verifyPlugin` downloads. Every job in ci.yml
pulls its own copy on its own runner, so `Initialize containers` measured 5m37s on a job whose
actual work is an 8-second vitest run, and 38 GB on a runner with ~25-30 GB free on the root
volume was also flirting with `No space left on device`.

The verifier is the only consumer of those IDEs and it runs only on a pull request from develop
into main, so it now downloads them when it runs. The set also moves — the verifier resolves from
the EAP/RC channels — so the baked copies stopped matching on JetBrains' release schedule, not
ours.

The warm-up itself was not warming anything. `dependencies --configuration compileClasspath`
resolves dependency metadata; it never triggers the artifact transform that EXTRACTS the IntelliJ
Platform, which is where the GB are. Measured: it leaves caches/*/transforms at 179 MB with no
extracted IDE in it. What warmed the platform was the `verifyPlugin` step, so removing that step
alone would have shipped a cold image that re-resolved the platform in every job — and the
`> /dev/null 2>&1 || true` meant nothing would have said so. It is now `testClasses`,
unredirected and without `|| true`, so a failed warm-up fails the image build.

Verified with the network disabled inside the image: `testClasses` compiles in 33s and the 84
frontend tests pass from the baked npm cache. 38.1 GB -> 8.07 GB, 3.61 GB compressed.

Also in this change:

- .dockerignore: the tree reaching `COPY` goes from 2.09 GB to 3.2 MB (node_modules, build/,
  .git). `node_modules` is now removed in the same layer `npm ci` creates it, since a later `rm`
  hides space rather than reclaiming it.
- python3 is installed explicitly. It was arriving transitively through dnf-plugins-core, which
  was itself unused — the Adoptium repo file is written with printf, not dnf config-manager, and
  curl is already in the base image — and bin/fake-claude, the stand-in the integration tests
  drive a real ClaudeSession against, is a `#!/usr/bin/env python3` script. Dropping the unused
  package without naming python3 would have broken the integration suite.
- One dnf transaction with tsflags=nodocs, git-core instead of git, /usr/share/locale removed.
- codeql.yml: the matrix is split so java-kotlin runs in the image and inherits the warm cache,
  dropping a setup-java that handed it a different JDK from the rest of the pipeline, while
  javascript-typescript stays on a bare runner where it is faster. Both display names are
  unchanged: they are required checks in .github/rulesets/main.json, and a renamed job stops
  applying its gate silently.
- No `:latest` anywhere; the image is `:base`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The release ran as: build, publish to the Marketplace, then stamp a tag on what had already gone
out. It now runs as the sequence the process actually describes — merge to main, cut and sign the
tag, open the GitHub Release, then build and publish FROM that tag.

Building from the tag rather than from `main` is the part that changes behaviour: `main` is a
moving ref, so a merge landing between the `guard` job and the `publish` job was silently included
in a release named after a different tree. `publish` now checks out `refs/tags/vX.Y.Z`.

Cutting the tag first used to be unsafe, and the comment saying so was right at the time: a tag
could exist for a version that was never published, and published tags are immutable here. What
makes it safe now is that the whole job sits behind the `marketplace` environment, so nothing —
including the tag — happens before a human approves. The residual case is a publish that fails
after the tag exists; the recovery is re-running the job on that tag, which is why the asset
upload carries `--clobber`.

This is deliberately NOT two workflows chained by the tag push. A tag pushed with the GITHUB_TOKEN
does not create a workflow run (the recursion guard), so chaining would need a PAT, a GitHub App
or a deploy key — a long-lived write credential — to buy an ordering one run already achieves.

`buildPlugin signPlugin publishPlugin` stays a single Gradle invocation. `publishPlugin` uploads
the signed archive only if `signPlugin.didWork` and falls back to the UNSIGNED one otherwise, so
splitting it to fit the new ordering is exactly how an unsigned plugin ships unnoticed.

The GitHub Release is created as a draft and undrafted once the artifacts are attached: created
final and empty, its download links would 404 for the length of the build.

Out of band but part of the same path: the `marketplace` environment's deployment branch policy
allowed only `tag: v*.*.*`, and the run's ref on the primary path is `refs/heads/main` — the
deployment would have been rejected before even asking for the approval. `main` was added to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
codeql.yml already triggers on every pull request into develop, so both jobs were running there and
nobody was obliged to read the result. Requiring them changes only that.

Affordable in a way the other main-only checks are not: no extra run is created. And a SAST finding
is the class of defect worth catching before the merge rather than at the release door, where it
arrives mixed in with everything else that landed on the branch since.

The contexts are the jobs' DISPLAY names. Renaming a job in codeql.yml does not fail this gate — it
silently stops applying it, which is why the names are duplicated in a comment beside them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
apply-rulesets.sh deleted the key named exactly `_comment`, so the moment a second annotation was
needed in one object the natural name — `_comment_codeql` — sailed through the filter and reached
the API, which rejected the whole ruleset with a bare 422 naming no property.

Observed, not hypothetical: it is how the CodeQL required check failed to apply. The failure reads
as "the ruleset is wrong" rather than "a comment leaked into the payload", which is the expensive
part. Now every key with the `_comment` prefix is stripped, so annotating a block twice is safe.

The key added in the previous commit is renamed back to the convention as well.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A release is a claim that develop is a finished state. An open pull request from Claude or from
Dependabot contradicts it: the change was meant to be in this release and sits one click away from
being in it. Merging past it does not lose the work — it ships a version whose CHANGELOG was
written as if the work had landed. For Dependabot it also means releasing with a known dependency
update unmerged, the one class of pending change an advisory gets written about.

A status check rather than a ruleset entry because it cannot be a ruleset entry: rulesets speak of
checks, signatures and approvals, and have no vocabulary for "no other pull request exists".
main.json requires this job by DISPLAY name, like every other gate.

The author match is anchored, not a substring, so a human whose username contains "claude" is not
caught by a release gate. It covers both renderings, since which one appears depends on how each
integration is installed: `app/<slug>` for an app, `<name>[bot]` for a bot user.

The cost is stated in the workflow rather than left to be discovered: Dependabot's resting state is
"has something open", so draining that queue becomes a release step. That is the intended trade,
and if the gate starts being routinely in the way the answer is to merge Dependabot more often, not
to widen the filter.

NB the check cannot be required until this job exists on develop — a required check that never
reports blocks the pull request forever. Merge first, run scripts/apply-rulesets.sh second.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One image served every job, and an image is a cost paid PER JOB: each one pulls its own
copy onto its own runner. So `Frontend tests` — whose work is an 8-second vitest run —
spent 1m05s pulling a JDK, a Gradle distribution and 3.4 GB of extracted IntelliJ
Platform it never opened.

There are now two, named for what they carry and pinned by version, not a floating tag:

  node-test:v1.0.0   462 MB. Node, npm, warm npm cache. For `Frontend tests` and
                     `Dependency audit`.
  jvm-test:v1.0.0    8.08 GB. The above plus the JDK, Gradle and the extracted platform.
                     For every job that runs Gradle: JVM tests, Static analysis, Plugin
                     verifier, CodeQL (java-kotlin), drift, and the release gate.

jvm-test is built FROM node-test, so it is not a second copy — the registry stores the
shared layers once — and it carries Node deliberately: `Static analysis`, `drift` and the
release gate each run Gradle AND npm in one job. Splitting those would add a whole extra
image pull, which is the cost this change exists to remove.

Two jobs now pull NOTHING. `Build plugin` downloads an artifact and runs `unzip`, `grep`
and `ls` — it does not build anything despite the name, and it was pulling GB to do it.
The bot-PR gate only calls `gh`. Both run on a bare runner.

Why the pull cannot simply be cached, since it is the obvious first idea: a `container:`
job pulls in `Initialize containers`, which runs BEFORE the job's first step, so there is
no point at which an `actions/cache` step could run first — and every job starts on a
fresh runner with no shared layer cache. Restoring a tarball instead is slower, not
faster: it moves the same bytes from a store further away than ghcr and capped at 10 GB
per repository. The only lever is how much each job downloads, which is what this does.

The images also run `dnf upgrade --refresh` before installing. That makes the build
non-reproducible, which is acceptable here precisely because the tag is explicit: what CI
runs is frozen at v1.0.0, and bumping it is the deliberate act that moves it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Since `CodeQL (java-kotlin)` moved into a container, every run logs:

  ERROR: ld.so: object '.../${LIB}_${PLATFORM}_trace.so' from LD_PRELOAD
         cannot be preloaded (cannot open shared object file): ignored

`$LIB` and `$PLATFORM` are glibc dynamic string tokens that ld.so expands at load time, so one
variable covers several ABIs. Measured locally rather than assumed: Fedora 44's loader expands
them to `lib64_x86_64` and loads the file happily; Ubuntu 24.04's does not resolve that form at
all. CodeQL is built and tested on Ubuntu runners, so the shipped filename matches Ubuntu's
expansion and Fedora asks for a name the bundle does not contain.

The step symlinks the name Fedora asks for onto the 64-bit tracer that is actually shipped, and
prints the directory listing first — that `ls` is the evidence the layout still matches, and it is
deliberately not guarded with `|| true` so a future CodeQL release that moves these files fails
loudly instead of quietly reverting to the current behaviour.

Whether the message was ever more than noise is NOT established, and this commit does not claim it
was: github/codeql-action#1113 records the same line as harmless with the real failure elsewhere.
It is removed because a permanent ERROR in a security gate's log trains you to skim past the one
that matters — and this gate has just become a required check on develop. The same run will settle
it: if extraction was actually being skipped, the build step's behaviour changes with the symlink
in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous attempt globbed the tool cache:

  ls -d /__t/CodeQL/*/x64/codeql/tools/linux64

and that matched TWO directories — the runner image ships one CodeQL bundle and the action had
downloaded another, so 2.26.1 and 2.26.2 sat side by side. The variable held two newline-separated
paths and every command after it failed:

  ls: cannot access '.../2.26.1/...'$'\n''.../2.26.2/...*_trace.so': No such file or directory

Picking one by sort order would have been a nicer-looking guess. LD_PRELOAD already names the exact
file the loader will be asked for, so the whole inference disappears: the step now reads it, takes
its dirname, and substitutes the tokens the way Fedora's ld.so resolves them.

The substitution is on the value read from the ENVIRONMENT, where `${LIB}` and `${PLATFORM}` are
literal characters that a shell assignment does not re-expand. Verified with a literal env value
rather than assumed — the first test of it was wrong (it built the string with double quotes, so
bash expanded both tokens to empty before the substitution ever ran) and looked like a real failure.

It also fails loudly if LD_PRELOAD is unset, since that means tracing was never initialised and this
step is patching a problem that no longer exists in the shape it was written for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`Initialize CodeQL` runs inside the container and writes to $GITHUB_ENV:

  LD_PRELOAD=/__t/CodeQL/<v>/x64/codeql/tools/linux64/${LIB}_${PLATFORM}_trace.so

`/__t` is the name the tool cache has INSIDE the container; on the host the same directory is
/opt/hostedtoolcache — which is precisely what this job exported back when it ran on a bare runner,
and why the message never appeared there. But $GITHUB_ENV is consumed by the runner process, which
lives on the HOST, so its helpers start with an LD_PRELOAD naming a path that does not exist from
where they stand, and ld.so logs "cannot be preloaded ... ignored" on every step.

The fix makes one string valid from both namespaces: symlink /opt/hostedtoolcache to /__t inside the
container, and rewrite the variable to use it. Verified inside the real jvm-test image before being
written here — the rewritten path resolves and the real tracer loads through it.

Three earlier hypotheses were wrong and are recorded in the workflow so nobody re-runs them:

  - Not a missing Fedora package. The bundle ships the full matrix (lib/lib64/lib32/x86_64-linux-gnu
    times x86_64/haswell/i686/xeon_phi) and lib64_x86_64_trace.so is present in the container, 0755.
  - Not a glibc difference. Fedora 44's loader expands the tokens and loads the real tracer fine in
    this exact image: AT_PLATFORM x86_64, every dependency satisfied.
  - Not a broken database. The tracing that matters happens inside the container, where the path was
    always valid — which is why the scan succeeded throughout.

This supersedes the symlink patch from df99ffe and 284233e, which the evidence showed was a no-op:
the file it created already existed.

Residual, stated rather than hidden: the ERROR still appears once at the start of the rewriting step
itself, since the new value cannot apply before the step that sets it. The steps that run the
compiler get the corrected value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ipelines

ci: release ordering, CodeQL gate and a release-readiness check
The recovery path the workflow documents did not work. If the publish failed after the tag was cut,
the intended fix was to re-run the job — but `fetch-depth: 0` fetches tags, so a bare `git tag -s`
exited non-zero on the second pass, the job died before reaching `gh release upload --clobber`, and
the release was stuck needing a human to delete a tag this repository treats as immutable. Which is
the exact situation immutability exists to prevent.

Re-running the whole WORKFLOW cannot substitute: `guard` would see the tag on the remote and
correctly report the version as already released, skipping verify and publish entirely. So the job
re-run is the only path, and it has to survive an existing tag. It now detects the ref, verifies the
signature that is already on it, and skips re-cutting.

Also corrects a justification that was simply false. The header claimed the tag is checked out
because `main` is a moving ref that a later merge could slip into the release. `actions/checkout`
defaults to `github.sha`, so every job was already pinned to the triggering commit — building from
the tag is a provenance statement, not a race fix. Both claims were flagged in the review of the
commit that introduced them; this is that follow-up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ag-idempotent

fix(release): make the tag step idempotent for job re-runs
@serialexperimentslainnnn
serialexperimentslainnnn merged commit ec6ba0b into main Aug 6, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant