Skip to content

Update stacklok/toolhive to v0.42.0 - #1086

Open
renovate[bot] wants to merge 8 commits into
mainfrom
renovate/stacklok-toolhive-0.x
Open

Update stacklok/toolhive to v0.42.0#1086
renovate[bot] wants to merge 8 commits into
mainfrom
renovate/stacklok-toolhive-0.x

Conversation

@renovate

@renovate renovate Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

This PR contains the following updates:

Package Update Change
stacklok/toolhive minor v0.41.0v0.42.0

After this PR opens, .github/workflows/upstream-release-docs.yml adds source-verified content edits for the new release. For stacklok/toolhive, the same workflow also syncs reference assets (CLI help, Swagger) and regenerates the CRD MDX pages.


Release Notes

stacklok/toolhive (stacklok/toolhive)

v0.42.0

Compare Source

🚀 Toolhive v0.42.0 is live!

AI-tool plugin management goes end to end — thv ai-plugin gains a full CLI, REST API, and registry catalog — and the skills supply chain gets Sigstore signature verification at install, sync, and upgrade time. Alongside that, a large batch of MCP dual-era correctness fixes lands: multiple clients can finally share a stdio server, and vMCP stops flapping between the Modern and Legacy revisions.

⚠️ Breaking Changes

  • Config CRD status fields removedstatus.referencingWorkloads and status.referenceCount (and the References printer column) are gone from all six config CRDs; replace any automation reading them with a workload field query (migration guide)
  • Cedar policy is now evaluated against the mutated MCP request — if you run a mutating webhook together with authorization, policy decisions and audit records can change on upgrade; re-audit your policies against the post-mutation shape first (migration guide)
  • Recovered HTTP panics no longer produce a log line — an unintended regression from the recovery-middleware migration; without Sentry configured a recovered panic is now silent apart from the 500 (migration guide)
  • Go API removals for out-of-tree importerspkg/telemetry/providers was deleted and two long-published optimizerdec constants were removed (migration guide)
Migration guide: Config CRD status fields removed

Who is affected: anyone reading status.referencingWorkloads or status.referenceCount from MCPOIDCConfig, MCPAuthzConfig, MCPExternalAuthConfig, MCPToolConfig, MCPWebhookConfig, or MCPTelemetryConfigkubectl users relying on the REFERENCES column, scripts and GitOps assertions using jsonpath/jq on those paths, Chainsaw/kuttl tests, kube-state-metrics custom-resource-state configs and the dashboards built on them, and Go code reading .Status.ReferencingWorkloads / .Status.ReferenceCount.

MCPWebhookConfig and MCPTelemetryConfig only ever had referencingWorkloads. MCPTelemetryConfig never had a References printer column, so its kubectl get output is unchanged.

Upgrade safety: these were derived values computed from workload specs — the source of truth (spec.*ConfigRef on workloads) is untouched, so nothing unrecoverable is lost. Applying the new schema does not rewrite or reject existing stored objects; residual values stay inert in etcd until each object's status is next written. No storage-version bump, no CRD delete/recreate, no migration job. Deletion protection is unchanged — every config controller still recomputes referrers live at deletion time and sets DeletionBlocked=True with reason ReferencedByWorkloads.

Before
$ kubectl -n toolhive-system get mcpoidcconfig
NAME       SOURCE   VALID   REFERENCES   AGE
my-oidc    inline   True    3            5d
After
$ kubectl -n toolhive-system get mcpoidcconfig
NAME       SOURCE   VALID   AGE
my-oidc    inline   True    5d

To list referrers, query the workloads by their config-ref:

kubectl -n toolhive-system get mcpservers,mcpremoteproxies,virtualmcpservers -o json \
  | jq -r --arg n my-oidc '.items[]
      | select((.spec.oidcConfigRef.name // .spec.incomingAuth.oidcConfigRef.name) == $n)
      | "\(.kind)/\(.metadata.name)"'

The reference paths per config kind, exactly as the operator's own indexers define them:

Config kind Workload kinds tracked Spec paths
MCPOIDCConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.oidcConfigRef.name; spec.incomingAuth.oidcConfigRef.name (vMCP)
MCPAuthzConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.authzConfigRef.name; spec.incomingAuth.authzConfigRef.name (vMCP)
MCPTelemetryConfig MCPServer, MCPRemoteProxy, VirtualMCPServer spec.telemetryConfigRef.name
MCPExternalAuthConfig MCPServer, MCPRemoteProxy spec.externalAuthConfigRef.name, or spec.authServerRef.name when spec.authServerRef.kind == "MCPExternalAuthConfig"
MCPToolConfig MCPServer spec.toolConfigRef.name
MCPWebhookConfig MCPServer spec.webhookConfigRef.name

Note kubectl --field-selector will not work for these paths — the operator's indexes are controller-runtime cache indexes, not API-server field selectors. Use -o json | jq or -o custom-columns.

Migration steps
  1. While still on v0.41.x, snapshot anything you may need: kubectl get mcpoidcconfigs,mcpauthzconfigs,mcpexternalauthconfigs,mcptoolconfigs,mcpwebhookconfigs,mcptelemetryconfigs -A -o json > /tmp/thv-config-refs-pre-0.42.json
  2. Grep your automation for referenceCount, referencingWorkloads, and the References/REFERENCES column — shell scripts, kubectl wait --for=jsonpath=, Chainsaw/kuttl assertions, Argo CD/Flux health checks, kube-state-metrics configs, Grafana panels, Kyverno/Gatekeeper rules.
  3. Rewrite each hit with the query for that config kind from the table above. For "is this config still in use?" checks, prefer the condition: kubectl -n NS get mcpoidcconfig my-oidc -o jsonpath='{.status.conditions[?(@.type=="DeletionBlocked")].message}'
  4. helm upgrade the operator-crds chart, then the operator chart. No pre/post hooks needed.
  5. Verify: kubectl -n toolhive-system get mcpoidcconfig shows NAME SOURCE VALID AGE, and deletion of a referenced config still leaves it with DeletionBlocked=True.
  6. Go consumers: drop .Status.ReferencingWorkloads / .Status.ReferenceCount reads. The WorkloadReference type (Kind, Name) is still exported if you want to keep your own list shape.

PR: #​5631 — completes the cleanup tracked in #​5607

Migration guide: Cedar policy now sees the post-mutation request

Who is affected: only workloads configured with at least one mutating webhook and either Cedar authorization or any consumer of audit / telemetry / usage metrics. Both are shipped, supported, non-mutually-exclusive configurations — thv run --webhook-config <file with a mutating: entry> --authz-config <file>, or MCPWebhookConfig.spec.mutating in the operator. Workloads with no mutating webhook see zero change; the republish is gated on the body actually having changed.

What was wrong: ParsingMiddleware parses the request body once and refuses to parse again. The mutating webhook replaced r.Body but passed the request through unchanged, so Cedar evaluated policy against the tool name and arguments that arrived while the backend executed the ones that ran. The audit half was reachable in the default configuration: the event type and target.name resolve through the parsed-request holder regardless of includeRequestData (which defaults to false), so the audit trail named a request that never executed. Telemetry and usage metrics drifted the same way.

Security framing, stated precisely: before v0.42.0, a client could reach a tool or argument set Cedar would have denied by sending a permitted request shape that the webhook rewrote into a forbidden one. A second bug narrowed this in practice: r.ContentLength was not refreshed alongside r.Body, so a mutation that shrank the body failed at the reverse proxy and one that grew it was truncated into invalid JSON. The bypass was live for length-preserving rewrites — which is exactly case/format normalization, and a webhook can pad JSON whitespace to hold length constant. That stale Content-Length is also fixed here.

Before
client request  ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
                                                │                      │
                                          Cedar reads ◄────────────────┘  (pre-mutation)
                                          audit reads                     backend runs post-mutation
After
client request  ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
                                                                        │
                                                       RepublishParsedMCPRequest (body changed)
                                                                        │
                                          Cedar reads ◄────────────────┘  (post-mutation)
                                          audit reads                     backend runs post-mutation
Migration steps
  1. Check whether you set --webhook-config with a mutating: entry (or MCPWebhookConfig.spec.mutating). If not, stop — no action needed.
  2. Read each mutating webhook's patch and enumerate what it rewrites: the JSON-RPC method, params.name, and/or params.arguments.
  3. Re-check your Cedar policies against the post-mutation shape — resource names (MCP::Tool::"<name>") and every when { context.arg_* } clause. Policies that were passing only because they never saw the rewrite will now deny, and vice versa.
  4. Update SIEM rules, dashboards, and saved queries keyed on audit type or target.name — for mutated requests those values change on upgrade.
  5. Expect two new fail-closed responses replacing what previously reached the backend: 400 if a webhook rewrites a single request into a JSON-RPC batch, and 500 if a webhook emits a body that is not a valid JSON-RPC request.

Gaps this deliberately does not close, all documented rather than fixed:

  • With includeRequestData: true, the recorded request payload is still the pre-mutation body (audit reads r.Body before the webhook), so event type/target name are post-mutation while the payload is not.
  • After a webhook renames a tool, the Mcp-Method/Mcp-Name headers forwarded to the backend still name the original tool. A conformant Modern backend rejects the mismatch, so it fails closed — but a mutating webhook should not rename tools on the Modern path.
  • The tool-call filter and rate limiter run outside ParsingMiddleware and still decide against the request as received, so --tools filtering remains bypassable by a webhook rename. Tracked in #​6134.

PR: #​6136 — Fixes #​6133

Migration guide: Recovered panics are no longer logged

This is an unintended regression, not a design decision. It is called out here because it costs you diagnostics silently, and a one-line fix is expected in a patch release.

Who is affected: any operator who relies on ToolHive's logs to diagnose a recovered HTTP panic — including log-based alerts, log-derived metrics, and support bundles. Everyone running without Sentry configured (the default) is affected most.

What changed: pkg/recovery became a thin shim over toolhive-core/recovery. Core's Middleware recovers panics silently unless a logger is injected via WithLogger, and ToolHive's shim passes only WithPanicHandler. The OTel span error recording and Sentry issue reporting are genuinely preserved — same span status (codes.Error, "panic recovered"), same sanitization, same raw value to Sentry, same ordering — but the slog.Error line and its stack trace are gone, and no other middleware picks them up.

Before (v0.41.0)
time=... level=ERROR msg="Panic recovered: runtime error: index out of range [3] with length 2
Stack trace:
goroutine 42 [running]:
runtime/debug.Stack()
..."
After (v0.42.0)
(no log output — the client receives 500 Internal Server Error and nothing is recorded locally)
Migration steps
  1. If you have alerts or log-based metrics matching Panic recovered, they will stop firing. Do not interpret the silence as "no panics" — re-point them at the 500-response rate or at Sentry until the log line returns.
  2. Configure Sentry if you have not already; ReportPanic still sends the raw panic value, so panics remain visible as Sentry Issues with full context.
  3. Traces are unaffected — the request span still carries RecordError plus an error status, so OTel-based panic detection keeps working.
  4. When the fix lands, the restored line will be structured (msg="panic recovered" with panic, method, path, and stack attributes) rather than the old single formatted string, so write any new log parser against that shape.

PR: #​6145

Migration guide: Go API changes

Who is affected: only out-of-tree Go code importing ToolHive packages. No CLI, REST API, or CRD surface changes here, and no in-tree caller is affected.

pkg/telemetry/providers was deleted (#​6146)

The package and its /otlp and /prometheus subpackages were removed and consumed from toolhive-core instead. The graduation is verbatim — every non-test file is byte-identical apart from two self-referential import paths — so all 12 options (WithServiceName, WithServiceVersion, WithOTLPEndpoint, WithHeaders, WithInsecure, WithCACertPath, WithTracingEnabled, WithMetricsEnabled, WithSamplingRate, WithEnablePrometheusMetricsPath, WithCustomAttributes, WithExtraSpanProcessors) plus NewCompositeProvider, ProviderOption, and CompositeProvider keep identical names and signatures. Nothing about emitted telemetry changes — resource attributes, service-name defaulting, OTLP exporter/TLS config, and Prometheus exporter registration all behave as before.

Before

import "github.com/stacklok/toolhive/pkg/telemetry/providers"

After

import "github.com/stacklok/toolhive-core/telemetry/providers"
Two optimizerdec constants were removed (#​6175)

pkg/vmcp/session/optimizerdec no longer exports CallToolArgToolName or CallToolArgParameters. Both have been part of the published API since v0.15.0. They existed to read the call_tool target out of a raw arguments map, a pattern that is now known-unsafe: encoding/json falls back to case-insensitive field matching, so a map index and a struct decode resolve different key sets.

Before

toolName, _ := args[optimizerdec.CallToolArgToolName].(string)
params, _ := args[optimizerdec.CallToolArgParameters].(map[string]any)

After

// Decode with the same call both dispatch sites use, so key matching cannot diverge.
in, err := schema.Translate[optimizer.CallToolInput](args)
registry.Provider gained three methods (#​6135)

ListAvailablePlugins(), GetPlugin(namespace, name), and SearchPlugins(query) were added to the interface. Implementations that embed registry.BaseProvider pick up no-op defaults and need no change; anything satisfying the old method set directly will no longer compile.

Migration: embed registry.BaseProvider in your provider struct, or implement the three methods.

🔄 Deprecations

  • pkg/audit's MCP event constants, LevelAudit, and NewAuditLogger are now transitional aliases for github.com/stacklok/toolhive-core/audit and will be removed once the migration's cleanup wave rewrites imports per subtree — prefer the toolhive-core/audit symbols in new code (#​6148)

🆕 New Features

  • Manage plugins for AI coding tools with the new thv ai-plugin command group — build, validate, push, install, list, info, uninstall, plus local build management via builds and builds remove — targeting Claude Code and Codex (#​5782)
  • The same plugin surface is available over REST at /api/v1beta/plugins (10 endpoints) with a matching Go HTTP client in pkg/plugins/client, so the CLI, API, and external tooling share one contract (#​5782)
  • thv ai-plugin install <name> now resolves a plain name against the configured registry instead of failing with a 404 hint, and new catalog routes let you browse and search plugins in a registry (#​6135)
  • Project-scoped skill installs now verify Sigstore signatures before anything is extracted or recorded, recording the signer identity as lock-file provenance: on first use and rejecting unsigned artifacts unless you pass --allow-unsigned (#​6129)
  • thv skill sync re-verifies each managed skill's stored Sigstore bundle offline against the lock file's recorded identity, treating a failed re-verification as drift so a CI gate catches signature changes exactly like content changes (#​6131)
  • thv skill upgrade refuses to move a skill to an artifact signed by a different identity — or to an unsigned one — reporting signer-change-blocked unless you explicitly rotate trust with --allow-signer-change (#​6132)
  • Git-installed skills get full gitsign commit-signature verification, with the chain of trust checked against embedded Fulcio roots and no network access (#​6121, #​6091)

The skills signing features above are all behind the experimental TOOLHIVE_SKILLS_LOCK_ENABLED gate and apply only to project-scoped installs. With the gate unset, thv skill install behaves exactly as in v0.41.0. Note that git (gitsign) provenance is recorded as provisional: true because the embedded Rekor transparency-log proof is not yet validated — signing time is checked only against the Fulcio certificate's own ~10-minute validity window. OCI provenance is not provisional.

🐛 Bug Fixes

  • Multiple MCP clients can now connect to a single stdio MCP server through ToolHive — the first handshake is cached and replayed instead of every client after the first getting duplicate "initialize" received, which also unblocks vMCP aggregating stdio backends (#​6153)
  • A client that retries initialize on a live connection behind the transparent proxy now receives a fresh session instead of a hard failure, because the proxy no longer forwards a session ID on initialize (#​6152)
  • vMCP gateways aggregating a dual-era backend such as github-mcp-server v1.6.0 no longer oscillate between the Modern and Legacy revisions and fail roughly half their health checks — a Modern promotion must now win a confirming server/discover probe rather than trusting the negotiated version alone (#​6158)
  • A vMCP backend redeployed from a hint-lying Legacy server to a genuinely Modern one now corrects its reported MCP revision within ~5 minutes instead of staying Legacy until the pod restarts (#​6185)
  • vMCP now relays backend log and progress notifications from Modern (2026-07-28) backends to the downstream client, opting in through the per-request io.modelcontextprotocol/logLevel _meta key that replaced the removed logging/setLevel RPC (#​6140)
  • vMCP clients on the Legacy revision now receive non-reserved backend _meta (trace ids, custom fields) on resources/read results, matching what the Modern path already delivered (#​6180)
  • A vMCP pod with the optimizer enabled no longer permanently loses its tool index while continuing to report itself healthy — the in-memory SQLite database is now pinned alive by a dedicated connection, so one cancelled request can't destroy it for the life of the process (#​6157)
  • find_tool's tool_keywords input now actually affects results instead of being decoded and dropped, and it drives the lexical BM25 arm while tool_description drives semantic matching (#​6124)
  • call_tool now accepts the common LLM malformation where tool_name is nested inside parameters, and a genuinely missing tool_name produces an error that states the expected shape and lists the parameter names received (#​6150)
  • On Windows, the discovery directory and server.json under %LOCALAPPDATA% are now protected with an explicit DACL granting only the ToolHive user and SYSTEM, and are ownership-validated before being trusted — POSIX mode bits are advisory on NTFS, so any local account with Modify could previously rewrite the npipe:// discovery URL and redirect the next MCP client (#​5951)
  • The authorization middleware now resolves a call_tool target through the same decoder dispatch uses, closing three case-sensitivity divergences that could skip a policy check or drop arguments (#​6175)

🧹 Misc

  • pkg/telemetry/providers (~2,900 LOC) is deleted in favour of the verbatim graduation in toolhive-core, with no change to emitted telemetry (#​6146)
  • pkg/recovery becomes a thin shim over toolhive-core/recovery, keeping ToolHive's OTel and Sentry wiring through a panic-handler hook (#​6145)
  • MCP histogram buckets are sourced from toolhive-core's semconv preset instead of a local literal; the boundaries are unchanged (#​6144)
  • Audit event constants, LevelAudit, and NewAuditLogger become aliases over toolhive-core/audit with byte-identical values, so the audit wire format is untouched (#​6148)
  • Pinned a regression test for vMCP elicitation failing fast when the client advertised the capability but holds no standalone SSE stream, and documented both delivery constraints (#​6182)
  • Fixed a port-selection TOCTOU race that flaked e2e tests under sharded CI by having the OIDC and LLM gateway mocks hold their own listener from construction (#​6142)
  • Pinned the ida-pro-mcp e2e image by digest after an upstream rebuild pulled in the breaking mcp Python SDK 2.0.0, and added test/e2e/images/** to the lifecycle suite's trigger filter so an image change can no longer skip the tests that consume it (#​6159)
  • Pinned the mcp-server-time e2e image by digest for the same upstream breakage, unblocking the proxy suites (#​6160)
  • Fixed the operator integration suites' timeout waiting for process kube-apiserver to stop flake — 32 of the job's last 51 failures — by awaiting manager shutdown before tearing down envtest (#​6179)

📦 Dependencies

Module Version
github.com/stacklok/toolhive-core v0.0.35 → v0.0.38
github.com/stacklok/toolhive-catalog v0.20260804.0
github.com/tailscale/hujson b80ff77
coverallsapp/github-action 8d6379e
github/codeql-action f205ea1
anthropics/claude-code-action v1.0.183

toolhive-core was bumped across #​6144, #​6146, and #​6180 rather than by a dependency PR; v0.0.38 also carries transitive bumps to aws-sdk-go-v2, go-containerregistry, moby/client, prometheus, and otel.

👋 Welcome to our newest contributor: @​Tanguille 🎉

Full commit log

What's Changed

New Contributors

Full Changelog: stacklok/toolhive@v0.41.0...v0.42.0

🔗 Full changelog: stacklok/toolhive@v0.41.0...v0.42.0


Configuration

📅 Schedule: (in timezone America/New_York)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

Rebasing: Never, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.


Docs update for toolhive v0.42.0

At a glance

Upstream stacklok/toolhive v0.41.0v0.42.0
Hand-written changes 2 commit(s)
Reference assets refreshed (separate commit)
Gaps 0
Release contributors 8 auto-assigned (see sidebar)
Action required Spot-check skill-authored prose for accuracy

Summary of changes

  • Added AI-tool plugins guide at docs/toolhive/guides-cli/ai-plugins.mdx for
    the new thv ai-plugin surface (build, validate, push, install, list, info,
    uninstall, builds), including the manifest format, Claude Code / Codex
    install paths, and troubleshooting.
  • Added a sidebar entry for the new AI-tool plugins guide in sidebars.ts.
  • Updated docs/toolhive/guides-cli/skills-management.mdx with three new
    sections covering the experimental lock file: pin-and-reconcile with
    thv skill sync, upgrades with thv skill upgrade, and Sigstore signature
    verification (--allow-unsigned, --allow-signer-change, coverage
    differences between OCI and Git installs). Added a matching troubleshooting
    entry.
  • Added "Ordering with Cedar authorization and audit" to
    docs/toolhive/guides-cli/webhooks.mdx documenting that Cedar policies,
    audit events, telemetry, and usage metrics see the post-mutation request,
    plus the new 400/500 fail-closed responses and the tool-filter and header
    gaps to be aware of.
  • Swept the removed config-CRD status fields
    (status.referencingWorkloads, status.referenceCount, and the
    REFERENCES printer column) out of three K8s guides and the
    MCPAuthzConfig CRD intro, replacing them with workload-side jq queries
    where a "which workloads reference this?" pattern was needed. Updated the
    CRD intro at the source (scripts/lib/crd-intros.mjs) and synced the
    generated mcpauthzconfig.mdx.
  • Added a v0.42.0 removal note to the referencingServers /
    referencingWorkloads section of docs/toolhive/guides-k8s/migrate-to-v1beta1.mdx
    so readers of that migration guide learn that both fields are now gone and
    see the current workload-query pattern.

Run cost

Session Turns Cost (USD)
Generation 115 $15.4505
Editorial review 61 $5.5362
Total 176 $20.9867
How this PR was built

Two Claude Opus sessions run per release: a generation pass
(upstream-release-docs skill, 6 phases) followed by a fresh-
context editorial pass (docs-review). Prettier/ESLint
auto-fixes are applied after.

Auto-synced paths — do not hand-edit these in review:

  • static/api-specs/
  • docs/toolhive/reference/cli/ (toolhive only)
  • docs/toolhive/reference/crds/

If a "Gaps needing human context" section is present above,
each entry includes a paste-ready Helper prompt for local
Claude
a reviewer can use to resolve the gap.

Signed-off-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
@renovate renovate Bot added the autogen-docs Pull requests generated by automated release workflows label Aug 5, 2026
@vercel

vercel Bot commented Aug 5, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
docs-website Ready Ready Preview Aug 6, 2026 3:17pm

Request Review

@github-actions
github-actions Bot marked this pull request as draft August 5, 2026 18:46
@claude

claude Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Claude finished @renovate[bot]'s task in 12m 41s —— View job


Upstream release docs: toolhive v0.42.0

Pushed as commit e940606 to renovate/stacklok-toolhive-0.x.

Phases

  • Phase 1: Discovery
  • Phase 2: Deep-dive verification against source (parallel agents, all findings source-cited)
  • Phase 3: Docs audit (displacement + inventory)
  • Phase 4: Implement doc changes
  • Phase 5: Format & lint (prettier and eslint pass; local docusaurus build blocked by sandbox network for MCP metadata fetch — will run cleanly in CI)
  • Phase 6: Commit, push, and SUMMARY.md

Summary of changes

  • Added AI-tool plugins guide at docs/toolhive/guides-cli/ai-plugins.mdx covering thv ai-plugin build / validate / push / install / list / info / uninstall / builds for Claude Code and Codex, plus the .claude-plugin/plugin.json manifest format and troubleshooting. Sidebar entry added after skills-management.
  • Extended skills-management.mdx with three new sections behind the TOOLHIVE_SKILLS_LOCK_ENABLED experimental gate: pin-and-reconcile with thv skill sync, upgrades with thv skill upgrade --allow-signer-change, and Sigstore signature verification (including the OCI-vs-Git provenance-assurance difference). Added a matching troubleshooting entry.
  • Added ordering note in webhooks.mdx — Cedar authorization, audit, telemetry, and usage metrics now see the post-mutation request; document the 400 (rewrite-to-batch) and 500 (invalid JSON-RPC body) fail-closed responses and the tool-filter and Mcp-Method header gaps.
  • Swept the removed config-CRD status fields (status.referencingWorkloads, status.referenceCount, and the REFERENCES printer column) out of authorization-k8s.mdx, auth-k8s.mdx, telemetry-and-metrics.mdx, and the MCPAuthzConfig CRD intro. Replaced "which workloads reference this?" prose with workload-side jq queries. Fixed the source (scripts/lib/crd-intros.mjs) and manually synced the generated mcpauthzconfig.mdx.
  • Added v0.42.0 removal note to the referencingServers section of migrate-to-v1beta1.mdx so readers of that migration guide learn the field they'd be migrating to has since been removed.

Notes for reviewers

  • The release notes claimed git-installed skills record provisional: true in the lock file. Source verification found the field is set on the in-memory verifier result but dropped when marshalling to Provenance in pkg/skills/options.go, so it does not appear on disk in v0.42.0. Documentation avoids that claim and instead notes that Git installs carry less transparency-log assurance than OCI, without promising a provisional field.
  • The generated mcpauthzconfig.mdx was updated by hand to match the new intro in crd-intros.mjs because this environment can't run node scripts/upstream-release/extract-crds.mjs. The next automated release run will regenerate it deterministically from the same source.
    · Commit e940606

- Add guide for the new thv ai-plugin surface (Claude Code, Codex)
- Add sync, upgrade, and signature verification sections to the skills
  management guide behind the experimental TOOLHIVE_SKILLS_LOCK_ENABLED gate
- Document that Cedar authorization now evaluates the post-mutation request
- Sweep removed config-CRD status fields (referencingWorkloads, REFERENCES
  column) out of K8s guides; replace with workload-side queries

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@claude

claude Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Claude finished @renovate[bot]'s task in 6m 29s —— View job


Editorial review of upstream release docs: toolhive v0.42.0

Pushed as commit 48c61ff to renovate/stacklok-toolhive-0.x.

  • Identify files changed by previous commit (7 eligible docs)
  • Run docs-review skill and cross-check claims against upstream source
  • Apply actionable fixes
  • Format and lint (prettier and eslint pass)
  • Commit and push

Findings and fixes

1. ai-plugins.mdx: default OCI tag was incorrect. The guide claimed the default tag is <name>:<version> from the manifest, or the raw digest when version is omitted. Source (pkg/plugins/pluginsvc/build.go L64-76) resolves the default tag as the plugin name alone (not name:version), falling back to the raw digest only when no name is set. The version field is captured in build metadata that thv ai-plugin builds and thv ai-plugin info report — it isn't part of the default tag. Fixed both mentions.

2. migrate-to-v1beta1.mdx: admonition overreached. The prior admonition said "both fields removed from all six configuration CRDs, along with the REFERENCES printer column." Per the upstream migration guide, MCPWebhookConfig and MCPTelemetryConfig only ever had referencingWorkloads (not referenceCount), and MCPTelemetryConfig never had the REFERENCES printer column, so its kubectl get output is unchanged. Reworded to reflect this.

3. skills-management.mdx: exit code 3 was mislabelled. The doc summarised the CI contract as "code 2 = drift, code 3 = a real failure." Source (cmd/thv/app/exitcode.go) defines code 3 as ExitCodePartialFailure (some skills succeeded, others failed) — the generic error is code 1. Expanded to distinguish 1, 2, and 3.

Not changed

The other four files (webhooks.mdx, auth-k8s.mdx, authorization-k8s.mdx, telemetry-and-metrics.mdx) checked out cleanly against upstream source and the style guide.
· Commit 48c61ff

github-actions Bot and others added 2 commits August 5, 2026 19:06
- ai-plugins: default build tag is <name>, not <name>:<version>
- skills-management: distinguish sync exit codes 1, 2, and 3
- migrate-to-v1beta1: reflect that MCPWebhookConfig and MCPTelemetryConfig
  never had referenceCount, and MCPTelemetryConfig never had the printer
  column

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

The `status.referencingServers` field (a plain `[]string`) has been replaced
with `status.referencingWorkloads` (an array of `{kind, name}` objects) on four
The `status.referencingServers` field (a plain `[]string`) was replaced with

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't true anymore as status.referencingWorkloads was removed this release. @copilot

Co-authored-by: ChrisJBurns <29541485+ChrisJBurns@users.noreply.github.com>
ChrisJBurns
ChrisJBurns previously approved these changes Aug 5, 2026
jhrozek
jhrozek previously approved these changes Aug 5, 2026
@samuv

samuv commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

@claude Remove the Sigstore signing work for skills. We’re going to add the documentation when the flag is removed.

@claude

claude Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Claude finished @samuv's task in 3m 54s —— View job


Removed the experimental skills signing docs ✅

Pushed as commit 456d1bb.

  • Reviewed what this PR added to skills-management.mdx
  • Removed the Sigstore signing / lock-file sections
  • Removed the matching troubleshooting entry and reference links
  • Verified the file now matches main, committed, and pushed

What changed

Removed everything this PR added to docs/toolhive/guides-cli/skills-management.mdx, all of which was gated behind the experimental TOOLHIVE_SKILLS_LOCK_ENABLED flag:

  • Pin and reconcile project skills (lock file, thv skill sync)
  • Upgrade project skills (thv skill upgrade)
  • Verify skill signatures (Sigstore/gitsign verification, --allow-unsigned, --allow-signer-change, verification coverage)
  • The Signature verification errors troubleshooting entry
  • The two Related information links to the thv skill sync / thv skill upgrade command references

The file is now byte-identical to main (the PR made no other edits to it), so it's already lint- and Prettier-clean. When the flag is dropped and the feature ships as generally available, the documentation can be re-added.

Left untouched (auto-generated)

The CLI reference pages are auto-generated from the ToolHive source and per the repo's CLAUDE.md must not be hand-edited — they're regenerated on every release sync. These reflect the actual CLI (the commands still exist upstream behind the flag), so I left them as-is:

  • docs/toolhive/reference/cli/thv_skill_sync.md, thv_skill_upgrade.md (added)
  • docs/toolhive/reference/cli/thv_skill.md, thv_skill_install.md (the --allow-unsigned flag / subcommand listing)

If you'd like those removed from the site too before an upstream release makes that happen, that needs an upstream change (e.g. marking the commands Hidden: true) or a human decision — flag it and I can help draft that.
· Branch

The lock-file, sync, upgrade, and Sigstore signature verification
content is gated behind TOOLHIVE_SKILLS_LOCK_ENABLED. Remove it until
the flag is dropped and the feature is documented as generally
available.

Co-authored-by: Samuele V <4377202+samuv@users.noreply.github.com>
samuv
samuv previously approved these changes Aug 6, 2026
aponcedeleonch
aponcedeleonch previously approved these changes Aug 6, 2026
@danbarr

danbarr commented Aug 6, 2026

Copy link
Copy Markdown
Collaborator

@claude the new ai plugins page was added with zero inbound links. This breaks the user journey; consider where this fits in the current journey and make sure the progression is maintained.

Also assess spots that refer to the ToolHive CLI as managing just MCP servers and skills that need to be updated to include plugins. The insertion of the page was too tactical and not holistic.

@claude

claude Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Claude finished @danbarr's task in 3m 4s —— View job


Integrate the AI plugins page into the user journey ✅

Pushed as commit 6997454.

  • Read the new AI plugins page and the CLI section structure
  • Diagnose the journey break
  • Add inbound links / restore forward momentum in the CLI journey
  • Update CLI descriptions that mention only MCP servers and skills to include plugins
  • Format (Prettier clean), lint (ESLint clean), commit, and push

What was wrong

ai-plugins.mdx was listed in sidebars.ts (after skills-management) but nothing linked to it in prose. The intended CLI flow is:

client-configuration → skills-management → ai-plugins → permissions

…yet skills-management's Next steps jumped ahead to the registry, and the CLI index "Where to start" never mentioned plugins. So a reader following the journey would sail right past the new page.

Changes

Restored forward momentum (inbound links):

  • guides-cli/index.mdx - added an "Extending your AI coding tools?" entry to Where to start pointing at the plugins guide.
  • guides-cli/skills-management.mdx - Next steps now leads with the plugins guide (the sibling bundling workflow), matching the sidebar order. The plugins page already linked back to skills, so this closes the loop.
  • guides-cli/client-configuration.mdx - Next steps now links to plugins, since registering a client is what routes a plugin install to the right directory.

Updated CLI-scoped descriptions that named only MCP servers and skills:

  • guides-cli/index.mdx front matter description - now "run and manage MCP servers, agent skills, and AI-tool plugins."
  • sidebars.ts - the ToolHive CLI category description now reads "managing MCP servers, skills, and plugins."
  • toolhive/index.mdx - the CLI edition bullet in the architecture overview now lists AI-tool plugins alongside agent skills.

Scope note

I deliberately left the platform-wide taglines ("run, govern, and connect MCP servers and agent skills" in index.mdx/faq.mdx) and the Registry Server / Cloud UI / Enterprise pages untouched. thv ai-plugin is a CLI-only capability today - the registry and platform don't manage plugins - so broadening those would overstate the current feature surface. If plugins later gain UI/registry support, those spots are the follow-up.
· Branch

The AI-tool plugins guide was reachable only from the sidebar, with
no inbound prose links, so the CLI journey skipped past it. Add
forward links from the CLI index, skills management, and client
configuration pages, and update CLI-scoped descriptions that named
only MCP servers and skills to include plugins.

Co-authored-by: Dan Barr <6922515+danbarr@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

autogen-docs Pull requests generated by automated release workflows

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants