Skip to content

budctl 0.3.1: stop reporting false blockers on clusters that work - #1

Merged
dittops merged 2 commits into
mainfrom
claude/cluster-readiness-check-417585
Sep 14, 2026
Merged

dittops merged 2 commits into
mainfrom
claude/cluster-readiness-check-417585

Conversation

@dittops

@dittops dittops commented Sep 14, 2026

Copy link
Copy Markdown
Member

Why

A budctl check against the tcs-vmware k3s cluster, which Bud was already serving, reported three blockers. None of them was stopping anything. This PR fixes the checks and catalogue entries behind each false finding, keeps the real ones, and bumps the installer pin to 0.3.1 so the release can be tagged from main.

Finding on tcs-vmware Was Cause Now
storage.csi-healthy BLOCK rancher.io/local-path matched every rancher/* k3s image, so a Pending ServiceLB pod counted as a broken driver PASS
egress.install: Altinity charts BLOCK The upstream chart repo was treated as install-time, but ArgoCD installs the packaged OCI chart with the subchart inside INFO (optional)
egress.install: Let's Encrypt BLOCK Checked whatever the TLS answer, with the wrong consequence New egress.acme, skipped unless TLS uses ACME
registry.from-cluster: azurecr BLOCK Registry retired Removed from the inventory
argocd.chart-repo-credential RISK Predicted a 401 without trying an anonymous pull; the charts project allows anonymous pulls PASS
storage.expansion RISK Asked local-path for allowVolumeExpansion, but local-path ignores sizes and has no resizer PASS, with the node's disk named as the ceiling
gpu.operator-functional not verified HAMi preloads a vGPU library needing libdl.so.2; busybox's shell died in the loader PASS on debian:trixie-slim; loader errors are named
Saved report header "cluster unreachable", "budctl —" The interactive view saved before main stamped the report Both paths use Report.Stamp

Unchanged, because they are real: registry.from-cluster for registry.cn-hangzhou.aliyuncs.com (HAMi's default scheduler image), the Helm/ArgoCD ownership overlap on argocd, and 2 nodes where the data stores expect 3.

Other changes

  • egress.proxy no longer treats k3s's NO_PROXY on helm-install pods as evidence of a proxy.
  • charts.classic wording updated now that only dapr is sourced from an upstream repo.
  • README: the GPU probe image and why it is not busybox.
  • Catalogue version 2026-09-15; install.sh pins 0.3.1.

Verification

  • go build, go vet, go test ./... (including the failure-coverage gate), cross-compiles for all four platforms, and sh -n install.sh.
  • The release guard passes locally: install.sh pins 0.3.1, and a binary built with that version reports budctl 0.3.1.
  • New tests cover each fix. The storage, ArgoCD, GPU and report-metadata tests were run against the pre-fix code and fail there.
  • Full rerun against tcs-vmware: 50 pass, 1 block (aliyuncs), 3 risk. Each fixed check passed with evidence: anonymous HEAD bud:1.2.8 returned ok, and the GPU probe printed BUDCTL_DEVICE_PRESENT.

Release

After merge, tag the merge commit v0.3.1 and push the tag. release.yml then tests, builds and publishes.

🤖 Generated with Claude Code

Ditto P S and others added 2 commits September 15, 2026 02:02
A readiness run against a k3s cluster Bud was already serving reported three
blockers, and none of them was stopping anything. This corrects the checks and
the catalogue behind each false finding, and keeps the real ones.

storage
- csi-healthy matched every rancher/* image to rancher.io/local-path, so a
  Pending k3s ServiceLB pod read as a broken storage driver. The vendor label
  is now a last resort, used only when the provisioner has no more specific
  name (driver.longhorn.io, openebs.io/local).
- csi-healthy no longer reports "registered on 0/N nodes" for a provisioner
  that is not a CSI driver.
- expansion no longer asks node-local classes for allowVolumeExpansion:
  local-path ignores the requested size and has no resizer, so the node's disk
  is the ceiling and the suggested patch changed nothing.

egress and charts
- The prometheus-community, bitnami, CloudNativePG, Altinity, Percona,
  SeaweedFS and OpenTelemetry repos are optional: every ApplicationSet installs
  those charts from registry.bud.studio as packaged OCI charts with the
  subcharts inside, so the cluster never fetches them. dapr stays install-time.
- Let's Encrypt moves to a new egress.acme check, skipped unless TLS is
  obtained through ACME, and it states the real consequence (no certificate)
  instead of "the sync stops at the first image or chart".
- egress.proxy no longer treats k3s's NO_PROXY on helm-install pods as a proxy.

argocd
- chart-repo-credential tries an anonymous pull before predicting a 401: a
  charts project served to anonymous clients needs no repository Secret.

gpu
- The functional probe defaults to debian:trixie-slim. HAMi preloads its vGPU
  library into every GPU container and it needs libdl.so.2, so busybox's shell
  died in the loader. A loader error is now named, with the flag to change it.

registry
- budimages.azurecr.io is retired and removed from the inventory.

report
- Reports saved from the interactive view carry the budctl version, catalogue,
  cluster and domain. Report.Stamp is shared by both output paths; the TUI used
  to save before main stamped the report, so the file said "cluster unreachable".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dittops
dittops merged commit 1ffe314 into main Sep 14, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant