Repository navigation
Conversation
The chart rendered a Deployment on a single plugin-cache claim. Scaling out meant ReadWriteMany, and replicas sharing a cache corrupt it: unpacking is serialised by an in-process lock only, so two pods missing the same plugin race on its remove-and-rename, and each applies cacheMaxBytes to the same bytes. On ReadWriteOnce the Deployment had to use strategy: Recreate instead, so every upgrade was a gap in service. The workload is now a StatefulSet, there for volumeClaimTemplates rather than identity: each replica gets a claim of its own, plugins-<fullname>-N, or an emptyDir with persistence.enabled=false. Upgrades roll, since a StatefulSet deletes a pod before starting its successor, and pods start in parallel because migrations already take an advisory lock. The headless governing Service publishes gRPC only, so the ServiceMonitor does not scrape each pod twice. - persistence.accessMode accepts ReadWriteOnce or ReadWriteOncePod; a shared mode is refused. - More than one replica, or an HPA allowed more than one, needs config.registry.s3.bucket: every claim starts empty and S3 is what fills it. - autoscaling.* renders an HPA on CPU, memory optional; replicaCount is ignored while it is on. An HPA with no metric, a target without the matching resource request, or minReplicas above maxReplicas is refused. - persistence.retentionPolicy sets the claims' whenDeleted and whenScaled, Retain by default. - persistence.existingClaim and templates/pvc.yaml are gone. Setting existingClaim fails the render rather than mounting an empty volume, and NOTES.txt names the old <fullname>-plugins claim when an upgrade leaves it unmounted. Every new value tolerates being absent, because --reuse-values from an older release carries none of them. tests/values-1.0.4.yaml holds the 1.0.4 values verbatim and render.sh renders the templates against them; a partial autoscaling map is refused by name. The chart README covers scaling, upgrading from the Deployment (including rebinding the old volume when there is no S3) and resizing the immutable claim templates. The service README says why replicas must not share a cache and that every limit is per pod. CI's token-rotation check compares pod UIDs, as a StatefulSet's replacement pod keeps its name. deploy/values-easyp-service.yaml adds values for a two-replica install against S3.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The chart rendered a Deployment on a single plugin-cache claim. Scaling out meant ReadWriteMany, and replicas sharing a cache corrupt it: unpacking is serialised by an in-process lock only, so two pods missing the same plugin race on its remove-and-rename, and each applies cacheMaxBytes to the same bytes. On ReadWriteOnce the Deployment had to use strategy: Recreate instead, so every upgrade was a gap in service.
The workload is now a StatefulSet, there for volumeClaimTemplates rather than identity: each replica gets a claim of its own, plugins--N, or an emptyDir with persistence.enabled=false. Upgrades roll, since a StatefulSet deletes a pod before starting its successor, and pods start in parallel because migrations already take an advisory lock. The headless governing Service publishes gRPC only, so the ServiceMonitor does not scrape each pod twice.
Every new value tolerates being absent, because --reuse-values from an older release carries none of them. tests/values-1.0.4.yaml holds the 1.0.4 values verbatim and render.sh renders the templates against them; a partial autoscaling map is refused by name.
The chart README covers scaling, upgrading from the Deployment (including rebinding the old volume when there is no S3) and resizing the immutable claim templates. The service README says why replicas must not share a cache and that every limit is per pod. CI's token-rotation check compares pod UIDs, as a StatefulSet's replacement pod keeps its name. deploy/values-easyp-service.yaml adds values for a two-replica install against S3.
What and why
Checks that are easy to miss
These are the ones this repository has actually been caught by. Delete the
lines that do not apply.
deploy/— ranbash deploy/charts/easyp-service/tests/render.sh.It reads more than the chart: certificate mounts, mimir's rule paths in
every compose file that mounts its config, the alert-name parity between
the chart and the compose stack, and whether the default image tag is
shaped like one that exists.
the runbooks page in easyp-tech/docs-fumadocs. The heading is the anchor
its
runbook_urlpointsat, so the test above fails without it.
easyp-svc config print --changedstill shows what a deployment actually overrides. The dev configs carry
only differences from the defaults, so a changed default silently changes
what those files mean.
both name them literally, and neither fails loudly when a name goes away.
An empty panel reads as "nothing is happening".
Before merging
Seven checks are required and the branch has to be up to date with
master.gh pr merge --squash --autowill wait for both and merge on its own.