diff --git a/MANUAL.md b/MANUAL.md index a60530c..67202dd 100644 --- a/MANUAL.md +++ b/MANUAL.md @@ -9,7 +9,11 @@ A short guide to installing ICC on Kubernetes with this Helm chart. - PostgreSQL reachable from the cluster (one server; ICC uses several databases on it, listed in section 2). - Valkey (or Redis), reachable from the cluster. -- Prometheus, reachable from the cluster. +- Prometheus, reachable from the cluster. It must scrape kubelet + `/metrics/cadvisor` and kube-state-metrics so ICC can read container resource + usage, requests, limits, and Pod labels. kube-state-metrics must export the + `platformatic.dev/monitor` Pod label. +- A Kubernetes resource metrics API (`metrics.k8s.io`) for the ICC HPA. - Prometheus Operator CRDs (`PodMonitor`, `ServiceMonitor`) installed. The chart renders both by default; to install without them, disable the monitors (`watt.monitor.enable` and the per-service `monitor.enable`). @@ -151,7 +155,8 @@ kubectl port-forward -n platformatic svc/icc 8080:80 ```sh helm upgrade platformatic oci://ghcr.io/platformatic/helm \ --version "^4.1.0" -n platformatic \ - -f my-values.yaml -f my-secrets.yaml + -f my-values.yaml -f my-secrets.yaml \ + --wait --timeout 10m helm uninstall platformatic -n platformatic ``` @@ -169,6 +174,10 @@ helm uninstall platformatic -n platformatic capabilities) is below 1.30. Upgrade the cluster. - `no matches for kind "PodMonitor"` (or `"ServiceMonitor"`): the Prometheus Operator CRDs are missing. Install them, or disable the monitors (section 1). +- ICC CPU metrics are empty or zero: confirm Prometheus contains + `container_cpu_usage_seconds_total`, `kube_pod_container_resource_limits` or + `kube_pod_container_resource_requests`, and + `kube_pod_labels{label_platformatic_dev_monitor="prometheus"}`. - ICC pod not ready: check `kubectl logs -n platformatic deploy/icc`; usually the database, Valkey, or Prometheus URL is unreachable. diff --git a/MIGRATING-v3-to-v4.md b/MIGRATING-v3-to-v4.md index ddb8ea0..1267fc3 100644 --- a/MIGRATING-v3-to-v4.md +++ b/MIGRATING-v3-to-v4.md @@ -14,6 +14,9 @@ Test the upgrade in a non-production cluster first. - Kubernetes 1.30 or newer. The chart refuses to install on older clusters. - Helm 3.13.2 or newer. - PostgreSQL, Valkey or Redis, and Prometheus reachable from the ICC Pod. + Prometheus must scrape kubelet `/metrics/cadvisor` and kube-state-metrics. + kube-state-metrics must export the `platformatic.dev/monitor` Pod label. +- A Kubernetes resource metrics API (`metrics.k8s.io`) for the ICC HPA. - A replacement for any Ingress or persistent storage created by chart v3. See [Resources removed or changed](#resources-removed-or-changed). @@ -114,8 +117,11 @@ services: scaler: algorithm_version: v1 - # Preserve the v3 values. The v3 defaults were 300 and 60. + # 300 preserves the chart v3 behavior. Chart v4 defaults to 15. + # With scaler v1, scale-up is blocked for this period after a scale + # operation; scale-down uses a shorter derived cooldown (60-180s). cooldown: 300 + # Preserve the chart v3 default. Chart v4 also defaults to 60. periodic_trigger: 60 machinist: @@ -281,7 +287,8 @@ helm upgrade "$RELEASE" oci://ghcr.io/platformatic/helm \ helm upgrade "$RELEASE" oci://ghcr.io/platformatic/helm \ --version "$CHART_VERSION" \ --namespace "$RELEASE_NAMESPACE" \ - -f v4-values.yaml + -f v4-values.yaml \ + --wait --timeout 10m ``` Review the dry-run for unexpected namespace changes. Compare it with the saved @@ -296,6 +303,24 @@ kubectl rollout status deployment/machinist --namespace platformatic helm test "$RELEASE" --namespace "$RELEASE_NAMESPACE" ``` +Run these queries in Prometheus. Each must return a value greater than zero: + +```promql +count(container_cpu_usage_seconds_total{container!="POD"}) +``` + +```promql +count( + kube_pod_container_resource_limits{resource="cpu", unit="core"} + or + kube_pod_container_resource_requests{resource="cpu", unit="core"} +) +``` + +```promql +count(kube_pod_labels{label_platformatic_dev_monitor="prometheus"}) +``` + Then confirm: - ICC `GET /api/status` returns 200.