Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 17 additions & 4 deletions .agents/skills/helm-dev-environment/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,10 @@ namespace with a `uri` key. For local manual testing, either create your own
PostgreSQL Secret or use the e2e PostgreSQL fixture manifest in
`e2e/kubernetes/postgres-fixture.yaml`.

Gateway pods in the `high-availability` profile use external PostgreSQL, so
they drain their supervisor sessions when deleted or rolled and can take up to
30 seconds to terminate.

For the `high-availability` profile, return to the repository root and apply the
GatewayClass and BackendTrafficPolicy manifest after Skaffold has installed
Envoy Gateway:
Expand Down Expand Up @@ -271,16 +275,24 @@ The kube e2e wrapper creates only one port-forward, to `svc/openshell`; it no
longer forwards the unauthenticated health listener or runs a `/readyz` e2e
target. `/readyz` remains covered by server unit/integration tests.

Use `mise run e2e:kubernetes:ha-rebalancing` for full-suite HA coverage. The
task creates an external PostgreSQL fixture, installs Envoy Gateway, applies
Use `mise run e2e:kubernetes:ha-rebalancing` for HA coverage. The task creates
an external PostgreSQL fixture, installs Envoy Gateway, applies
`deploy/kube/manifests/envoy-gateway-openshell.yaml`, enables the chart
`GRPCRoute`, and runs the full Kubernetes e2e suite, including
`kubernetes_ha_rebalancing`. That coverage validates sandbox create/watch and
`GRPCRoute`, and runs the CLI conformance profile and the
`kubernetes_ha_rebalancing` tests. That coverage validates sandbox create/watch and
exec through the Envoy proxy while gateway replicas scale up, scale down, and
rotate. It also keeps a long-running sandbox alive and runs upload/download
operations while gateway pods roll, so file sync exercises the same relay retry
path as interactive sessions.

`supervisor_sessions_redistribute_across_gateway_pod_rolls` restarts the
gateway Deployment, reads `openshell_server_draining` from the terminating
pods, and checks that the new pods' `openshell_server_supervisor_sessions` add
up to the Ready sandbox count. It scrapes each pod through
`kubectl get --raw /api/v1/namespaces/<ns>/pods/<pod>:9090/proxy/metrics`.
Gateway pods take up to 30 seconds to terminate because each drains its

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[general obs] A lone PostgreSQL replica still drains even though no peer can receive its sessions, adding up to 15 seconds.

sessions, and the HA tests are allowed ten minutes each.

If you reuse an existing Skaffold cluster for the full kube suite, make sure the
chart has `server.hostGatewayIP` set so sandbox pods can resolve
`host.openshell.internal` back to the test host. The e2e wrapper detects this on
Expand Down Expand Up @@ -446,6 +458,7 @@ for dependencies still declared in `Chart.yaml`.
| `deploy/helm/openshell/ci/values-cert-manager.yaml` | cert-manager PKI overlay (opt-in; disables pkiInitJob) |
| `deploy/helm/openshell/ci/values-gateway.yaml` | Envoy Gateway GRPCRoute + Gateway overlay |
| `deploy/helm/openshell/ci/values-high-availability.yaml` | HA test overlay (`replicaCount: 2` with external PostgreSQL Secret) |
| `deploy/helm/openshell/ci/values-autoscaling.yaml` | Render-only overlay for the optional gateway HorizontalPodAutoscaler (helm lint and helm-unittest) |
| `deploy/helm/openshell/ci/values-keycloak.yaml` | Keycloak OIDC overlay |
| `deploy/helm/openshell/ci/values-spire.yaml` | SPIFFE/SPIRE provider token grant overlay |
| `deploy/helm/openshell/ci/values-spire-stack.yaml` | SPIRE hardened chart values for local dev |
Expand Down
4 changes: 4 additions & 0 deletions .config/nextest.toml
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,10 @@ kubernetes-ha = { max-threads = 1 }
[[profile.e2e-kubernetes.overrides]]
filter = "test(/gateway_(scale_and_rollout|pod_rolls)$/)"
test-group = "kubernetes-ha"
# Each rolled gateway pod drains its supervisor sessions for up to the 30s
# termination grace period, so these tests get ten minutes, matching
# HA_SYNC_TIMEOUT in kubernetes_ha_rebalancing.rs.
slow-timeout = { period = "60s", terminate-after = 10 }

# Relative to the profile store dir (`e2e/rust/target/nextest/e2e-kubernetes/`).
[profile.e2e-kubernetes.junit]
Expand Down
18 changes: 18 additions & 0 deletions TESTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,23 @@ mise run test:rust # cargo test --workspace

Rust validation checks tracked Cargo lockfiles; run `mise run rust:lockfiles:check` to check them directly. If one is stale, refresh it with Cargo using its adjacent manifest, review the diff, and commit the update.

### PostgreSQL-backed tests

Tests that need a real PostgreSQL server, such as advisory-lock concurrency
across two stores, are `#[ignore]`d and named `postgres_*`. Run them with:

```shell
mise run test:rust:postgres
```

The task starts a disposable PostgreSQL container with Docker or Podman
(set `CONTAINER_ENGINE` to choose), runs the tests one at a time, and removes
the container. Each test works in its own temporary schema. To use your own
disposable database, set `OPENSHELL_TEST_POSTGRES_URL`. Never point it at a
database that a running gateway uses: the tests take fleet-wide advisory locks.
CI does not run these tests; the Kubernetes HA e2e suite covers PostgreSQL end
to end.

### Native Windows validation

Use `mise run --skip-tools pre-commit` with the existing Rust/MSVC toolchain.
Expand Down Expand Up @@ -428,6 +445,7 @@ Available task variants:
|---|---|
| `e2e:kubernetes` | Default Rust e2e against Helm-deployed gateway |
| `e2e:kubernetes:db` | All database backend scenarios (SQLite + external PostgreSQL) |
| `e2e:kubernetes:ha-rebalancing` | Two gateway replicas behind Envoy with external PostgreSQL: scale, pod deletion, rollout drain and session redistribution, and file sync during pod rolls |
| `e2e:kubernetes:sidecar` | Supervisor sidecar topology overlay |
| `e2e:kubernetes:credential-drivers` | Kubernetes Secrets and Vault credential storage |
| `e2e:kubernetes:workspace-managed` | Managed workspace mode (auto-created namespaces) |
Expand Down
Loading
Loading