feat(workers): added health probe for workers and handling sigterms and sigint signals - #72
Merged
Merged
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
10 of 12 tasks
tsdk02
added a commit
that referenced
this pull request
Sep 4, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow superposition's chart: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. - A pre-install/pre-upgrade migration Job, disabled by default until the api image supports a migrate-and-exit invocation. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. Requires #72 for the worker's /health and /ready endpoints. Closes #73
tsdk02
added a commit
that referenced
this pull request
Sep 4, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. - A pre-install/pre-upgrade migration Job, disabled by default until the api image supports a migrate-and-exit invocation. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. Requires #72 for the worker's /health and /ready endpoints. Closes #73
tsdk02
approved these changes
Sep 5, 2026
tsdk02
added a commit
that referenced
this pull request
Sep 5, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tsdk02
added a commit
that referenced
this pull request
Sep 7, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. The release workflow now keeps the chart in step with the app. After cog bumps the version it rewrites Chart.yaml's version and appVersion from the new tag, amends the bump commit and moves the tag onto it -- cog tags the pre-amend commit, and both the docker and helm-chart jobs check out by tag. image.tag therefore defaults to appVersion and only needs setting to pin a different build. A new job packages the chart and pushes it to oci://ghcr.io/juspay/helm-charts, the same registry as the images and the way hyperswitch-helm already consumes sibling charts. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tsdk02
added a commit
that referenced
this pull request
Sep 7, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. The release workflow now keeps the chart in step with the app. After cog bumps the version it rewrites Chart.yaml's version and appVersion from the new tag, amends the bump commit and moves the tag onto it -- cog tags the pre-amend commit, and both the docker and helm-chart jobs check out by tag. image.tag therefore defaults to appVersion and only needs setting to pin a different build. A new job packages the chart and pushes it to oci://ghcr.io/juspay/helm-charts, the same registry as the images and the way hyperswitch-helm already consumes sibling charts. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tsdk02
added a commit
that referenced
this pull request
Sep 25, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. The release workflow now keeps the chart in step with the app. After cog bumps the version it rewrites Chart.yaml's version and appVersion from the new tag, amends the bump commit and moves the tag onto it -- cog tags the pre-amend commit, and both the docker and helm-chart jobs check out by tag. image.tag therefore defaults to appVersion and only needs setting to pin a different build. A new job packages the chart and pushes it to oci://ghcr.io/juspay/helm-charts, the same registry as the images and the way hyperswitch-helm already consumes sibling charts. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd sigint signals
tsdk02
force-pushed
the
worker-health
branch
from
September 25, 2026 17:30
9524669 to
5c868f9
Compare
tsdk02
added a commit
that referenced
this pull request
Sep 28, 2026
Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. The release workflow now keeps the chart in step with the app. After cog bumps the version it rewrites Chart.yaml's version and appVersion from the new tag, amends the bump commit and moves the tag onto it -- cog tags the pre-amend commit, and both the docker and helm-chart jobs check out by tag. image.tag therefore defaults to appVersion and only needs setting to pin a different build. A new job packages the chart and pushes it to oci://ghcr.io/juspay/helm-charts, the same registry as the images and the way hyperswitch-helm already consumes sibling charts. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tsdk02
added a commit
that referenced
this pull request
Sep 28, 2026
* feat(helm): add Helm chart for Kubernetes deployment Invokr runs as two workloads sharing one PostgreSQL database, so the chart ships an api Deployment (HTTP :8080, also serving the dashboard) and a worker Deployment (ops server :9090, no inbound traffic), each with its own ConfigMap, probes, HPA and PodDisruptionBudget. Structure and conventions follow the stack's other service charts: a flat values map rendered into env vars via envFrom, a global.imageRegistry override so an ECR mirror is one value rather than a chart fork, optional Istio VirtualService and DestinationRule, and a helm-docs generated README. Adapted where Invokr differs: - Three ConfigMaps rather than one, so worker settings never reach an API pod. - configsToData takes the values map as an argument instead of hardcoding a path, which is what lets one helper render all three. - Only nil is skipped when rendering config, so `false` and `0` survive. A truthiness guard would silently drop kms_enabled: false. - Component labels on every resource, so the two Deployments and their Services cannot select each other's pods. - Per-workload config checksums, so a worker-only values change does not roll the API pods. - Two ServiceMonitors: the API serves /metrics under the configured path prefix, the worker serves it unprefixed on its ops port. Schema migrations run as a pre-install/pre-upgrade hook Job. Helm blocks on hook Jobs, so a failed migration aborts the release before any pod is replaced -- the failure mode is "nothing changed", not "half changed". The Job runs the api image at the chart's appVersion, so it cannot apply a different version's migrations than the code about to start. Istio injection is disabled on it, or Envoy outlives the container and Helm hangs on the hook. INVOKR_DB_MIGRATION_MODE is set on the Job rather than in the shared ConfigMap: the application defaults to `none`, and only this Job should ever migrate. Values are validated at render time: the chart fails with the specific missing value rather than deploying something that cannot start, and rejects a PodDisruptionBudget whose minAvailable would make nodes undrainable. INVOKR_LISTEN_ADDR is derived from api.service.targetPort, and the path prefix feeds the probes, ingress paths, ServiceMonitor path and dashboard config from one value, so neither can drift. The release workflow now keeps the chart in step with the app. After cog bumps the version it rewrites Chart.yaml's version and appVersion from the new tag, amends the bump commit and moves the tag onto it -- cog tags the pre-amend commit, and both the docker and helm-chart jobs check out by tag. image.tag therefore defaults to appVersion and only needs setting to pin a different build. A new job packages the chart and pushes it to oci://ghcr.io/juspay/helm-charts, the same registry as the images and the way hyperswitch-helm already consumes sibling charts. Requires #72 for the worker's /health and /ready endpoints. Closes #73 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: helm hooks, ingress path prefix, migration job * feat(helm): add extraEnv and extraEnvFrom for non-INVOKR_ variables `configs` and `secrets` are always rendered as INVOKR_<KEY>, which is a deliberate guarantee rather than a convenience. It also meant a variable the process needs under any other name could not be set at all. The blocking case is `kms.enabled`. The AWS SDK reads AWS_REGION and falls back to us-east-1 when it is unset, so a key in any other region fails to decrypt and every pod crash-loops on startup. AWS_ENDPOINT_URL was equally unreachable, which made the LocalStack path in crates/common/src/kms.rs untestable from Kubernetes despite being written for exactly that. Applied to the api, the worker and the migration Job -- the Job runs the same image with the same INVOKR_KMS_ENABLED, so it must decrypt too. Verified end to end against LocalStack in-cluster: both workloads log "Using custom AWS endpoint", the Secret holds ciphertext, and removing AWS_ENDPOINT_URL correctly fails the pre-upgrade hook without replacing any running pod. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(helm): correct NOTES.txt output and guard CPU-less autoscaling Three defects, each invisible to `helm lint` and `--dry-run=server` and each found by running what the chart actually prints. NOTES.txt section 2 printed nothing at all in the documented default configuration. ingress.yaml derives a host's paths from `configs.path_prefix` when `paths` is omitted, but NOTES.txt ranged over `.paths` directly, so an empty range emitted no URL -- and because the outer `if ingress.enabled` was true, the port-forward fallback was skipped as well. It now mirrors ingress.yaml's own if/else. NOTES.txt section 3 documented `{"name": "your-org"}`, which the API rejects with `missing field slug`. CreateOrganization and CreateWorkspace both require name and slug. The corrected payloads were run verbatim against a live deployment. An HPA computes utilization as a percentage of the CPU request, and `resources` defaults to {}. Enabling autoscaling therefore produced an HPA reporting <unknown> utilization that could never scale, with nothing to say so -- the same silent-misconfiguration shape the chart's other guards exist to prevent. `dig` is used because a direct field access panics on the empty default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(just): make kms-dev actually load .env.kms The recipe checked that .env.kms existed and then never used it -- the `cp .env.kms .env` line was commented out. Both binaries read config via dotenvy::dotenv(), which loads .env, so `just kms-dev` started services with the kms feature compiled in but INVOKR_KMS_ENABLED unset: plaintext mode, never touching KMS at all. Export the file instead of copying it over .env. dotenvy does not override variables already present in the environment, so these win without clobbering the developer's own .env. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(helm): target main, with the migration hook off until the image supports it The chart was stacked on the embedded-migrations branch. Rebased onto main and decoupled from it. The release workflow was the blocker: the helm-chart job declared `needs: [tag-release, create-manifest]`, and create-manifest arrived with the multi-arch work on that branch. Against main it is an unknown job, which invalidates the whole workflow — every release would fail, not just the chart step. It now depends on docker, which is the real prerequisite: the chart must not publish pointing at a version whose images do not exist. migration.enabled now defaults to false. main's api binary ignores argv, so `args: ["migrate"]` starts an API server that never exits; the pre-install hook would hang until Helm's timeout and fail the release with a message that says nothing about migrations. Failing loudly would be tolerable, hanging is not. The Job and its hook-scoped prerequisites are kept, disabled. They are correct and reviewable now, and flipping the default is one line once the image gains a migrate-and-exit path — cheaper than deleting and re-adding them, and it keeps the whole deployment story in one PR. Three docs claimed behaviour main does not have: NOTES.txt told operators to run `INVOKR_DB_MIGRATION_MODE=run <api-image> migrate`, and the chart README said migrations are compiled into the image with a dry-run mode. NOTES.txt now gives the psql loop that actually works, and the README says plainly that migrations are a manual step today. database.host also drops out of the required-values list, since the guard that demands it only fires when the hook is enabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Shubhranshu Sanjeev <shubhranshu.sanjeev@juspay.in>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.