Technical Lead - DevOps Engineering at Wix, Dublin. I keep a multi-region Kubernetes platform for untrusted user code running: isolation, cold starts, scaling, cost and on-call, one site per gVisor sandbox.
When the bug is in open source, I fix it upstream, and I write up how I found it at catchkill9.dev.
%%{init: {"themeVariables": {"fontSize": "18px"}}}%%
flowchart LR
kubelet -->|CRI| ctr["containerd"]
ctr --> shim["gVisor shim"]
shim --> sandbox
subgraph sandbox["Pod sandbox, one site"]
code["Untrusted<br/>user code"] -->|syscalls| sentry["gVisor Sentry"]
end
sentry --> kernel["Host kernel"]
style sandbox fill:#fff8e1,stroke:#bf8700,color:#4d2d00
style code fill:#fff1c2,stroke:#bf8700,color:#4d2d00
style kubelet fill:#326ce5,stroke:#1f4fb4,color:#ffffff
style ctr fill:#8250df,stroke:#5e34a8,color:#ffffff
style shim fill:#8250df,stroke:#5e34a8,color:#ffffff
style sentry fill:#8250df,stroke:#5e34a8,color:#ffffff
style kernel fill:#1a7f37,stroke:#0f5323,color:#ffffff
| containerd | |
|---|---|
| Bound snapshot GC so a hung snapshotter can't block new pods while the node still reports Ready | merged, then reverted for a client-side fix |
| Configurable deadline on proxy snapshotter calls so one hung call can't block the whole snapshotter | open |
| SOCI snapshotter | |
|---|---|
| Close the span cache reader in file Verify so pod-start bursts can't leak the snapshotter past its fd limit | merged |
| Security Profiles Operator | |
|---|---|
| Handle AppArmorProfile owners in node status, my first upstream PR | merged |
All of it with the backstory: catchkill9.dev/oss



