I work across AI agents, cloud-native AI and inference acceleration: Agent Loop / RSI, disaggregated P/D deployments, and tiered KV cache reuse.
The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.
- Bison — Enterprise GPU billing & multi-tenant platform
- Inference Cookbook — Inference frameworks, deep-dived
- Cloud Native Cookbook — Cloud-native engineering, deep-dived




