This repo focuses on latency-aware resource optimization for Kubernetes
-
Updated
Oct 8, 2025 - Python
This repo focuses on latency-aware resource optimization for Kubernetes
Adaptive C++/CUDA runtime that profiles workloads at submission time and dynamically routes to CPU, GPU, or batched execution based on arithmetic intensity and transfer cost
Add a description, image, and links to the workload-profiling topic page so that developers can more easily learn about it.
To associate your repository with the workload-profiling topic, visit your repo's landing page and select "manage topics."