Adaptive inference scheduling for AI agents — per-round model, effort and provider routing for coding harnesses: a Command Code mod or a local OpenAI-compatible proxy.
-
Updated
Sep 30, 2026 - TypeScript
Adaptive inference scheduling for AI agents — per-round model, effort and provider routing for coding harnesses: a Command Code mod or a local OpenAI-compatible proxy.
Reinforcement learning for LLM inference scheduling. DQN agent learns to balance throughput, TTFT, latency, and memory pressure vs FIFO/SJF/priority baselines.
WNIDIA v5.1 · 异构算力纳管与可信推理调度平台 | Hackathon project
To associate your repository with the inference-scheduling topic, visit your repo's landing page and select "manage topics."