Popular repositories Loading
-
Annotated-LLM-Runtime
Annotated-LLM-Runtime PublicFrom-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path comment…
-
WarpGroup-backend
WarpGroup-backend PublicA high-performance C++ backend for extreme-context LLM inference. It replaces item-count batching with dynamic, VRAM-aware First-Fit Decreasing (FFD) bin packing. By using PyBind11 for async queuei…
-
sionna-munich-ai-ran
sionna-munich-ai-ran PublicGPU-accelerated Network Digital Twin for AI-RAN experiments in Munich. Powered by NVIDIA Sionna RT and PyTorch for ray-traced RF simulation. Includes real-world Munich simulation, Mitsuba 3 backend…
-
ILCP-for-Agents
ILCP-for-Agents PublicInductive Latent Context Persistence (ILCP) for Agentic AI. This infrastructure persists, routes, and reuses LLM latent context across multi-agent DAGs. By eliminating redundant prefix-prefill comp…
Python 2
-
VRAM-Conductor
VRAM-Conductor PublicC++17 + CUDA orchestrator for running multiple LLM agents on one constrained GPU. `lmxd` daemon does NVML-seeded admission control to stop llama.cpp OOM crashes; `LayerStreamer` + `PinnedHostPool` …
C++ 2
If the problem persists, check the GitHub status page or contact support.
