Roofline profiling for Deep Learning models
-
Updated
May 4, 2024 - Python
Roofline profiling for Deep Learning models
Kernel-level profiling of batch-1 decode on consumer GPU to prove that decode is memory-bandwidth-bound against the hardware roofline, then beating the baseline with a fused dequant+GEMV CUDA kernel. (regime-aware attribution at the end)
To associate your repository with the roofline-profiler topic, visit your repo's landing page and select "manage topics."