Hydra-2P, KDA, and CADR: fused PyTorch/Triton kernels for routing across model depth and sequence time.
-
Updated
Jul 18, 2026 - Python
Hydra-2P, KDA, and CADR: fused PyTorch/Triton kernels for routing across model depth and sequence time.
Runnable C++20 experiments for RoPE, KV cache, speculative decoding, attention layouts, online softmax, and GEMM.
Implementing FlashAttention from scratch to understand standard attention, online softmax, tiling, memory efficiency, and GPU-aware attention algorithms.
Row-wise CUDA softmax kernels: shared-memory reduction, warp-shuffle reduction, and online softmax benchmarked against cuDNN.
Benchmark-driven Triton kernel experiments covering vector operations, reductions, online softmax, and GPU optimization.
To associate your repository with the online-softmax topic, visit your repo's landing page and select "manage topics."