Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
-
Updated
Mar 19, 2026 - Python
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
Open source skill library for AI coding agents to write, optimize, and debug high performance compute kernels across CUDA, Triton, and quantized workloads.
Custom Linux kernels purpose-built for Apple Mac hardware
Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.
Repo containing artifacts for Neurips 2025 tutorial- How to Build Agents to Generate Kernels for Faster LLMs (and Other Models!)
Agent-queryable ROCm kernel optimization knowledge base for AMD Instinct MI300/gfx942 and MI350/MI355X/gfx950, packaged for Codex CLI and Claude Code with merged-PR provenance, real-silicon validation, and a maintainer-controlled pull-request evidence pipeline.
Evidence-driven CUDA, CUTLASS, Triton and GPU workload optimization for ChatGPT · 使用 ChatGPT 驱动 GPU workload 性能优化
Custom AWS Transform agent that migrates PyTorch/Triton kernels to AWS Trainium NKI (@nki.jit) and compiles, numerically verifies, and profiles every candidate on a real Trainium device before opening a PR.
Noeris — autonomous kernel fusion discovery + Triton autotuning for LLM kernels and Gemma layer deeper fusion (A100/H100 wins).
Extended TileLang as a unified DSL to enable high-performance kernel development for Near-Memory Computing, Distributed Memory AI Accelerators, and Networked Accelerators.
METAL-SCI: a scientific compute benchmark for evolutionary LLM kernel search on Apple Silicon Metal
可验证的 CUDA 学习主线:SGEMM、通用 GPU 算子、性能优化与轻量推理组件
Compiler MVP that detects Transformer fusion patterns, generates optimized CUDA kernels with WMMA Tensor Cores, and executes them on real GPU hardware — 10.5 TFLOPs on RTX 2070, correctness validated against PyTorch.
RWKV-7 FP8 quantized inference - 6.4x decode speedup on Blackwell GPUs with <0.3% accuracy loss. Full FP8 E4M3 weight quantization with fused Triton kernels.
Reward-hardened evaluation for LLM-generated GPU kernels, built on KernelBench.
Hand-written CUDA vs Mojo GPU kernels benchmarked on consumer Ampere (RTX 3090, sm_86), with roofline analysis
Automatic Triton kernel generation and optimization for Intel GPU, powered by Claude Code.
Open reproduction of NVIDIA's AVO paper (arXiv:2603.24517): evolutionary search where an autonomous coding agent IS the variation operator — Vary(P)=Agent(P,K,f). Runs on the Claude Code or Codex session you already have.
Add a description, image, and links to the kernel-optimization topic page so that developers can more easily learn about it.
To associate your repository with the kernel-optimization topic, visit your repo's landing page and select "manage topics."