I am an M.S. student in Computer Science at UC San Diego, advised by Prof. Hao Zhang in Hao AI Lab. I worked as a Student Researcher Intern at ByteDance Seed Infra. I hold a B.S. from ShanghaiTech University, where I was advised by Prof. Kewei Tu.
I work on efficient AI systems: LLM serving, GPU/TPU kernel optimization, sparse attention, and agents for performance engineering. I enjoy connecting model architecture, on-chip dataflow, and distributed execution to make models more efficient in practice.
- FastAFD — Currently contributing to a new LLM inference engine that targets high-throughput workloads through Attention–FFN disaggregation.
- AutoPallas & TPU Inference at ByteDance — During my internship, I was code owner of an internal agent system for TPU kernel optimization, with interactive task specification, configurable single-/multi-agent loops, parallel optimization exploration, a strict correctness gate, and reusable optimization trajectories. I also contributed to Seed’s internal vllm tpu-inference framework and optimized Pallas paged-attention and fused kernels, using HLO/LLO analysis and XProf profiling with XLA as a performance baseline.
- HiLS-Attention — Core contributor; built the official SGLang serving backend and reference inference system for learned sparse attention, supporting 512K-token context.
- FlashMHF / FlashFFN — First-author paper accepted at NeurIPS 2026, on software–hardware co-design for efficient FFN architectures. I wrote Flash-style kernels in ThunderKittens/CUDA and Triton that keep intermediates in SRAM. The project reduces peak memory by 3–5× and achieves up to 1.08× inference speedup over the SwiGLU baseline while improving model quality.
- FastVideo — One of the code owners, contributing training infrastructure, custom GPU kernels, inference optimization, and quantization-aware distillation for video generation and world models.
Technical focus: LLM Serving GPU/TPU Kernels Coding Agents Sparse Attention Software–Hardware Co-design


