Skip to content
View alexzms's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Organizations

@FoundationResearch

Block or report alexzms

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
alexzms/README.md

Hi, I'm Minshen Zhang. You can also call me Alex. 👋

Github Linkedin Gmail

I am an M.S. student in Computer Science at UC San Diego, advised by Prof. Hao Zhang in Hao AI Lab. I worked as a Student Researcher Intern at ByteDance Seed Infra. I hold a B.S. from ShanghaiTech University, where I was advised by Prof. Kewei Tu.

I work on efficient AI systems: LLM serving, GPU/TPU kernel optimization, sparse attention, and agents for performance engineering. I enjoy connecting model architecture, on-chip dataflow, and distributed execution to make models more efficient in practice.

Current work

  • FastAFD — Currently contributing to a new LLM inference engine that targets high-throughput workloads through Attention–FFN disaggregation.
  • AutoPallas & TPU Inference at ByteDance — During my internship, I was code owner of an internal agent system for TPU kernel optimization, with interactive task specification, configurable single-/multi-agent loops, parallel optimization exploration, a strict correctness gate, and reusable optimization trajectories. I also contributed to Seed’s internal vllm tpu-inference framework and optimized Pallas paged-attention and fused kernels, using HLO/LLO analysis and XProf profiling with XLA as a performance baseline.

Research and open source

  • HiLS-Attention — Core contributor; built the official SGLang serving backend and reference inference system for learned sparse attention, supporting 512K-token context.
  • FlashMHF / FlashFFN — First-author paper accepted at NeurIPS 2026, on software–hardware co-design for efficient FFN architectures. I wrote Flash-style kernels in ThunderKittens/CUDA and Triton that keep intermediates in SRAM. The project reduces peak memory by 3–5× and achieves up to 1.08× inference speedup over the SwiGLU baseline while improving model quality.
  • FastVideo — One of the code owners, contributing training infrastructure, custom GPU kernels, inference optimization, and quantization-aware distillation for video generation and world models.

Technical focus: LLM Serving GPU/TPU Kernels Coding Agents Sparse Attention Software–Hardware Co-design

Pinned Loading

  1. FoundationResearch/FlashMHF FoundationResearch/FlashMHF Public

    Flash Multi-Head Feed-Forward Networks

    Python 8

  2. SGLang-HiLS SGLang-HiLS Public

    SGLang with the HSA (Hierarchical Sparse Attention) backend for HiLS-Attention — for speed benchmarking

    Python 6

  3. hao-ai-lab/FastVideo hao-ai-lab/FastVideo Public

    A unified inference and post-training framework for accelerated video generation.

    Python 4.5k 484

  4. huggingface/transformers huggingface/transformers Public

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    Python 167k 34.7k