LLM post-training · alignment · agents 🤖 — MSc @ BJTU
- 🎓 MSc in Software Engineering @ Beijing Jiaotong University; BEng @ Nantong University.
- 🔬 I work on LLM post-training (SFT / LoRA / DPO), alignment & preference optimization (PPO / GRPO / reward modeling), LLM agents (LangChain / LangGraph), and efficient inference (vLLM).
- 📝 Co-authored a research paper.
- ✍️ I write notes & blog posts at excelius.xyz.
- ByteDance — Multimodal LLM Algorithm Intern · present
- Baidu — LLM Post-training Algorithm Intern
- Tsinghua University, Institute of Vehicle Power & Intelligent Energy — LLM Application & Full-stack Intern
- Languages: Python, C++, TypeScript / JavaScript
- LLM / Post-training: PyTorch, HuggingFace Transformers, LLaMA-Factory, ms-swift, vLLM
- Alignment / RL: SFT, LoRA, DPO, PPO, GRPO, Reward Modeling (GenRM, LLM-as-a-Judge)
- Agents: LangChain, LangGraph (ReAct, Tool Calling, Plan-and-Execute)
- Infra: multi-node multi-GPU training, Git, Docker
- dive-into-transformer-pytorch ⭐6 — A Transformer language model built from scratch in PyTorch; trains an autoregressive classical-Chinese text generator on Dream of the Red Chamber, with multi-GPU data-parallel training, checkpointing, and loss visualization.
- BERT_BiLSTM_CRF ⭐5 — Chinese named-entity recognition on a literary-prose corpus, built on BERT + BiLSTM + CRF (BJTU NLP coursework).
- Self-DeepResearch — An autonomous deep-research agent on LangGraph + Tavily: plan → search → reflect → generate, iterating to a cited report; Vue frontend with a streaming backend.
- marginalia — A deep paper-reading skill for Claude Code / Codex that reconstructs the author's reasoning, stress-tests its weakest assumptions, and publishes structured notes to a Feishu knowledge base.
- PyTorchClassics — Classic deep-learning models re-implemented from scratch in PyTorch (MLP, LeNet, attention …) — an ongoing study collection.
|
|
|
||||
|
|
|
|
|||
|
|
|||||
