Hands-on Qwen3-1.7B code post-training lab: SFT, LoRA, DPO, PPO, GRPO, and reproducible EvalPlus benchmarks.
-
Updated
Sep 27, 2026 - Python
Hands-on Qwen3-1.7B code post-training lab: SFT, LoRA, DPO, PPO, GRPO, and reproducible EvalPlus benchmarks.
Compute-efficient QLoRA specialization of Qwen3-4B for Python, with executable EvalPlus evidence and reproducible local training.
Cluster-Aware Sequential Refinement (CASR) for test-free multi-model code generation and model-pool selection.
Executable-code evaluation and verifier-guided inference, with reproducible benchmarks and separate model-weight retention audits.
Research code and artifacts for How LLMs Fail in Code Generation: a cross-benchmark, fine-grained error taxonomy (SDDT 2026).
To associate your repository with the evalplus topic, visit your repo's landing page and select "manage topics."