The implementation for our paper: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning.
-
Updated
Aug 5, 2026 - Python
The implementation for our paper: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning.
Based on the virtual world built with IsaacSim, accomplish the training and deployment of real-world reinforcement learning for Hil-Serl.
基于ms-swift的Qwen3_8B 金融推理两阶段后训练(LoRA SFT->GRPO)。
Interactive Autonomous Navigation Research Lab for UGVs featuring localization, path planning, DWA, MPC, Adaptive MPC, reinforcement learning hooks, safety supervision, live visualization, benchmarking, replay, and automated report generation.
To associate your repository with the reforcement-learning topic, visit your repo's landing page and select "manage topics."