Skip to content

Pinned Loading

  1. understand-r1-zero understand-r1-zero Public

    Understanding R1-Zero-Like Training: A Critical Perspective

    Python 1.3k 63

  2. zero-bubble-pipeline-parallelism zero-bubble-pipeline-parallelism Public

    Forked from NVIDIA/Megatron-LM

    Zero Bubble Pipeline Parallelism

    Python 466 31

  3. lorahub lorahub Public

    [COLM 2024] LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

    Python 669 44

  4. oat oat Public

    🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.

    Python 669 63

  5. stde stde Public

    Official implementation of Stochastic Taylor Derivative Estimator (STDE) NeurIPS2024

    Python 127 9

  6. feedback-conditional-policy feedback-conditional-policy Public

    Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"

    Python 65 2

Repositories

Showing 10 of 101 repositories
  • envpool Public

    C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.

    sail-sg/envpool's past year of commit activity
    C++ 1,515 Apache-2.0 141 16 0 Updated Aug 31, 2026
  • NDA Public

    Code for "Nonparametric Data Attribution for Diffusion Models"

    sail-sg/NDA's past year of commit activity
    Jupyter Notebook 27 1 2 0 Updated May 27, 2026
  • jrystal Public

    A JAX-based Differentiable Density Functional Theory Framework for Materials

    sail-sg/jrystal's past year of commit activity
    Python 52 Apache-2.0 1 5 3 Updated Apr 20, 2026
  • odc Public

    On demand communication

    sail-sg/odc's past year of commit activity
    Python 33 4 2 4 Updated Apr 16, 2026
  • TeamHOI Public

    [CVPR 2026] TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size

    sail-sg/TeamHOI's past year of commit activity
    Python 47 MIT 4 0 0 Updated Mar 12, 2026
  • Stable-RL Public

    Rethinking the Trust Region in LLM Reinforcement Learning

    sail-sg/Stable-RL's past year of commit activity
    Python 73 Apache-2.0 7 1 5 Updated Mar 2, 2026
  • oat Public

    🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.

    sail-sg/oat's past year of commit activity
    Python 669 Apache-2.0 63 6 1 Updated Jan 29, 2026
  • sail-sg/LifelongSafetyAlignment's past year of commit activity
    Python 12 0 0 0 Updated Jan 13, 2026
  • feedback-conditional-policy Public

    Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"

    sail-sg/feedback-conditional-policy's past year of commit activity
    Python 65 2 0 0 Updated Jan 5, 2026
  • InfNeRF Public

    InfNeRF: Towards Infinite Scale NeRF Rendering with O(log n) Space Complexity

    sail-sg/InfNeRF's past year of commit activity
    Python 12 Apache-2.0 1 1 0 Updated Jan 3, 2026