ByteDance Seed and the Institute for AI Industry Research (AIR) at Tsinghua University have released DAPO (Decoupled Clip and Dynamic sAmpling Policy Optimization), an open-source reinforcement learning (RL) system designed for large-scale Large Language Model (LLM) training. The system includes the algorithm, code infrastructure, and a curated dataset.
Model ReleasesByteDanceByteDance SeedTsinghua UniversityDAPODAPO-Qwen-32B
ByteDance and Tsinghua Release DAPO: An Open-Source RL System for Large-Scale LLM Training
In evaluations on the AIME 2024 math competition, the DAPO-Qwen-32B model—trained on a Qwen2.5-32B base model—achieved a score of 50. According to the researchers, this performance outperforms the previous state-of-the-art result, DeepSeek-R1-Zero-Qwen-32B, which scored 47 points, while using only 50% of the training steps.
The DAPO algorithm introduces four key technical approaches to improve large-scale RL for long Chain-of-Thought (CoT) reasoning scenarios:
- Clip-Higher: Decouples the lower and higher clipping ranges to enhance policy entropy and prevent entropy collapse.
- Dynamic Sampling: Filters out prompts with perfect or zero accuracy to maintain effective gradients and improve sample efficiency.
- Token-Level Policy Gradient Loss: Applies loss calculation at the token level to ensure longer sequences have sufficient influence on gradient updates.
- Overlong Reward Shaping: Implements a penalty mechanism for truncated samples to reduce reward noise and stabilize training.
The researchers have fully open-sourced the training recipe, including the DAPO-Math-17k dataset, to improve the reproducibility of large-scale RL training. The implementation is built on the verl framework.
Sources
- DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air (Hacker News Frontpage, 2026-09-20)
- arXiv