English

Model ReleasesByteDanceByteDance SeedTsinghua UniversityDAPODAPO-Qwen-32B

ByteDance and Tsinghua Release DAPO: An Open-Source RL System for Large-Scale LLM Training

ByteDance Seed and the Institute for AI Industry Research (AIR) at Tsinghua University have released DAPO (Decoupled Clip and Dynamic sAmpling Policy Optimization), an open-source reinforcement learning (RL) system designed for large-scale Large Language Model (LLM) training. The system includes the algorithm, code infrastructure, and a curated dataset.

In evaluations on the AIME 2024 math competition, the DAPO-Qwen-32B model—trained on a Qwen2.5-32B base model—achieved a score of 50. According to the researchers, this performance outperforms the previous state-of-the-art result, DeepSeek-R1-Zero-Qwen-32B, which scored 47 points, while using only 50% of the training steps.

The DAPO algorithm introduces four key technical approaches to improve large-scale RL for long Chain-of-Thought (CoT) reasoning scenarios:

  • Clip-Higher: Decouples the lower and higher clipping ranges to enhance policy entropy and prevent entropy collapse.
  • Dynamic Sampling: Filters out prompts with perfect or zero accuracy to maintain effective gradients and improve sample efficiency.
  • Token-Level Policy Gradient Loss: Applies loss calculation at the token level to ensure longer sequences have sufficient influence on gradient updates.
  • Overlong Reward Shaping: Implements a penalty mechanism for truncated samples to reduce reward noise and stabilize training.

The researchers have fully open-sourced the training recipe, including the DAPO-Math-17k dataset, to improve the reproducibility of large-scale RL training. The implementation is built on the verl framework.

Sources

  1. DAPO: An Open-Source RL System from ByteDance Seed and Tsinghua Air (Hacker News Frontpage, 2026-09-20)
  2. arXiv