English

NewsMichael Noukhovitch

New "Never Give Up" Method Aims to Solve Hard RL Problems in LLMs

Researchers Michael Noukhovitch et al. have introduced "Never Give Up," a method designed to address the Matthew Effect in reinforcement learning (RL) for Large Language Models (LLMs). The Matthew Effect refers to a phenomenon where RL performance improvements are disproportionately concentrated on easy tasks, while the model fails to make significant progress on harder problems.

The "Never Give Up" approach optimizes compute allocation by using an adaptive sampling strategy. Instead of using a fixed number of samples (k), the method quickly filters tasks that are solved easily and reallocates computing resources to harder problems by sampling additional completions when an initial set fails to provide a solution.

In evaluations using GSM8k and code RL benchmarks like Manufactoria, the method showed improved performance on difficult subsets compared to standard Group Relative Policy Optimization (GRPO). The researchers suggest that this efficiency helps mitigate the stagnation often seen in traditional RL training for complex reasoning and coding tasks.

Sources

  1. Learning to solve hard problems in RL for LLMs by never giving up (Hacker News Frontpage, 2026-09-15)