A new open-source benchmarking tool named livenerf has been released to address the growing concern regarding "model nerfing"—the observation that frontier AI models may become less capable or undergo quantization after their initial launch.
Livenerf: A New Deterministic Benchmark to Detect Post-Launch Model Degradation
Unlike traditional evaluations that rely on anecdotal "vibes," livenerf employs a deterministic methodology to track performance drift over time. The tool uses frozen prompts, pinned CLI versions, and exact graders to ensure that any observed changes in performance are due to model updates rather than variations in the testing environment. It is built upon Inspect, the open-source evaluation framework from the UK AI Security Institute.
The benchmark is currently tracking Claude Opus 5.5, specifically as it is served through Claude Code on a subscription. The project follows a pre-registered protocol to ensure transparency, aiming to detect statistically significant changes in core tasks—such as GPQA, MMLU-Pro, and math problems—over a 30-day period. Notably, the tool also monitors output token counts as a secondary signal, as reductions in token counts can often indicate a model is "thinking" less before accuracy metrics even begin to shift.
Sources
- Livenerf: Has Opus 5.5 been nerfed yet? (Hacker News Frontpage, 2026-09-29)