English

News

Open-source Harnesses Provide No Advantage for Coding Agents with Strong LLM Backbones, Apple Research Finds

Apple Machine Learning researchers have found that open-source state-of-the-art harnesses provide no advantage over a single session of a minimal-harness coding agent baseline, under equal time budgets and the same frontier LLM backbone.

In the paper "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?", authors Kirill Brilliantov, Alejandro Hernández-Cano, and Emmanuel Abbé state that while modern machine learning engineering (MLE) agents often deploy on elaborate machinery—such as multi-agent orchestrators and dedicated retrieval subagents—these layers may become redundant in certain coding agent settings. The researchers argue that performance is primarily driven by the underlying LLM backbone.

The study focused on coding agents that have direct access to execution environments through read, write, and bash primitives. The findings suggest that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

Sources

  1. How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? (Apple Machine Learning, 2026-10-01)
  2. arXiv