English

NewsUniversity of California, BerkeleyArena

Research Finds AI Coding Agent "Harnesses" Significantly Impact Cost, While Contribution to Success Rate Remains Limited

This article is a translation. Read the Japanese original

AI coding agents manage processes such as which tools to allow the model to use and what information to provide as context through tools called "harnesses" when running AI models.

A joint research team from the University of California, Berkeley, and the AI ranking site "Arena" conducted a comparative study across 2 types of benchmarks, covering 21 patterns that combine 7 models with 3 harnesses (Claude Code, Codex CLI, and Pi).

The results of the research showed that the difference in task success rates due to variations in the harness remained within ±2% for SWE-bench Lite and generally within ±5% for Terminal-Bench 2.0. On the other hand, the study demonstrated that costs could differ by as much as five times. For example, when comparing Claude Fable 5, Claude Code cost approximately twice as much as Pi.

One reason for this cost discrepancy is the difference in the amount of initial context provided to the model by the harness. It was confirmed that the initial context for Claude Code was more than ten times that of Pi in some cases. The research team refers to the state of paying extra costs to achieve results of the same quality as a "Harness Tax."

Furthermore, the study found that harnesses provided by model providers do not necessarily deliver optimal performance. The research team stated that the key is not the model provider, but rather selecting the harness that best maintains the balance between cost and success rate in actual work.

Sources

  1. How important are harnesses for AI coding agents? (GIGAZINE, 2026-09-17)