A research team at Fiddler has found that selecting a cheaper AI model for programming tasks can actually lead to higher total costs. In a study comparing Anthropic's Claude Haiku 4.5 and Claude Sonnet 5, Haiku's cost per successful task was found to be 8.4 to 14.1 times higher than Sonnet's, despite Haiku's token prices being half that of Sonnet.
PLUS ULTRAPricing & LimitsClaude Haiku 4.5Claude Sonnet 5
Cheaper AI Models Can Increase Coding Task Costs Due to Verbose and Fabricated Responses
PLUS ULTRA by Amenoyomi
The study focused on Python debugging tasks using an agentic loop, where the model executes commands to inspect files, edit code, and run tests. The research identified that the cost disparity was driven by the length of the responses and specific inefficient behaviors. Haiku's median output was 361 tokens, significantly higher than Sonnet's 16 tokens.
The researchers observed that Haiku often engaged in "fabricating" session responses. In one instance, a single response contained 21,492 tokens describing 60 shell commands, even though the agent loop only executes one command per turn. Haiku predicted the outcomes of these commands without waiting for the actual execution results, generating a massive amount of unexecuted text. These long responses not only increase immediate output costs but also raise input costs in subsequent turns as the entire conversation history is passed back to the model.
While the study notes that these results may vary depending on the specific coding tasks and the testing harness used, it highlights that per-token pricing is an insufficient metric for evaluating the total cost of AI-driven coding agents. The researchers suggest that organizations should measure success based on the cost per completed task in environments that closely mimic their actual workflows.
PLUS ULTRAby Amenoyomi
The cost inefficiency in cheaper models often stems from a structural failure in how they handle the agent loop. In a standard agentic workflow, a model is expected to issue a single command, stop, and wait for the system to return the actual execution result. However, some models engage in "fabrication," where they predict a plausible outcome for the first command and immediately write that invented result into their response, continuing this process for multiple subsequent commands without any real-world verification.
This behavior transforms a single turn into a simulated session. In one observed case, a model generated over 21,000 tokens describing 60 shell commands and their predicted results in one response, even though the system could only execute the first command. Because the model declares the task complete based on a state that never existed, the remaining reasoning and commands are rendered useless, while the tokens used to generate them are still billed as output.
The financial impact of fabrication extends beyond the initial output. Because the entire conversation history is passed back to the model in every subsequent turn, a single massive, fabricated response creates a "snowball effect" on input costs. For example, one study showed that a model's input context grew 13.7 times between the first and second turns due to this verbosity, forcing the user to pay for the same redundant text in every following request.
These findings suggest that token unit prices are a poor proxy for the actual cost of completing a task. Because a single unusually long response can drastically shift the total expense, the only reliable metric for evaluating coding agents is the total cost per verified successful outcome in an environment that mimics actual production workflows.
Sources
- 単価の安いAIモデルを使うとかえってコストが高くつくことがあるという研究結果 (GIGAZINE, 2026-09-22)
- Fiddler blog