Microsoft has released ThinkingBox on Hugging Face, a new framework and dataset designed to evaluate the reliability of AI agents. Unlike traditional benchmarks that focus on the sentences an agent generates, ThinkingBox grades agents based on the actual records and side effects they leave behind in a backend database.
Microsoft Releases ThinkingBox: Evaluating AI Agent Consistency through Database States
The benchmark utilizes isolated Model Context Protocol (MCP) tool sessions to run agents through 507 stateful business workflows. To test for consistency, each task is repeated 20 independent times. This approach aims to identify cases where an agent produces a correct-sounding response while failing to update a database correctly or causing unintended side effects.
According to the research, there is a significant gap between perceived capability and actual reliability. In a study involving 121,680 valid trials across 12 LLM models, approximately 65.6% of attempts failed executable checks. Notably, 67.24% of these failures still resulted in the agent terminating "cleanly," invoking a state-changing tool, and reporting no final tool error, despite leaving behind incorrect field values, unintended extra effects, or missing required effects.
The benchmark results highlight the difference between "pass@1" (the single-attempt score) and consistency across 20 runs. While models such as GPT-6 Astra and Claude Opus 5.5 retained a high percentage of their single-attempt scores, others like DeepSeek-V4-Pro and Kimi-K2.6 showed much lower consistency, retaining only about 8% of their initial performance over 20 repeats.
ThinkingBox and the ThinkingBox-Bench dataset are now available on Hugging Face. The environment allows developers to verify if agents can perform repetitive tasks reliably by checking the terminal state of a system rather than relying on the model's summary.
Sources
- The Agent Said It Was Done. The Database Disagreed. (Hugging Face Blog, 2026-10-03)
- ThinkingBox paper