English

Product LaunchesIBM Research

IBM Research Introduces Consistency Guidelines via ALTK-Evolve to Stabilize AI Agent Performance

IBM Research has introduced "consistency guidelines," a new type of guideline within their ALTK-Evolve system designed to bridge the "consistency gap" in AI agents. While standard benchmarks often report average success rates (Mean@k), they frequently overlook whether an agent can perform the same task reliably across multiple runs (Pass^k).

The research highlights that even capable models like GPT-4.1 can exhibit significant variability. For instance, on the AppWorld benchmark, a ReAct agent achieved a Mean@5 of 77.4% but a Pass^5 of only 53.0%, representing a 24.4-point consistency gap. This inconsistency often stems from "flat" probability distributions, where a model's decision-making at certain steps is nearly a coin flip, making it vulnerable to minor perturbations.

To combat this, the new workflow employs a two-stage pipeline:

  1. Detection: The "Consistency Analyzer" replays decision steps through controlled resampling to identify steps where the model's output is unstable.
  2. Generation: Targeted guidelines are automatically generated to stabilize these specific decision points.

Evaluations on the AppWorld test set showed that applying these guidelines reduced the consistency gap by approximately half, narrowing it from 24.4pp to 12.0pp. Notably, the aggregate Pass^5 success rate rose from 53.0% to 69.0% without sacrificing the overall Mean@5 accuracy. The ALTK-Evolve toolkit, including the new Consistency Analyzer, is available as an open-source project.

Sources

  1. Your Agent Aced the Task. Will It Do It Again? (Hugging Face Blog, 2026-09-15)
  2. altk-evolve