IBM Tool Targets Hidden Inconsistency in AI Agent Performance

IBM researchers have developed a consistency analyzer and new guideline type to address the gap between average success rates and reliable task completion in AI agents. On the AppWorld benchmark, a GPT-4.1 ReAct agent succeeded on 77.4% of runs but only completed 53.0% of tasks consistently across five repetitions. The proposed approach aims to reduce this variability without sacrificing overall accuracy.
The Consistency Analyzer works by resampling individual decision points within a single recorded agent trajectory, requesting multiple completions per step rather than re-running entire tasks end-to-end. This design requires no ground-truth labels and operates from just one trace, making it lightweight for practical deployment. The resulting consistency guidelines are a new addition to the ALTK-Evolve framework, which previously distilled reusable guidelines from past trajectories to boost average success. The new guidelines target flip-prone decisions specifically, and the team reports they generalize across tasks rather than patching single trajectories. On AppWorld's hard tasks, the consistency gap reached roughly 30 percentage points, underscoring how much variability standard Mean@k metrics conceal.
This work could reshape how enterprises evaluate AI agents before deployment, particularly in finance, legal review, and other high-stakes domains where a single failed run carries real cost. If consistency metrics become standard practice, organizations may demand more rigorous testing before trusting agents with critical workflows, potentially slowing adoption but improving reliability. It could also pressure model providers to optimize for repeatable performance rather than headline accuracy, shifting competitive benchmarks toward dependability.