English

NewsAmazon

Amazon Researchers Propose Ising Model Approach to Address Correlations in LLM-as-a-Judge

Amazon researchers have introduced a method to improve the reliability of "LLM-as-a-judge" pipelines by accounting for correlations between different judge models. In current systems, when multiple LLMs evaluate the same task, their agreement might appear stronger than it is if the judges share training lineages, prompt templates, or model families. This can lead to a majority vote that reinforces shared biases rather than providing independent evidence.

The researchers proposed using an Ising model—a statistical model that can represent pairwise dependencies—to treat a judge panel as a network. This approach learns both the reliability of individual judges and the relationships between them, allowing the system to discount redundant agreement and distinguish between genuine consensus and shared mistakes.

In evaluations across three tasks—relevance classification, toxicity detection, and summarization assessment—the dependence-aware aggregation method outperformed weighted-majority-vote baselines by 9% to 14% when using 10-judge panels. The method works in an unsupervised setting, meaning it can learn from judge outputs without requiring human reference labels for training.

Sources

  1. When LLM judges agree, should we believe them? (Hacker News Frontpage, 2026-09-14)