AI agents are fast evolving, but evaluation is lagging.
Stanford HAI’s 2026 AI Index shows OSWorld task success rising from 12% to about 66%, yet agents still fail roughly one in three attempts on structured benchmarks, while responsible AI benchmarks continue to lag behind model capabilities.
AI teams need evaluations that can distinguish meaningful differences between increasingly capable outputs.
When the evaluation is poorly built, or when benchmarks report their results inconsistently, it's hard to know what actually drove the score.
Across 6,523 AI evaluations involving 215 annotators, Hugo Inc. found that task design drove far more variation in model preference than demographic differences among raters.
"We've audited pools that were rebalanced three times and still produced muddy signal," says Sean Neighbors, Strategic Accounts Executive at Hugo.
"Our data shows if the task isn't designed to force a real distinction, no rater roster will fix it."
Why AI Evaluations Fail to Distinguish Good Outputs
Hugo’s audit points to task design, not rater demographics, as the bigger driver of evaluation quality:
- Demographic groups differed by only 2–3 percentage points in model preference: male annotators selected Model A 44% of the time, compared with 42% among female annotators, and other demographic splits showed similarly small differences.
- Task design produced 70 times more variance than rater demographics.
- Tie rates varied sharply by task type, from 86% on instruction-following to 16% on overall-preference tasks.
While a diverse rater pool can expose cultural references, regional language, and safety issues a narrower one would miss, it can’t fix inconsistent preference data that poor task design has made hard to read.
“High tie rates do not always mean the outputs are equal,” says Neighbors "A poorly designed task can make meaningful differences hard to see."
Hugo recommends auditing task design before assuming persistent ties stem from annotator behavior.
The issue shows up beyond Hugo's own data.
In Stanford's 2025 AI Index, BetterBench examined 24 prominent AI benchmarks and found that 14 did not report statistical significance and 17 lacked scripts for result replication.
The study also found documentation and construction problems that limited reproducibility and the benchmarks' usefulness for judging model performance.
The findings put the focus on the evaluation itself, because raters can only make useful judgments when the task gives them enough information to distinguish between outputs.
How to Improve Human AI Evaluation Feedback
Task design is only half the picture. Rater decisiveness is the other.
The way a task is built determines whether raters can tell two outputs apart. A rater’s decisiveness determines how often they are willing to choose one.
Hugo found that master’s degree holders had a 16.6% tie rate compared with 19.8% for bachelor’s degree holders, while male annotators were 3-4 percentage points more decisive.
Their model preferences remained broadly aligned.
The difference was in decisiveness, not preference. Some groups were simply more likely to choose one output instead of returning a tie.
“When a task asks raters to choose between two strong outputs, you need people who can make a clear call when the criteria supports it.” Neighbors says.
“But that is a selection decision that only matters after the task itself is well designed.”
An effective evaluation process needs a way for raters to express meaningful differences between outputs.
And this comes down to two layers worked through in order.
Layer 1: Evaluation design
Start with the task.
Criteria should be clear enough for raters to apply consistently, and instructions should make the priorities clear.
The rubric must also let raters register their preference or separate genuine indifference from uncertainty about the task.
Binary choices can push them toward ties that become valid when two outputs truly cannot be separated.
Layer 2: Rater design
Once the rubric works, teams can assess how decisively the rater group responds.
A group that readily returns ties will produce high agreement but thin preference data.
That works when the goal is to confirm two outputs are equivalent, but not when the task needs a clear winner.
A more decisive group can produce sharper signals, if the rubric gives them a valid basis for choosing.
Stanford's 2026 AI Index makes the stakes clear: AI agents are improving fast, but the ability to reliably judge them isn't keeping up.
How to evaluate model outputs, what good feedback looks like, and where expert judgment fits in the process are all open questions that are driving a wider conversation across the field.
Hugo hosted a series of offline conversations in New York City.
Join the catch-up event and learn more about reasoning, multimodal evaluation, reward hacking and the people behind better model feedback.






