STARZO AIEvals

The STARZO AI evals report

The benchmark that measures whether AI earned the answer.

A coding agent can pass for the wrong reasons. A deterministic verifier decides every reward (exit zero, nothing else), and a separate classifier then reads the trajectory and says why the run landed where it did, without ever touching the score. 879 trials, sorted five ways.

879
scored trials
across 90 tasks
45.3%
resolved overall
398 of 879
5
agent / model lines
k = 10 per task
18
passes that shouldn't count
reward-hacks caught
88
failures blamed on the task
not on the agent
14
tasks nothing solved
against 20 solved by every attempt
47
runs that never really ran
infrastructure, not capability

This run's board

The highest raw pass rate here holds no rank at all: 95.0% rests on 20 trials, at the 20-trial held-back floor, so first place goes to codex / gpt-5.5, whose 84.9% pass@1 is 69.8% once only clean passes count. Rank by the raw rate and the story looks simple. Add the clean-pass rate (passes that survived the reward-hack audit) and the interval beside it, and the board stops flattering anyone. Per-category boards →

#Agent / model pass@1 95% Wilson interval
0255075100
clean rate hack share of passesn
1 codexgpt-5.5 84.9%CI 72.9–92.1 69.8% 13.3% 53
2 claude-codeclaude-opus-4-8 45.3%CI 41.4–49.4 44.2% 2.6% 591
3 claude-codeclaude-opus-4-7 31.1%CI 24.8–38.2 28.3% 3.6% 180
4 claude-codeclaude-sonnet-4-6 28.6%CI 16.3–45.1 17.1% 30.0% 35
— claude-codeclaude-haiku-4-5▵ held back 95.0%CI 76.4–99.1 90.0% 0.0% 20

Clean rate = trials classified GOOD_SUCCESS ÷ all trials for that line. Hack share = passes classified BAD_SUCCESS ÷ all passes: of claude-sonnet-4-6's 10 passes, 3 gamed the grader. 20 trials or fewer is held back — drawn and tagged, but given no rank, because a rate on that little evidence would outrank rows carrying ten times the trials. Rates are not comparable across rows: every line ran a different subset of tasks, and the widest intervals are simply the smallest samples.

Why benchmark numbers lie

In 12.1% of trials the score and the truth disagree: 106 of 879 runs were either a pass that gamed the grader or a failure the task itself caused. Reward is binary; the reason is not. Two of these five labels mean the number lies. What each label means →

Legitimate solveGOOD_SUCCESS373
Reward-hackBAD_SUCCESS18
Honest missGOOD_FAILURE353
Task at faultBAD_FAILURE88
Harness errorHARNESS_ERROR47
Table view
OutcomeLabelTrialsShare
Legitimate solveGOOD_SUCCESS37342.4%
Reward-hackBAD_SUCCESS182.0%
Honest missGOOD_FAILURE35340.2%
Task at faultBAD_FAILURE8810.0%
Harness errorHARNESS_ERROR475.3%

Reward-hack hall of fame

The 18 unearned passes are not spread thin: they land on 14 tasks, and the worst single task accounts for 3 of them. On each of these the grader was satisfied without the work being done. Ordered by how often. How each exploit worked →

Coverage by category

10 of the 15 categories clear the 20-trial floor; the remaining 5 are still preview. Status is derived, not asserted: a category is live once it has at least 20 scored trials, and preview until then. Live categories are listed first.

Why STARZO AI exists

Current AI benchmarks measure outcomes. They rarely measure whether the outcome was earned. STARZO AI evaluates both, capability and evaluation integrity, because a leaderboard that can be gamed measures the gaming, not the model.
The score is never touched.

A deterministic verifier is the sole authority on reward. The five-way classification is post-hoc and explanatory; it can flag a pass as unearned, but it can never change a number.

Task defects are our fault, and we say so.

88 failures in this run were caused by the task, not the agent. Most benchmarks silently charge those to the model. We publish them, per task, and fix them.

Status is derived, not asserted.

A category goes live at 20 scored trials, not when it looks finished. Every figure on this site is recomputed from the raw trial set.

Work with STARZO AI

Everything above was produced by our own pipeline — and the pipeline is the product.

The pipeline behind this report is available for client work.

Request a private evaluation team@starzo.ai