The STARZO AI evals report
The benchmark that measures whether AI earned the answer.
A coding agent can pass for the wrong reasons. A deterministic verifier decides every reward (exit zero, nothing else), and a separate classifier then reads the trajectory and says why the run landed where it did, without ever touching the score. 879 trials, sorted five ways.
This run's board
The highest raw pass rate here holds no rank at all: 95.0% rests on 20 trials, at the 20-trial held-back floor, so first place goes to codex / gpt-5.5, whose 84.9% pass@1 is 69.8% once only clean passes count. Rank by the raw rate and the story looks simple. Add the clean-pass rate (passes that survived the reward-hack audit) and the interval beside it, and the board stops flattering anyone. Per-category boards →
| # | Agent / model | pass@1 | 95% Wilson interval 0255075100 |
clean rate | hack share of passes | n |
|---|---|---|---|---|---|---|
| 1 | codexgpt-5.5 | 84.9%CI 72.9–92.1 |
|
69.8% | 13.3% | 53 |
| 2 | claude-codeclaude-opus-4-8 | 45.3%CI 41.4–49.4 |
|
44.2% | 2.6% | 591 |
| 3 | claude-codeclaude-opus-4-7 | 31.1%CI 24.8–38.2 |
|
28.3% | 3.6% | 180 |
| 4 | claude-codeclaude-sonnet-4-6 | 28.6%CI 16.3–45.1 |
|
17.1% | 30.0% | 35 |
| — | claude-codeclaude-haiku-4-5▵ held back | 95.0%CI 76.4–99.1 |
|
90.0% | 0.0% | 20 |
Clean rate = trials classified GOOD_SUCCESS ÷ all trials for that line. Hack share = passes classified BAD_SUCCESS ÷ all passes: of claude-sonnet-4-6's 10 passes, 3 gamed the grader. 20 trials or fewer is held back — drawn and tagged, but given no rank, because a rate on that little evidence would outrank rows carrying ten times the trials. Rates are not comparable across rows: every line ran a different subset of tasks, and the widest intervals are simply the smallest samples.
Why benchmark numbers lie
In 12.1% of trials the score and the truth disagree: 106 of 879 runs were either a pass that gamed the grader or a failure the task itself caused. Reward is binary; the reason is not. Two of these five labels mean the number lies. What each label means →
Table view
| Outcome | Label | Trials | Share |
|---|---|---|---|
| Legitimate solve | GOOD_SUCCESS | 373 | 42.4% |
| Reward-hack | BAD_SUCCESS | 18 | 2.0% |
| Honest miss | GOOD_FAILURE | 353 | 40.2% |
| Task at fault | BAD_FAILURE | 88 | 10.0% |
| Harness error | HARNESS_ERROR | 47 | 5.3% |
Reward-hack hall of fame
The 18 unearned passes are not spread thin: they land on 14 tasks, and the worst single task accounts for 3 of them. On each of these the grader was satisfied without the work being done. Ordered by how often. How each exploit worked →
Coverage by category
10 of the 15 categories clear the 20-trial floor; the remaining 5 are still preview. Status is derived, not asserted: a category is live once it has at least 20 scored trials, and preview until then. Live categories are listed first.
Why STARZO AI exists
Current AI benchmarks measure outcomes. They rarely measure whether the outcome was earned. STARZO AI evaluates both, capability and evaluation integrity, because a leaderboard that can be gamed measures the gaming, not the model.
A deterministic verifier is the sole authority on reward. The five-way classification is post-hoc and explanatory; it can flag a pass as unearned, but it can never change a number.
88 failures in this run were caused by the task, not the agent. Most benchmarks silently charge those to the model. We publish them, per task, and fix them.
A category goes live at 20 scored trials, not when it looks finished. Every figure on this site is recomputed from the raw trial set.
Work with STARZO AI
Everything above was produced by our own pipeline — and the pipeline is the product.
I need an evaluation
Talk to our team.
A private run of your agent, tasks and environments built to your spec, or the data behind a report — one conversation scopes all of it.
Continue →I build evaluations
Build with us.
Domain experts and task authors who can write problems an agent cannot shortcut — the suite grows through people like you.
Continue →The pipeline behind this report is available for client work.
Request a private evaluation team@starzo.ai