STARZO AIEvals

Results

398 of 879 trials came back at exactly 1.0 — 45.3% of everything scored. Not a race.

The highest percentage on the board is 95.0%, from claude-code / claude-haiku-4-5, a row held back from ranking on 20 trials across 2 of 90 tasks. The row with the widest coverage, claude-code / claude-opus-4-8 on 56 tasks, resolved 268 of 591 for 45.3%. Those two numbers are measuring different corpora, which is why the board is ordered by trial count and not by score.

Read the intervals, not the ranking: every row ran a different number of trials on a different slice of the corpus, a row at or below 20 trials is tagged and held back from ranking, and at 10 trials per task one flipped verdict moves a category cell ten points.

45.3%
resolved overall
398 of 879 trials
5
agent lines
4 ranked, 1 held back
56/90
widest coverage
claude-code / claude-opus-4-8
±14pp
widest interval
n = 35

The board

The dot is the point estimate; the bar behind it is the 95% Wilson interval on that row's own trials. Under each bar, the same trials split five ways — a pass that gamed the grader and a failure the task caused are both scored, and both mislead. Every segment carries its label in the key; colour is a second cue, never the only one.

point estimate95% Wilson interval▵ held back from ranking
agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
45.3%41.4–49.4
268/591
56 of 90
261269
claude-codeclaude-opus-4-7
31.1%24.8–38.2
56/180
19 of 90
516742
codexgpt-5.5
84.9%72.9–92.1
45/53
12 of 90
37
claude-codeclaude-sonnet-4-6
28.6%16.3–45.1
10/35
4 of 90
61259
claude-codeclaude-haiku-4-5▵ held back
95.0%76.4–99.1
19/20
2 of 90
18

Outcome mix under each bar — five fixed classes, always in this order

Legitimate solveGOOD_SUCCESS
Reward-hackBAD_SUCCESS
Honest missGOOD_FAILURE
Task at faultBAD_FAILURE
Harness errorHARNESS_ERROR
Table view · leaderboard numbers
Status
claude-codeclaude-opus-4-845.3%41.449.426859156ranked
claude-codeclaude-opus-4-731.1%24.838.25618019ranked
codexgpt-5.584.9%72.992.1455312ranked
claude-codeclaude-sonnet-4-628.6%16.345.110354ranked
claude-codeclaude-haiku-4-595.0%76.499.119202held back

Published order: most trials first. Sortable heads carry an arrow.

Interval bounds are the 95% Wilson score interval on that row's own trials, in percentage points. Resolved counts a reward of exactly 1.0 — partial credit is not a pass.

Table view · outcome counts per agent
claude-code / claude-opus-4-826172693816591
claude-code / claude-opus-4-7512674218180
codex / gpt-5.537643353
claude-code / claude-sonnet-4-663125935
claude-code / claude-haiku-4-518010120

Published order: most trials first. Sortable heads carry an arrow.

Every category, and the agents inside it

One axis, one measure, 15 categories. The bar in the pass@1 column is the category rate; the columns to its right are that rate for a single agent line. Open a category name and the same rows are redrawn as a board, with the interval and the outcome mix each agent actually produced there.

pass@1 by category, then the same rate split by agent line. A dash means that pairing never ran. Select a category name to open its agent board in place.

CategoryTasksTrials pass@1claude-codeclaude-opus-4-8claude-codeclaude-opus-4-7codexgpt-5.5claude-codeclaude-sonnet-4-6claude-codeclaude-haiku-4-5
live88557.6%49/8550.0%35/70—100.0%5/5—90.0%9/10
live88035.0%28/8035.0%28/80————
live1918031.1%56/180—31.1%56/180———
live2015851.9%82/15847.7%62/130—71.4%20/28——
live58172.8%59/8172.8%59/81————
live1010543.8%46/10535.0%21/60—100.0%5/533.3%10/30100.0%10/10
preview315100.0%15/15——100.0%15/15——
preview11040.0%4/1040.0%4/10————
live66543.1%28/6546.7%28/60——0.0%0/5—
live33033.3%10/3033.3%10/30————
preview1100.0%0/100.0%0/10————
live22025.0%5/2025.0%5/20————
live22040.0%8/2040.0%8/20————
preview11040.0%4/1040.0%4/10————
preview11040.0%4/1040.0%4/10————
All categories9087945.3%398/87945.3%268/59131.1%56/18084.9%45/5328.6%10/3595.0%19/20

Status is derived from trial count, not task count: a category goes live at 20 scored trials and stays in preview below that.

The pipeline behind this report is available for client work.

Request a private evaluation team@starzo.ai