Results
398 of 879 trials came back at exactly 1.0 — 45.3% of everything scored. Not a race.
The highest percentage on the board is 95.0%, from claude-code / claude-haiku-4-5, a row held back from ranking on 20 trials across 2 of 90 tasks. The row with the widest coverage, claude-code / claude-opus-4-8 on 56 tasks, resolved 268 of 591 for 45.3%. Those two numbers are measuring different corpora, which is why the board is ordered by trial count and not by score.
Read the intervals, not the ranking: every row ran a different number of trials on a different slice of the corpus, a row at or below 20 trials is tagged and held back from ranking, and at 10 trials per task one flipped verdict moves a category cell ten points.
The board
The dot is the point estimate; the bar behind it is the 95% Wilson interval on that row's own trials. Under each bar, the same trials split five ways — a pass that gamed the grader and a failure the task caused are both scored, and both mislead. Every segment carries its label in the key; colour is a second cue, never the only one.
Outcome mix under each bar — five fixed classes, always in this order
Table view · leaderboard numbers
| Status | ||||||||
|---|---|---|---|---|---|---|---|---|
| claude-code | claude-opus-4-8 | 45.3% | 41.4 | 49.4 | 268 | 591 | 56 | ranked |
| claude-code | claude-opus-4-7 | 31.1% | 24.8 | 38.2 | 56 | 180 | 19 | ranked |
| codex | gpt-5.5 | 84.9% | 72.9 | 92.1 | 45 | 53 | 12 | ranked |
| claude-code | claude-sonnet-4-6 | 28.6% | 16.3 | 45.1 | 10 | 35 | 4 | ranked |
| claude-code | claude-haiku-4-5 | 95.0% | 76.4 | 99.1 | 19 | 20 | 2 | held back |
Published order: most trials first. Sortable heads carry an arrow.
Interval bounds are the 95% Wilson score interval on that row's own trials, in percentage points. Resolved counts a reward of exactly 1.0 — partial credit is not a pass.
Table view · outcome counts per agent
| claude-code / claude-opus-4-8 | 261 | 7 | 269 | 38 | 16 | 591 |
| claude-code / claude-opus-4-7 | 51 | 2 | 67 | 42 | 18 | 180 |
| codex / gpt-5.5 | 37 | 6 | 4 | 3 | 3 | 53 |
| claude-code / claude-sonnet-4-6 | 6 | 3 | 12 | 5 | 9 | 35 |
| claude-code / claude-haiku-4-5 | 18 | 0 | 1 | 0 | 1 | 20 |
Published order: most trials first. Sortable heads carry an arrow.
Every category, and the agents inside it
One axis, one measure, 15 categories. The bar in the pass@1 column is the category rate; the columns to its right are that rate for a single agent line. Open a category name and the same rows are redrawn as a board, with the interval and the outcome mix each agent actually produced there.
pass@1 by category, then the same rate split by agent line. A dash means that pairing never ran. Select a category name to open its agent board in place.
| Category | Tasks | Trials | pass@1 | claude-codeclaude-opus-4-8 | claude-codeclaude-opus-4-7 | codexgpt-5.5 | claude-codeclaude-sonnet-4-6 | claude-codeclaude-haiku-4-5 |
|---|---|---|---|---|---|---|---|---|
| live | 8 | 85 | 57.6%49/85 | 50.0%35/70 | — | 100.0%5/5 | — | 90.0%9/10 |
Algorithms, data structures, bug-fixes, and API/systems implementation, graded against hidden test suites the agent never sees. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
50.0%38.6–61.4
35/70
7
claude-codeclaude-haiku-4-5▵ held back
90.0%59.6–98.2
9/10
1
codexgpt-5.5▵ held back
100.0%56.6–100.0
5/5
1
| ||||||||
| live | 8 | 80 | 35.0%28/80 | 35.0%28/80 | — | — | — | — |
Structural and numerical solvers (FD/FE, contact dynamics) in C++, checked against multi-binary hidden references. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
35.0%25.5–45.9
28/80
8
| ||||||||
| live | 19 | 180 | 31.1%56/180 | — | 31.1%56/180 | — | — | — |
Provision, diagnose, and guardrail cloud infrastructure (AWS), verified against the deployed state. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-7
31.1%24.8–38.2
56/180
19
| ||||||||
| live | 20 | 158 | 51.9%82/158 | 47.7%62/130 | — | 71.4%20/28 | — | — |
RTL / digital-logic design checked under hardened simulation and formal-equivalence harnesses. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
47.7%39.3–56.2
62/130
13
codexgpt-5.5
71.4%52.9–84.7
20/28
7
| ||||||||
| live | 5 | 81 | 72.8%59/81 | 72.8%59/81 | — | — | — | — |
Quantitative reasoning across math, physics, and the natural sciences with deterministic, checkable answers. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
72.8%62.3–81.3
59/81
5
| ||||||||
| live | 10 | 105 | 43.8%46/105 | 35.0%21/60 | — | 100.0%5/5 | 33.3%10/30 | 100.0%10/10 |
Interactive simulation and game-logic tasks graded on exact state transitions and rule fidelity. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
35.0%24.2–47.6
21/60
6
claude-codeclaude-sonnet-4-6
33.3%19.2–51.2
10/30
3
claude-codeclaude-haiku-4-5▵ held back
100.0%72.2–100.0
10/10
1
codexgpt-5.5▵ held back
100.0%56.6–100.0
5/5
1
| ||||||||
| preview | 3 | 15 | 100.0%15/15 | — | — | 100.0%15/15 | — | — |
Real-world bug-fix tasks distilled from open-source issues (esbuild, klauspost/compress, rust-lang/semver), graded against each project's own tests. agent / model 0255075100 pass@1 · 95% CI resolved tasks codexgpt-5.5▵ held back
100.0%79.6–100.0
15/15
3
| ||||||||
| preview | 1 | 10 | 40.0%4/10 | 40.0%4/10 | — | — | — | — |
Train and tune models to a target metric on real engineering datasets, graded on held-out performance against a solved-reward threshold. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
| ||||||||
| live | 6 | 65 | 43.1%28/65 | 46.7%28/60 | — | — | 0.0%0/5 | — |
Long-horizon, from-scratch numpy ML research: implement reverse-mode autodiff and a training method (quantization-aware, adversarial-robust, semi-supervised, worst-group), then train to a sealed held-out metric. Tasks published blinded. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
46.7%34.6–59.1
28/60
6
claude-codeclaude-sonnet-4-6▵ held back
0.0%0.0–43.4
0/5
1
| ||||||||
| live | 3 | 30 | 33.3%10/30 | 33.3%10/30 | — | — | — | — |
Physics-informed and surrogate modeling: PDE forecasting and CFD/FEA prediction, graded on quantitative predictive accuracy. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8
33.3%19.2–51.2
10/30
3
| ||||||||
| preview | 1 | 10 | 0.0%0/10 | 0.0%0/10 | — | — | — | — |
Machine-learning-for-engineering tasks (DeepCAD command-history canonicalization and equivalence), graded on a continuous partial-credit score against verifier-owned labels rather than all-or-nothing pass/fail. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
0.0%0.0–27.8
0/10
1
| ||||||||
| live | 2 | 20 | 25.0%5/20 | 25.0%5/20 | — | — | — | — |
Modeling and statistical inference on real datasets (federated learning, event studies), graded against held-out ground truth. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
25.0%11.2–46.9
5/20
2
| ||||||||
| live | 2 | 20 | 40.0%8/20 | 40.0%8/20 | — | — | — | — |
Bias-correction and outlier-robust analysis in R: recover trustworthy estimates from messy data, checked against reference results. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
40.0%21.9–61.3
8/20
2
| ||||||||
| preview | 1 | 10 | 40.0%4/10 | 40.0%4/10 | — | — | — | — |
Applied product analytics: causal impact and decision analysis on real product datasets (R). agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
| ||||||||
| preview | 1 | 10 | 40.0%4/10 | 40.0%4/10 | — | — | — | — |
Population PK/PD modeling with nonlinear mixed-effects: fit drug-exposure models and recover the correct parameters. agent / model 0255075100 pass@1 · 95% CI resolved tasks claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
| ||||||||
| All categories | 90 | 879 | 45.3%398/879 | 45.3%268/591 | 31.1%56/180 | 84.9%45/53 | 28.6%10/35 | 95.0%19/20 |
Status is derived from trial count, not task count: a category goes live at 20 scored trials and stays in preview below that.
The pipeline behind this report is available for client work.
Request a private evaluation team@starzo.ai