SyncValsfield notes

Results · run 2026-07-06 · SyncVals 0.1.0 · commit 2f94510

Five agent lines on one axis. Not a race.

Each row is pass@1 over the trials that row actually ran, and no two rows ran the same set of tasks. claude-code / claude-opus-4-8 was measured on 56 of 90 tasks; claude-code / claude-haiku-4-5 saw 2. Ranking the percentages against each other would mostly rank the task selections. Read each interval on its own terms.

  • Ordered by trials, not by score. The board follows trial count downwards, which keeps the widest evidence at the top and puts the thinnest at the bottom where it belongs.
  • A wide bar means few trials. The widest interval here spans 29 percentage points on 35 trials. That is arithmetic about sample size, not a claim about consistency.
  • 20 trials or fewer is held back. Those rows are drawn, tagged, and excluded from ranking — visible because hiding them would hide the coverage gap too.
5
agent lines
4 ranked, 1 held back
45.3%
resolved overall
398 of 879 trials
56/90
widest coverage
claude-code / claude-opus-4-8
±14pp
widest interval
n = 35

The board

The dot is the point estimate; the bar behind it is the 95% Wilson interval on that row's own trials. Under each bar, the same trials split five ways — a pass that gamed the grader and a failure the task caused are both scored, and both mislead. Every segment carries its label in the key; colour is a second cue, never the only one.

point estimate95% Wilson interval▵ held back from ranking
agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
45.3%41.4–49.4
268/591
56 of 90
Legitimate solve261Reward-hack7Honest miss269Task at fault38Harness error16
claude-codeclaude-opus-4-7
31.1%24.8–38.2
56/180
19 of 90
Legitimate solve51Reward-hack2Honest miss67Task at fault42Harness error18
codexgpt-5.5
84.9%72.9–92.1
45/53
12 of 90
Legitimate solve37Reward-hack6Honest miss4Task at fault3Harness error3
claude-codeclaude-sonnet-4-6
28.6%16.3–45.1
10/35
4 of 90
Legitimate solve6Reward-hack3Honest miss12Task at fault5Harness error9
claude-codeclaude-haiku-4-5▵ held back
95.0%76.4–99.1
19/20
2 of 90
Legitimate solve18Honest miss1Harness error1

Outcome mix under each bar — five fixed classes, always in this order

Legitimate solveGOOD_SUCCESS
Reward-hackBAD_SUCCESS
Honest missGOOD_FAILURE
Task at faultBAD_FAILURE
Harness errorHARNESS_ERROR
Table view · leaderboard numbers
AgentModelpass@1CI low CI highResolvedTrials TasksStatus
claude-codeclaude-opus-4-845.3%41.449.426859156ranked
claude-codeclaude-opus-4-731.1%24.838.25618019ranked
codexgpt-5.584.9%72.992.1455312ranked
claude-codeclaude-sonnet-4-628.6%16.345.110354ranked
claude-codeclaude-haiku-4-595.0%76.499.119202held back

Interval bounds are the 95% Wilson score interval on that row's own trials, in percentage points. Resolved counts a reward of exactly 1.0 — partial credit is not a pass.

Table view · outcome counts per agent
Agent / modelLegitimate solveReward-hackHonest missTask at faultHarness errorTrials
claude-code / claude-opus-4-826172693816591
claude-code / claude-opus-4-7512674218180
codex / gpt-5.537643353
claude-code / claude-sonnet-4-663125935
claude-code / claude-haiku-4-518010120

The same rows, split by category

Sub-boards stay closed so the page reads top to bottom; open one to see which agents were actually measured there. Intervals inside a category are wide by construction. A row of 10 trials can only land on one of 11 possible percentages, so treat these as coverage evidence first and performance evidence second.

Software Engineering live 57.6% pass@1 · 49/85 resolved · 8 tasks · 3 agent lines

Algorithms, data structures, bug-fixes, and API/systems implementation, graded against hidden test suites the agent never sees.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
50.0%38.6–61.4
35/70
7
Legitimate solve34Reward-hack1Honest miss28Task at fault6Harness error1
claude-codeclaude-haiku-4-5▵ held back
90.0%59.6–98.2
9/10
1
Legitimate solve8Honest miss1Harness error1
codexgpt-5.5▵ held back
100.0%56.6–100.0
5/5
1
Legitimate solve4Harness error1
Mechanical Engineering live 35.0% pass@1 · 28/80 resolved · 8 tasks · 1 agent line

Structural and numerical solvers (FD/FE, contact dynamics) in C++, checked against multi-binary hidden references.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
35.0%25.5–45.9
28/80
8
Legitimate solve28Honest miss50Harness error2
Cloud Operations live 31.1% pass@1 · 56/180 resolved · 19 tasks · 1 agent line

Provision, diagnose, and guardrail cloud infrastructure (AWS), verified against the deployed state.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-7
31.1%24.8–38.2
56/180
19
Legitimate solve51Reward-hack2Honest miss67Task at fault42Harness error18
Electrical Engineering live 51.9% pass@1 · 82/158 resolved · 20 tasks · 2 agent lines

RTL / digital-logic design checked under hardened simulation and formal-equivalence harnesses.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
47.7%39.3–56.2
62/130
13
Legitimate solve62Honest miss53Task at fault6Harness error9
codexgpt-5.5
71.4%52.9–84.7
20/28
7
Legitimate solve14Reward-hack5Honest miss4Task at fault3Harness error2
STEM live 72.8% pass@1 · 59/81 resolved · 5 tasks · 1 agent line

Quantitative reasoning across math, physics, and the natural sciences with deterministic, checkable answers.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
72.8%62.3–81.3
59/81
5
Legitimate solve58Reward-hack1Honest miss14Task at fault8
Game live 43.8% pass@1 · 46/105 resolved · 10 tasks · 4 agent lines

Interactive simulation and game-logic tasks graded on exact state transitions and rule fidelity.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
35.0%24.2–47.6
21/60
6
Legitimate solve19Reward-hack2Honest miss30Task at fault8Harness error1
claude-codeclaude-sonnet-4-6
33.3%19.2–51.2
10/30
3
Legitimate solve6Reward-hack3Honest miss7Task at fault5Harness error9
claude-codeclaude-haiku-4-5▵ held back
100.0%72.2–100.0
10/10
1
Legitimate solve10
codexgpt-5.5▵ held back
100.0%56.6–100.0
5/5
1
Legitimate solve5
Debugging preview 100.0% pass@1 · 15/15 resolved · 3 tasks · 1 agent line

Real-world bug-fix tasks distilled from open-source issues (esbuild, klauspost/compress, rust-lang/semver), graded against each project's own tests.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
codexgpt-5.5▵ held back
100.0%79.6–100.0
15/15
3
Legitimate solve14Reward-hack1
ML Engineering preview 40.0% pass@1 · 4/10 resolved · 1 task · 1 agent line

Train and tune models to a target metric on real engineering datasets, graded on held-out performance against a solved-reward threshold.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
Legitimate solve4Honest miss6
Post Training live 43.1% pass@1 · 28/65 resolved · 6 tasks · 2 agent lines

Long-horizon, from-scratch numpy ML research: implement reverse-mode autodiff and a training method (quantization-aware, adversarial-robust, semi-supervised, worst-group), then train to a sealed held-out metric. Tasks published blinded.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
46.7%34.6–59.1
28/60
6
Legitimate solve25Reward-hack3Honest miss31Harness error1
claude-codeclaude-sonnet-4-6▵ held back
0.0%0.0–43.4
0/5
1
Honest miss5
Scientific ML live 33.3% pass@1 · 10/30 resolved · 3 tasks · 1 agent line

Physics-informed and surrogate modeling: PDE forecasting and CFD/FEA prediction, graded on quantitative predictive accuracy.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8
33.3%19.2–51.2
10/30
3
Legitimate solve10Honest miss20
ML for Engineering preview 0.0% pass@1 · 0/10 resolved · 1 task · 1 agent line

Machine-learning-for-engineering tasks (DeepCAD command-history canonicalization and equivalence), graded on a continuous partial-credit score against verifier-owned labels rather than all-or-nothing pass/fail.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
0.0%0.0–27.8
0/10
1
Honest miss10
Data Science live 25.0% pass@1 · 5/20 resolved · 2 tasks · 1 agent line

Modeling and statistical inference on real datasets (federated learning, event studies), graded against held-out ground truth.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
25.0%11.2–46.9
5/20
2
Legitimate solve5Honest miss14Task at fault1
Data Science: Robustness live 40.0% pass@1 · 8/20 resolved · 2 tasks · 1 agent line

Bias-correction and outlier-robust analysis in R: recover trustworthy estimates from messy data, checked against reference results.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
40.0%21.9–61.3
8/20
2
Legitimate solve8Honest miss9Task at fault3
Product Data Science preview 40.0% pass@1 · 4/10 resolved · 1 task · 1 agent line

Applied product analytics: causal impact and decision analysis on real product datasets (R).

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
Legitimate solve4Honest miss2Task at fault4
Pharmacometrics preview 40.0% pass@1 · 4/10 resolved · 1 task · 1 agent line

Population PK/PD modeling with nonlinear mixed-effects: fit drug-exposure models and recover the correct parameters.

agent / model
0255075100
pass@1 · 95% CI
resolved
tasks
claude-codeclaude-opus-4-8▵ held back
40.0%16.8–68.7
4/10
1
Legitimate solve4Honest miss2Task at fault2Harness error2

pass@1 per category

One axis, one measure, fifteen categories. The bar in the pass@1 column is the category rate; the columns to its right are that rate for a single agent line, which is where the sample sizes get small. Check the resolved count under a cell before quoting the percentage above it — at 10 trials, one flipped verdict moves the figure ten points.

pass@1 by category, then the same rate split by agent line. A dash means that pairing never ran.
CategoryTasksTrials pass@1claude-codeclaude-opus-4-8claude-codeclaude-opus-4-7codexgpt-5.5claude-codeclaude-sonnet-4-6claude-codeclaude-haiku-4-5
Software Engineeringlive88557.6%49/8550.0%35/70100.0%5/590.0%9/10
Mechanical Engineeringlive88035.0%28/8035.0%28/80
Cloud Operationslive1918031.1%56/18031.1%56/180
Electrical Engineeringlive2015851.9%82/15847.7%62/13071.4%20/28
STEMlive58172.8%59/8172.8%59/81
Gamelive1010543.8%46/10535.0%21/60100.0%5/533.3%10/30100.0%10/10
Debuggingpreview315100.0%15/15100.0%15/15
ML Engineeringpreview11040.0%4/1040.0%4/10
Post Traininglive66543.1%28/6546.7%28/600.0%0/5
Scientific MLlive33033.3%10/3033.3%10/30
ML for Engineeringpreview1100.0%0/100.0%0/10
Data Sciencelive22025.0%5/2025.0%5/20
Data Science: Robustnesslive22040.0%8/2040.0%8/20
Product Data Sciencepreview11040.0%4/1040.0%4/10
Pharmacometricspreview11040.0%4/1040.0%4/10
All categories9087945.3%398/87945.3%268/59131.1%56/18084.9%45/5328.6%10/3595.0%19/20

Status is derived from trial count, not task count: a category goes live at 20 scored trials and stays in preview below that. Category rates mix agents together, so a category that only one agent attempted is reporting that agent, not the field.