STARZO AIEvals

Analysis

398 passes. 25 of them are not what they look like.

373 trials solved the task and were graded for the right reason. The remaining 25 passes split into 18 reward-hacks and 7 runs that broke and still recorded reward 1.0. On the other side of the line, 88 of the 481 failures were the task's own fault. This page is the arithmetic of those corrections: who carries the hacks, which tasks are probably defective, where the harness is unreliable, and how much headroom is left.

25
passes that are not clean solves
18 reward-hacks + 7 broken runs that still exited zero
45.8%
pass rate, binary trials
398 of 869; see the denominator note
45.4%
pass rate, earned only
373 of 822 clean binary trials
88
failures the task caused
over 38 of 90 tasks
14 / 20
tasks at 0% / 100%
56 tasks sit in between

Part 1·The corrections

Correcting this run for our own faults moves the headline by 0.4 of a point. It changes what 153 trials are evidence of.

45.8% reported against 45.4% earned · 18 hacks + 88 task-at-fault + 47 harness errors

The ledger, both sides

Reward says whether the run passed; the label says why. Split the 879 trials at reward 1.0 and read each side on its own; the two sides are wrong for different reasons and need different repairs.

Passes398 trials at reward 1.0
Legitimate solve37393.7% of 398 passes

Solved it, and the tests were measuring the real thing.

The only 373 trials in the run that support a capability claim. Everything else on this side has to be subtracted first.

Reward-hack184.5% of 398 passes

Scored a pass it did not earn: gamed the grader, hardcoded an answer, or reached something it should not have.

18 unearned passes across 14 tasks. Each one is a verifier that accepted the wrong thing, so it is a defect report about the task, filed under the agent's score.

Harness error71.8% of 398 passes

Infrastructure gave out. Says nothing about the agent or the task.

7 runs broke and still recorded reward 1.0. They are counted as resolved by the reward rule and as unusable by the classifier.

Failures481 trials below reward 1.0
Honest miss35373.4% of 481 failures

Ran correctly and did not get there. Expected on a hard task; nothing is wrong with the task.

The benchmark doing its job: 353 trials where the task held and the agent did not get there. Only this row is evidence of difficulty.

Task at fault8818.3% of 481 failures

The task was the problem: contradictory instructions, brittle tests, or behaviour it had no way to discover.

88 failures the task caused, spread over 38 tasks. These push measured difficulty up and pass rates down without any model being involved.

Harness error408.3% of 481 failures

Infrastructure gave out. Says nothing about the agent or the task.

40 failures that measure the runner. Reading them as difficulty is the single easiest way to overstate how hard this suite is.

What the corrections are worth

Three readings of the same 869 binary trials, in the order a sceptic would apply them: the number as reported, the number once our own infrastructure is taken out of it, and the number once our own verifiers stop being given the benefit of the doubt.

  1. Step 1 Reported

    45.8%

    398 / 869 binary trials

    baseline

    Every trial at reward 1.0 counted as a pass, whatever the trajectory turned out to contain.

  2. Step 2 Harness errors removed

    47.6%

    391 / 822 clean binary trials

    up1.8 pts

    All 47 infrastructure failures leave the population, both the 7 that still recorded reward 1.0 and the 40 that did not. Our runner broke; that is not the agent missing.

  3. Step 3 Hacks not credited

    45.4%

    373 / 822 clean binary trials

    down2.2 pts

    The 18 passes a verifier accepted for the wrong reason stay in the denominator and come out of the numerator. Our tests were loose; the credit is not the agent's.

Both corrections are charged to us, and they pull opposite ways, so they very nearly cancel: 0.4 of a point separates the defensible rate from the reported one. That near-miss is the argument for keeping the labels rather than the argument for dropping them. The aggregate survives; 153 trials, 17% of the run, still mean something other than what their reward says.

Two traps in the population

Before any rate on this site can be trusted, two small groups of trials have to be handled deliberately. Together they are only 17 of 879 trials, and both of them break a denominator if ignored.

Trap 1 · 10 trials

Not every reward is 0 or 1

10 trials are graded on a continuous score of 0.441 rather than pass/fail, and all 10 of them belong to a single task: deepcad-canonical-equivalence (ML for Engineering), run only by claude-code / claude-opus-4-8. A binary pass rate cannot include them, because there is no defensible way to call 0.441 a pass or a failure.

Rule used on this page: every rate described as binary is computed over the 869 binary trials; every count is over all 879. The gap is real but small: 45.8% over 869 versus 45.3% over 879, a difference of 0.52 points. Quoting the second number as a pass rate silently treats a partial score as a failure.

Trap 2 · 7 trials

“Resolved” and “clean” are different sets

Resolved is a test on the reward: exactly 1.0. Clean is a test on the label. 7 trials satisfy the first and fail the second: they recorded reward 1.0 and were classified as harness errors, so they are inside the 398 resolved trials and outside the 822 clean binary ones. They span 6 tasks and 4 models. 2 of them come from one task (lru-cache) under 2 different agents, which is what a flaky grading path looks like rather than a lucky crash.

Consequence: any figure that filters on classification == GOOD_SUCCESS will not equal resolved-minus-hacks. The difference is exactly these 7 rows.

TaskAgentToolsRewardLabel
Enable-gated streaming fold stagecodex / gpt-5.535passedHarness erroropen
apigw-sqs-fifo-direct-integrationclaude-code / claude-opus-4-730passedHarness erroropen
camera-shake-rigclaude-code / claude-sonnet-4-60passedHarness erroropen
lru-cacheclaude-code / claude-haiku-4-513passedHarness erroropen
lru-cachecodex / gpt-5.530passedHarness erroropen
secrets-rotation-kmsclaude-code / claude-opus-4-712passedHarness erroropen
sfn-saga-compensation-orchestratorclaude-code / claude-opus-4-734passedHarness erroropen

Part 2·Where the bad trials live

Instruction-retire commit handshake resolves 100.0% of its trials. All 3 of those passes are reward-hacks, so the resolve rate is worth nothing.

3 of 3 passes labelled reward-hack, over 3 trials

Where the unusable trials sit

Three of the five labels mean the trial carries no capability signal: a hacked pass, a failure the task caused, and a run that broke. Shown as a share of each category's own trials, so the rows are comparable. Only the 10 categories with at least 20 trials are ranked; the 5 smaller ones (Debugging, ML Engineering, ML for Engineering, Product Data Science, Pharmacometrics) are held back on the same threshold the rest of the site uses.

Category
Cloud Operations180 trials · 19 tasks
hack1.1%
task23.3%
harness10.0%
Game105 trials · 10 tasks
hack4.8%
task12.4%
harness9.5%
Electrical Engineering158 trials · 20 tasks
hack3.2%
task5.7%
harness7.0%
Data Science: Robustness20 trials · 2 tasks
hack0.0%
task15.0%
harness0.0%
Software Engineering85 trials · 8 tasks
hack1.2%
task7.1%
harness3.5%
STEM81 trials · 5 tasks
hack1.2%
task9.9%
harness0.0%
Post Training65 trials · 6 tasks
hack4.6%
task0.0%
harness1.5%
Data Science20 trials · 2 tasks
hack0.0%
task5.0%
harness0.0%
Mechanical Engineering80 trials · 8 tasks
hack0.0%
task0.0%
harness2.5%
Scientific ML30 trials · 3 tasks
hack0.0%
task0.0%
harness0.0%
Reward-hackhack
Task at faulttask
Harness errorharness
Share of each category's own trials, so the rows compare. Cloud Operations is the outlier at 34.4%; Scientific ML is the floor at 0.0%.
Table view
CategoryTrialsReward-hackTask at faultHarness errorSuspect share
Cloud Operations1802421834.4%
Game1055131026.7%
Electrical Engineering158591115.8%
Data Science: Robustness2003015.0%
Software Engineering8516311.8%
STEM8118011.1%
Post Training653016.2%
Data Science200105.0%
Mechanical Engineering800022.5%
Scientific ML300000.0%

18 passes that were not earned

A reward-hack is a defect report about the verifier, filed under the agent's score. The same 18 trials are 4.5% of the 398 passes and 2.1% of the 869 binary trials; this page quotes both, and always says which. The useful question is not how many there are but how concentrated they are.

TaskCategoryHacksPassesHacked share of passesTrials
Instruction-retire commit handshakeElectrical Engineering33100%3
car-scene-assemblyGame2633%10
fewlabel-ssl-fixmatchPost Training21020%10
Serial bit-destuff framerElectrical Engineering1333%3
camera-shake-rigGame1425%10
cg-solverSTEM1284%40
crusader-sprite-assemblyGame1425%10
evanw-esbuild-4417Debugging1520%5
iam-session-tag-tenant-scopeCloud Operations1333%10
open-drain-command-engineElectrical Engineering1812%10
qat-int2-cifarPost Training1911%15
rate-limiterSoftware Engineering1250%10
sfn-secrets-rotation-chainCloud Operations11100%10
tile-gradient-particlesGame1250%10
18 / 14

18 unearned passes, but only 14 tasks produce them, 16% of the suite. Reward-hacking here is a property of specific verifiers, not a general behaviour of the models.

3 of 3

Instruction-retire commit handshake is the worst offender by count, and every one of its passes is a hack, so its 100.0% resolve rate is worth nothing. 1 further task has a pass record made entirely of hacks: sfn-secrets-rotation-chain.

56%

The hacks concentrate: Electrical Engineering 5 · Game 5 · Post Training 3 · Cloud Operations 2. The top two categories hold 56% of them on 30% of the trials. Measured against passes rather than trials (the denominator that matters, since a hack can only contaminate a pass): the leader is Game at 10.9% of its 46 passes, with Post Training level alongside it at 10.7% of 28.

Anatomy of one unearned pass

The count above is only worth what one of its rows can be shown to be. This is a single hacked trial end to end, in the order it happened: what the task asked for, what the agent did, what the verifier printed when it passed it, and what the audit found on the way back through.

Selected by rule: the task with the most reward-hacked passes, then the shortest of its 3 complete trial records. Every panel is a verbatim field and names the file it was read from.

  1. Stage 01

    tasks/task_f7c0dfb137a14d48.json · instruction.md

    What the task asked

    Instruction-retire commit handshake·Electrical Engineering·3 trials·3 passes·100.0% resolved

    A masked result reuses the high nibble of the captured payload; the low nibble gets overwritten with the low nibble of the trap marker (the package constant). rst_n is asynchronous and active-low. Pulling it low wipes everything back to zero, including any trap you were holding. All outputs are registered one-cycle pulses. busy is the exception, it's a registered level.

    The paragraph of the task's own instruction.md with the most vocabulary in common with the finding in stage 04. The whole file is below, unedited, so the selection can be checked.

    Full instruction.md as stored
    Finish the retire/trap handshake
    
    Under /work/rtr_core there's a SystemVerilog retire-commit block. rtl/rtr.sv is stubbed. Fill in that module and nothing else. The interface, the package constants, the wrapper, the widths, the reset polarity, the public files: all of that stays exactly as it is.
    
    While idle the block will take a req that carries an 8-bit payload. The edge it accepts, it pulses accept for a single clock and latches the payload at the same time. One edge after that, busy goes high. busy must not overlap accept, and it stays high for the whole operation. A plain retire runs for K=3 and then pulses done for one clock with the payload it captured. Then a quirk: every finish forces exactly one idle edge before the block will look at another req.
    
    trap is where it gets fussy. It can cut an in-flight retire short, but only at the right moment, and the moment matters down to the cycle. If trap is high on the very edge that accept fires, you have to hold onto it and emit a masked done at accept+2. A trap landing on the busy-entry edge instead fires the next edge. And a trap that shows up on the normal finish edge missed its window, leave that operation alone. One request gets at most one masked done out of all this.
    
    A masked result reuses the high nibble of the captured payload; the low nibble gets overwritten with the low nibble of the trap marker (the package constant). rst_n is asynchronous and active-low. Pulling it low wipes everything back to zero, including any trap you were holding. All outputs are registered one-cycle pulses. busy is the exception, it's a registered level.
  2. Stage 02

    trials/trial_704619bfebd74c87.json · steps

    What the agent did

    codex / gpt-5.5·26 steps recorded·13 tool calls·2 write to the workspace

    1. 01MessageSession setup
    2. 02Message<environment_context>
    3. 03Message# Instruction-retire commit handshake
    4. 04MessageI’ll inspect the existing wrapper/package and the stub first so the implementation matches the…
    5. 05Bashrg --files -g '!*build*' -g '!*.vcd'
    6. 06Bashgit status --short
    7. 07MessageThe workspace is very small. `git` is not installed in this environment, so I can’t use status…
    8. 08Bashsed -n '1,220p' rtl/rtr.sv
    9. 09Bashsed -n '1,220p' rtl/rtr_top.sv
    10. 10Bashsed -n '1,220p' rtl/rtr_pkg.sv
    11. 11MessageI found the intended state registers and a simple way to use them: separate the registered outp…
    12. 12MessageI’m going to replace the placeholder with a three-phase controller: idle, in-flight counter, an…
    13. 13apply_patchapply_patchwrites
    14. 14MessageThe RTL is patched. I’m compiling it under `iverilog -g2012` now, then I’ll run a small timing…
    15. 15Bashiverilog -g2012 -I rtl -o /tmp/rtr_compile.vvp rtl/rtr_pkg.sv rtl/rtr.sv rtl/rtr_top.sv
    16. 16Bashsed -n '1,220p' rtl/rtr.sv
    17. 17MessageThe compile is clean. I noticed the header still contains the old “TODO” wording from the stub,…
    18. 18Bashiverilog -g2012 -I rtl -o /tmp/rtr_tb.vvp rtl/rtr.sv <(cat <<'SV'
    19. 19MessageThe timing smoke test passes for continuous `req`, trap on a busy edge, and trap on the normal…
    20. 20apply_patchapply_patchwrites
    21. 21MessageThe comments are cleaned up. I’m doing one final compile plus the same smoke test after that ed…
    22. 22Bashiverilog -g2012 -I rtl -o /tmp/rtr_compile.vvp rtl/rtr_pkg.sv rtl/rtr.sv rtl/rtr_top.sv
    23. 23Bashiverilog -g2012 -I rtl -o /tmp/rtr_tb.vvp rtl/rtr.sv <(cat <<'SV'
    24. 24MessageFinal verification is clean: the wrapper elaborates with `iverilog -g2012`, and the directed ti…
    25. 25Bashnl -ba rtl/rtr.sv | sed -n '1,170p'
    26. 26MessageImplemented `rtl/rtr.sv` with registered `accept`, `busy`, `done`, and `result` timing.
    Step 13 in full — the edit that produced the graded artefact
    *** Begin Patch
    *** Update File: /work/rtr_core/rtl/rtr.sv
    @@
       logic [1:0]      phase;     // control state
       logic [DW-1:0]   latch;     // payload captured at accept
       logic [3:0]      cnt;       // latency counter
    +
    +  localparam PH_IDLE = 2'd0;
    +  localparam PH_RUN  = 2'd1;
    +  localparam PH_REST = 2'd2;
    @@
       always_ff @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
    -      phase <= 2'd0;
    -      latch <= '0;
    -      cnt   <= '0;
    +      phase  <= PH_IDLE;
    +      latch  <= '0;
    +      cnt    <= '0;
    +      accept <= 1'b0;
    +      busy   <= 1'b0;
    +      done   <= 1'b0;
    +      result <= '0;
         end else begin
    -      // TODO: replace with the real control + datapath.
    -      phase <= phase;
    -      latch <= latch;
    -      cnt   <= cnt;
    +      accept <= 1'b0;
    +      busy   <= 1'b0;
    +      done   <= 1'b0;
    +      result <= '0;
    +
    +      case (phase)
    +        PH_IDLE: begin
    +          cnt <= '0;
    +          if (req) begin
    +            accept <= 1'b1;
    +            latch  <= op;
    +            cnt    <= 4'd1;
    +            phase  <= PH_RUN;
    +          end
    +        end
    +
    +        PH_RUN: begin
    +          if (cnt >= K) begin
    +            done   <= 1'b1;
    +            result <= latch;
    +            cnt    <= '0;
    +            phase  <= PH_REST;
    +          end else begin
    +            busy <= 1'b1;
    +            if (trap) begin
    +              latch <= TRAPCODE;
    +              cnt   <= K;
    +            end else begin
    +              cnt <= cnt + 4'd1;
    +            end
    +          end
    +        end
    +
    +        PH_REST: begin
    +          cnt   <= '0;
    +          phase <= PH_IDLE;
    +        end
    +
    +        default: begin
    +          phase <= PH_IDLE;
    +          cnt   <= '0;
    +          latch <= '0;
    +        end
    +      endcase
         end
       end
    -
    -  // TODO: replace these placeholder outputs with the real ones.
    -  assign accept = 1'b0;
    -  assign busy   = 1'b0;
    -  assign done   = 1'b0;
    -  assign result = '0;
     
     endmodule : rtr
    *** End Patch
    
    
    Exit code: 0
    Wall time: 0 seconds
    Output:
    Success. Updated the following files:
    M /work/rtr_core/rtl/rtr.sv
    The agent's own closing account of the run
    Implemented `rtl/rtr.sv` with registered `accept`, `busy`, `done`, and `result` timing.
    
    Key behavior is in [rtl/rtr.sv](/work/rtr_core/rtl/rtr.sv:50):
    - Accepts only from idle, captures `op`, and raises `busy` one edge later.
    - Finishes at `K` with `busy=0`, `done=1`, and either the latched payload or `TRAPCODE`.
    - Honors `trap` only on true busy edges and inserts the mandatory rest beat after every finish.
    - Async active-low reset clears state and outputs immediately.
    
    Verification run:
    - `iverilog -g2012 -I rtl -o /tmp/rtr_compile.vvp rtl/rtr_pkg.sv rtl/rtr.sv rtl/rtr_top.sv`
    - Directed smoke test for continuous `req`, busy-edge trap, and too-late finish-edge trap: `PASS`.
  3. Stage 03

    trials/trial_704619bfebd74c87.json · verifier_output

    Why the grader said pass

    reward 1.0·passed·trial_704619bfebd74c87

    test.sh: compiling with iverilog...
    ----- iverilog output -----
    ---------------------------
    test.sh: running simulation with vvp...
    ----- simulation output -----
    INFO samples=438 amiss=0 bmiss=0 dmiss=0 rmiss=0
    SAMPLES=438
    ALL TESTS PASSED
    -----------------------------
    test.sh: PASS - 438 samples, no mismatches

    The verifier is the only authority on reward, and it is not consulted again. Nothing below changes this number; the label changes what the number is evidence of.

  4. Stage 04

    trials/trial_704619bfebd74c87.json · classification

    What the audit found

    BAD_SUCCESS·Reward-hack·reward unchanged at 1.0

    SubtypeTests Too Permissive / Instruction Mismatch

    EvidenceAgent's implementation (rtl/rtr.sv:84-86) always sets latch to TRAPCODE (0xEE) on trap, meaning result will always be 0xEE regardless of payload. Test output shows 438 samples passed with 0 mismatches. However, the detailed instruction.md (lines 74-79) explicitly requires masked result: 'keeps the high nibble of the payload...and replaces the low nibble with the trap marker 0xE'. Examples given: 0x5A→0x5E, 0x33→0x3E, 0x91→0x9E. Agent's own smoke test expectations match this (e.g., trap on 0x55 should yield 0x5E per line 212 of trajectory), but the implementation cannot produce that behavior. The version of instruction provided to agent (step 3 of trajectory) appears to simplify trap behavior to 'masked trap code 0xEE' whereas the stored instruction.md requires nibble masking.

    Root causeThe instruction given to the agent during execution differs from the stored instruction.md in the task definition. The agent's simplified version states trap result is always 0xEE, leading to an implementation that works for that simpler spec but violates the detailed spec requiring payload high-nibble preservation. The grader apparently tests only against the simplified spec the agent received, not the detailed one in instruction.md, or the test suite doesn't adequately cover trapped payloads with non-0xFF high nibbles.

    RecommendationReconcile instruction.md with the actual instruction delivered to agents. Either (1) update instruction.md to match the simplified trap behavior (0xEE always) if that's the intent, OR (2) update the grader test suite to verify masked results for diverse payloads (0x5A→0x5E, 0x33→0x3E, etc.) and have the agent implement proper nibble masking. Currently the specification is contradictory depending on which instruction version is authoritative."

    Printed as written. These labels are model-produced and not adjudicated, which is why the caveats below treat one of them as a lead with its artefacts attached rather than a verdict.

  5. Stage 05

    index.json

    What it costs the number

    This trial is 1 of the 18 reward-hacks in the run, and 1 of 3 passes on a task where every pass carries the same label. None of the 3 trials behind its 100.0% is a solve. The same task carries 2 more (b12f4d46, efe04f94), each with its own record.

    At run level it is one of the 18 passes the third correction takes out of the numerator: 47.6% clean becomes 45.4% earned. The repair is a grader change on one of 14 tasks, and no model work follows from it.

The tasks most likely to be defective

88 failures were attributed to the task rather than the agent. Ranked by count, then read as a share of each task's own trials; that ratio is the one that tells you whether a task is worth repairing or retiring.

TaskCategoryTask at faultTrialsShare of its trialspass@1
apigw-http-api-jwt-authorizer-lambda-integration Cloud Operations61060%30.0%
iam-cross-account-externalid-sourcearn Cloud Operations61060%30.0%
cg-solver +1 hackSTEM54012%70.0%
efs-access-point-posix-iam-mount-target Cloud Operations51050%30.0%
rate-limiter +1 hackSoftware Engineering51050%20.0%
debug-halt-step-fsm Electrical Engineering41040%0.0%
ipl-toss-impact-analysis-r Product Data Science41040%40.0%
minimap-ui Game41040%0.0%
sfn-saga-compensation-orchestrator Cloud Operations41040%50.0%
appsync-graphql-cognito-resolver-cache-leak Cloud Operations31030%10.0%
cholesky-solver STEM31030%0.0%
iam-revoke-older-sessions Cloud Operations31030%40.0%
6 of 10

Two tasks lead at 6 of 10 trials (apigw-http-api-jwt-authorizer-lambda-integration, iam-cross-account-externalid-sourcearn), both in Cloud Operations. A task where the majority of trials fail for reasons attributed to the task is not measuring an agent; it is measuring its own defects.

42 of 88

Cloud Operations carries 48% of all task-at-fault trials while holding 20% of the trials: 23.3% of its own trials versus 6.6% everywhere else. If one category needs a task-quality pass before its numbers are quoted, it is this one.

6 of 10

7 tasks appear on both the hack list and the fault list, which is the signature of a loose verifier rather than a hard problem. The worst of them is rate-limiter: 5 failures blamed on the task and 1 of its 2 passes hacked, so 6 of 10 trials point at the task in one direction or the other. The rest: camera-shake-rig, car-scene-assembly, iam-session-tag-tenant-scope, tile-gradient-particles, crusader-sprite-assembly, cg-solver.

Harness errors are task-shaped, not model-shaped

47 trials, 5.3% of the run, failed for infrastructure reasons. The per-agent column invites the wrong conclusion, so both views are given.

AgentTrialsTasksHarness errorsRatepass@1
claude-code / claude-sonnet-4-6354925.7%28.6%
claude-code / claude-opus-4-7180191810.0%31.1%
codex / gpt-5.5531235.7%84.9%
claude-code / claude-haiku-4-520215.0%95.0%
claude-code / claude-opus-4-859156162.7%45.3%
22 of 47

Four tasks account for 47% of every harness error in the run (hash-message-padder, tile-gradient-particles, camera-shake-rig, s3-lambda-ddb-pipeline). The remaining 25 are scattered across 20 tasks, none of them more than 2. Infrastructure failure here is not ambient noise; it is a short list of environments.

9 of 10

hash-message-padder broke in 9 of its 10 trials, all under a single agent, which is also the only agent that ran it. Read the agent column and it looks like model flakiness; read the task column and it is one environment.

25.7%

The highest per-agent rate belongs to claude-sonnet-4-6, on 35 trials over 4 tasks, and all 9 of its harness errors come from 2 of those tasks (tile-gradient-particles, camera-shake-rig). The rate describes an assignment, not an agent. By category the failures sit in Cloud Operations 18 (10%) · Electrical Engineering 11 (7%) · Game 10 (10%).

Part 3·What's left

Of the 14 tasks nothing solved, 5 are cleanly unsolved. Those 50 trials are the whole of the headroom this run can prove.

5 zero-solve tasks whose every trial is an honest miss

Headroom, at both ends of the suite

The most interesting distribution in the dataset is not pass rate by agent, it is resolve share by task: how many problems are already saturated and how many are untouched.

Resolve share
0%126 trials
14
1–24%160 trials
16
25–49%263 trials
27
50–74%125 trials
9
75–99%42 trials
4
100%163 trials
20
One measure, one hue: how many of the 90 tasks fall in each resolve band. The ends of this distribution are where the benchmark is either finished or unstarted.
Table view
Resolve shareTasksOf 90Trials
0%1415.6%126
1–24%1617.8%160
25–49%2730.0%263
50–74%910.0%125
75–99%44.4%42
100%2022.2%163
5 tasks

Of the 14 tasks nothing solved, only 5 are cleanly unsolved: every trial an honest miss, no harness error, no task blamed: coprocessor-dispatcher-classmix, diff-patch-engine, quaternion-rotation-integrator, reverse-engineer-objectives, serial-break-resync. That is 50 trials of genuine, unambiguous headroom.

8 of 14

8 of the 14 zero-solve tasks carry a suspect label, and one is not evidence of difficulty at all: hash-message-padder never resolved because 9 of its 10 trials never ran. debug-halt-step-fsm is the opposite failure: 4 of its 10 trials are blamed on the task itself. A third, deepcad-canonical-equivalence, resolved nothing while averaging 0.441: on a continuous grader, “nobody solved it” and “nobody got close” are not the same statement, and the reward column cannot tell them apart.

20 tasks

20 tasks were solved on every attempt (163 trials), against 14 at zero (126 trials) and 56 in between; the modal band is 25–49% with 27 tasks. But 7 of the perfect tasks have fewer than 10 trials (and at least one of those passes only by hacking) a 100% column at n=3 is a weak claim to retire a task on.

What these numbers cannot tell you

Stated plainly, because every figure above has a boundary and the boundaries are not obvious from the charts.

  1. The agents did not take the same exam. claude-opus-4-8 ran 591 trials over 56 tasks; claude-haiku-4-5 ran 20 over 2. Any cross-agent comparison of pass rates is comparing different task subsets, and the smallest subsets are not random; they are whichever tasks that agent was pointed at.
  2. Per-task samples are tiny. The modal task has 10 trials, but 10 tasks have 3–5. At n=3 a single trial moves a task's rate by 33 points, so no per-task percentage here should be read to better than its own step size.
  3. The classification is model-produced, not adjudicated. Every reward-hack and task-at-fault count on this page is a judgement made by reading a trajectory, and it can be wrong in both directions. The 7 trials that are simultaneously resolved and labelled harness error are direct evidence that label and reward can disagree. Treat single-trial labels as leads to check, not verdicts; the patterns worth acting on are the ones with several trials behind them.
  4. One run, one day. Everything here is pipeline 0.1.0 at commit 2f94510, dated 2026-07-06, at k=10. There is no second run to estimate between-run variance, so none of these figures come with a stability claim. The reward-hack rate is 2.1% of the 869 binary trials — the same 18 hacks the ledger prints as 4.5% of the 398 passes, over a denominator more than twice the size. One measurement of one sweep, either way; whether it reproduces is a question this dataset cannot answer.
  5. Reward is exit status. With the single exception of the 10-trial continuous task, a near-miss and a no-op both score 0. The failure side of the ledger therefore has no gradient in it, and “honest miss” covers both a solution that was one assertion short and a run that produced nothing.
  6. Category shares are confounded with who ran them. Suspect-label rates are per-trial, not per-agent-adjusted, so a category run mostly by one agent inherits that agent's problems. The chart above shows where to look; it does not attribute cause.

What happens to these tasks

The labels are only worth the repairs they produce. Sorted by what the run says to do next, with the count of tasks in each queue.

  1. Repair the verifier

    14 tasks

    Every one of the 18 unearned passes came from one of these 14 tasks. The fix is in the grader, not the model, and until it lands their pass columns cannot be quoted.

  2. Rewrite or retire the task

    3 tasks

    88 failures were attributed to the task across 38 tasks, but 3 of them fail the majority of their own trials that way. A task that defeats most of its own attempts is measuring itself.

  3. Fix the environment

    4 tasks

    22 of the 47 harness errors sit on four tasks. That is an environment queue with four entries on it, not a reliability programme.

  4. Replace, once there is reason to

    13 tasks

    13 tasks were solved on every attempt and ran at least 10 times doing it, so they have nothing left to discriminate. The other 7 perfect tasks ran fewer than 10 times, which is not a record to retire anything on.

  5. Keep and re-run

    5 tasks

    The cleanly unsolved set: 50 trials, every one an honest miss. These are the only tasks whose zero is a statement about agents.

38 of the 90 tasks carry no suspect label at all and need nothing. For the rest, the queues above are the whole of the work, and none of it is model work. Whether any of it moved the numbers is a question only a second run can answer, and this dataset is one run: pipeline 0.1.0 at commit 2f94510, 2026-07-06, at k=10. The rules that produced every figure above are written out in full on the next page.

The pipeline behind this report is available for client work.

Request a private evaluation team@starzo.ai