STARZO AIEvals

Independent evaluations for AI coding agents.

Most evaluations tell you your coding agent passed.
We tell you whether the pass was earned.

We build containerised tasks with graders the agent never sees, run it across them at scale, and check every pass against the trajectory that produced it.

879 graded trials· 90 tasks· 15 domains· every pass audited

4.5%
of passes were not earned
Every pass is read back against the trajectory that produced it.
398
passes audited
5
agent lines
481
failures explained

Evidence: the full evaluation, published

We publish the whole run rather than a claim, and the awkward figures are printed beside the flattering ones.

18
Reward-hack audit trail
Passes that gamed the grader, hardcoded an answer, or reached something they should not have. Each one is named.
88
Task-defect honesty
Failures charged to our tasks rather than to your model — 18% of every failure in the run, listed per task.
10×
Statistical discipline
Trials per task, with Wilson intervals on every rate and coverage read off the data rather than asserted.
0
Verifier isolation
Grader files the agent can read. The classifier that labels a trial never touches its score, so neither can contaminate the other.

What the audit changes

Every line's reported pass rate, next to the rate that is left once passes that gamed the grader are removed. Ranked lines only.

ModelReported pass@1After the auditDifference
gpt-5.5codex · 53 trials
84.9%
69.8%
−15.1
claude-opus-4-8claude-code · 591 trials
45.3%
44.2%
−1.1
claude-opus-4-7claude-code · 180 trials
31.1%
28.3%
−2.8
claude-sonnet-4-6claude-code · 35 trials
28.6%
17.1%
−11.5
Read the report →

90 tasks across 15 domains

Grouped into 5 families. Counts and resolve rates are this run's own, and every domain links through to its tasks.

Software & cloud

30 tasks· 43% resolved

Hardware & mechanics

28 tasks· 46% resolved

Reasoning & simulation

15 tasks· 56% resolved

Resolve rate is the share of a family's trials that returned a reward of exactly 1.0, across all 5 agent lines. One task is scored on a continuous 0–1 scale, so it counts toward the task totals but not the rates.

How we operate

The commitments a buyer is entitled to check before commissioning anything. Each one is visible in the report rather than asserted here.

  1. The grader is never in the agent's reach

    Verifier files sit outside the container surface the agent can read or write, so a trial cannot be passed by editing what marks it.

    How grading is isolated →

  2. Scoring and labelling are separate systems

    The verifier alone sets reward. The five-way classification that explains a trial is post-hoc and cannot change a score, so neither can contaminate the other.

    The five labels →

  3. Our own defects are counted against us

    88 failures in this run are charged to our tasks rather than to the model under test — 18% of every failure, published per task.

    Task-defect attribution →

  4. Rates carry their uncertainty, or they are withheld

    Every rate is published with a 95% Wilson interval, and a line with 20 trials or fewer is drawn but given no rank.

    How the intervals are computed →

Get in touch

Two doors into STARZO AI, depending on which side of the evaluation you sit on.

Each panel opens a message. Or write to us directly at team@starzo.ai.