Independent evaluations for AI coding agents.
Most evaluations tell you your coding agent passed.
We tell you whether
the pass was earned.
We build containerised tasks with graders the agent never sees, run it across them at scale, and check every pass against the trajectory that produced it.
879 graded trials· 90 tasks· 15 domains· every pass audited
Evidence: the full evaluation, published
We publish the whole run rather than a claim, and the awkward figures are printed beside the flattering ones.
What the audit changes
Every line's reported pass rate, next to the rate that is left once passes that gamed the grader are removed. Ranked lines only.
90 tasks across 15 domains
Grouped into 5 families. Counts and resolve rates are this run's own, and every domain links through to its tasks.
Software & cloud
30 tasks· 43% resolved
Hardware & mechanics
28 tasks· 46% resolved
Machine learning
11 tasks· 40% resolved
Data science & statistics
6 tasks· 35% resolved
Resolve rate is the share of a family's trials that returned a reward of exactly 1.0, across all 5 agent lines. One task is scored on a continuous 0–1 scale, so it counts toward the task totals but not the rates.
How we operate
The commitments a buyer is entitled to check before commissioning anything. Each one is visible in the report rather than asserted here.
The grader is never in the agent's reach
Verifier files sit outside the container surface the agent can read or write, so a trial cannot be passed by editing what marks it.
Scoring and labelling are separate systems
The verifier alone sets reward. The five-way classification that explains a trial is post-hoc and cannot change a score, so neither can contaminate the other.
Our own defects are counted against us
88 failures in this run are charged to our tasks rather than to the model under test — 18% of every failure, published per task.
Rates carry their uncertainty, or they are withheld
Every rate is published with a 95% Wilson interval, and a line with 20 trials or fewer is drawn but given no rank.
Get in touch
Two doors into STARZO AI, depending on which side of the evaluation you sit on.
I need an evaluation
Talk to our team.
A private run of your agent, tasks and environments built to your spec, or the data behind a report — one conversation scopes all of it.
Continue →I build evaluations
Build with us.
Domain experts and task authors who can write problems an agent cannot shortcut — the suite grows through people like you.
Continue →Each panel opens a message. Or write to us directly at team@starzo.ai.