SyncValsfield notes

Task index · run 2026-07-06 · 90 problems

90 problems, each with a verifier the agent never sees.

A task is one self-contained problem: a starter workspace the agent opens cold, a reference solution held outside its image, and a grading program that decides reward on its own. The agent gets the workspace and the brief. It does not get the solution, the tests, or a second opinion. Everything below is per task — how many times it was attempted, how often a run came back at exactly 1.0, and which of the five outcomes it produced most.

90
tasks in the corpus
879 scored trials across them
15
categories
10 live, 5 preview
14
tasks nobody solved
no agent reached reward 1.0 once

What a task is made of

Four pieces, and the split between them is the whole design: the agent can read two of them and nothing it writes can reach the other two.

starter workspace

The filesystem the agent wakes up in — sources, fixtures, build files, and nothing that answers the question for it.

hidden reference solution

Every task ships with a worked solution kept out of the agent's image. It proves the problem is solvable and fixes what counts as done.

deterministic verifier

Grading is a program, not an opinion. The same submission returns the same reward on every run, and the agent never reads it.

k = 10 trials per task

Each agent attempts a task repeatedly, so one lucky pass cannot carry a task or a category.

pass@1 below is the share of a task's trials that returned a reward of exactly 1.0. 1 of the 90 tasks is scored on a continuous 0–1 scale instead, and is shown with a mean score — a pass@1 figure would be meaningless for it.

Fifteen categories

Bars are resolve rate — one hue, one measure. A category reads live once it has 20 or more scored trials and preview below that; the threshold counts trials, not tasks, so three tasks run five times each still read preview.

Software Engineeringlive

Algorithms, data structures, bug-fixes, and API/systems implementation, graded against hidden test suites the agent never sees.

8tasks
85trials
49/85resolved
57.6% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
session-token-verify1010/10100.0%Legitimate solvetask page ↗
window-aggregate-store1010/10100.0%Legitimate solvetask page ↗
lru-cache1514/1593.3%Legitimate solvetask page ↗
occ-conditional-store109/1090.0%Legitimate solvetask page ↗
idempotency-middleware104/1040.0%Honest misstask page ↗
rate-limiter102/1020.0%Task at faulttask page ↗
diff-patch-engine100/100.0%Honest misstask page ↗
resilient-http-client100/100.0%Honest misstask page ↗

Mechanical Engineeringlive

Structural and numerical solvers (FD/FE, contact dynamics) in C++, checked against multi-binary hidden references.

8tasks
80trials
28/80resolved
35.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
rk4-orbit-integrator1010/10100.0%Legitimate solvetask page ↗
truss2d-solver106/1060.0%Legitimate solvetask page ↗
pipeflow-colebrook-solver104/1040.0%Honest misstask page ↗
beam-deflection-solver103/1030.0%Honest misstask page ↗
heat1d-conduction-solver103/1030.0%Honest misstask page ↗
projectile-drag-integrator102/1020.0%Honest misstask page ↗
collision2d-impulse-solver100/100.0%Honest misstask page ↗
quaternion-rotation-integrator100/100.0%Honest misstask page ↗

Cloud Operationslive

Provision, diagnose, and guardrail cloud infrastructure (AWS), verified against the deployed state.

19tasks
180trials
56/180resolved
31.1% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
ecs-fargate-secrets-kms-exec-role76/785.7%Legitimate solvetask page ↗
secrets-rotation-kms106/1060.0%Legitimate solvetask page ↗
sfn-saga-compensation-orchestrator105/1050.0%Legitimate solvetask page ↗
athena-workgroup-result-encryption-cmk-enforced52/540.0%Honest misstask page ↗
ddb-outbox-eventbridge-fanout104/1040.0%Honest misstask page ↗
iam-revoke-older-sessions104/1040.0%Legitimate solvetask page ↗
ecr-image-scan-lifecycle-immutable-tags-replication93/933.3%Honest misstask page ↗
glue-etl-catalog-security-configuration-kms93/933.3%Honest misstask page ↗
apigw-http-api-jwt-authorizer-lambda-integration103/1030.0%Task at faulttask page ↗
efs-access-point-posix-iam-mount-target103/1030.0%Task at faulttask page ↗
iam-cross-account-externalid-sourcearn103/1030.0%Task at faulttask page ↗
iam-session-tag-tenant-scope103/1030.0%Honest misstask page ↗
apigw-sqs-fifo-direct-integration102/1020.0%Honest misstask page ↗
iam-permissions-boundary-ceiling102/1020.0%Honest misstask page ↗
s3-lambda-ddb-pipeline102/1020.0%Harness errortask page ↗
s3-sqs-image-pipeline-kms102/1020.0%Honest misstask page ↗
appsync-graphql-cognito-resolver-cache-leak101/1010.0%Honest misstask page ↗
cognito-m2m-httpapi-jwt-scope-gated101/1010.0%Honest misstask page ↗
sfn-secrets-rotation-chain101/1010.0%Honest misstask page ↗

Electrical Engineeringlive

RTL / digital-logic design checked under hardened simulation and formal-equivalence harnesses.

20tasks
158trials
82/158resolved
51.9% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
apb-lite-slave-regfile1010/10100.0%Legitimate solvetask page ↗
bcd-tens-rollover1010/10100.0%Legitimate solvetask page ↗
bus-slave-abort-ack1010/10100.0%Legitimate solvetask page ↗
Enable-gated streaming fold stage33/3100.0%Legitimate solvetask page ↗
Instruction-retire commit handshake33/3100.0%Reward-hacktask page ↗
page-program-suspend1010/10100.0%Legitimate solvetask page ↗
round-robin-arbiter-grant1010/10100.0%Legitimate solvetask page ↗
Serial bit-destuff framer33/3100.0%Legitimate solvetask page ↗
Wait-state register-file completer33/3100.0%Legitimate solvetask page ↗
open-drain-command-engine108/1080.0%Legitimate solvetask page ↗
dualmaster-membridge105/1050.0%Legitimate solvetask page ↗
stream-frame-hold104/1040.0%Honest misstask page ↗
byte-serial-round-scheduler102/1020.0%Honest misstask page ↗
serial-receiver-framed101/1010.0%Honest misstask page ↗
coprocessor-dispatcher-classmix100/100.0%Honest misstask page ↗
debug-halt-step-fsm100/100.0%Honest misstask page ↗
hash-message-padder100/100.0%Harness errortask page ↗
Multi-cycle signed divider with a start/valid handshake30/30.0%Honest misstask page ↗
Resynchronising serial byte receiver30/30.0%Task at faulttask page ↗
serial-break-resync100/100.0%Honest misstask page ↗

STEMlive

Quantitative reasoning across math, physics, and the natural sciences with deterministic, checkable answers.

5tasks
81trials
59/81resolved
72.8% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
adaptive-quadrature1010/10100.0%Legitimate solvetask page ↗
anova-stats1111/11100.0%Legitimate solvetask page ↗
cubic-spline1010/10100.0%Legitimate solvetask page ↗
cg-solver4028/4070.0%Legitimate solvetask page ↗
cholesky-solver100/100.0%Honest misstask page ↗

Gamelive

Interactive simulation and game-logic tasks graded on exact state transitions and rule fidelity.

10tasks
105trials
46/105resolved
43.8% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
game-of-life-step1515/15100.0%Legitimate solvetask page ↗
car-scene-assembly106/1060.0%Legitimate solvetask page ↗
combo-score-system106/1060.0%Legitimate solvetask page ↗
camera-shake-rig104/1040.0%Harness errortask page ↗
checkpoint-system104/1040.0%Honest misstask page ↗
crusader-sprite-assembly104/1040.0%Honest misstask page ↗
minimap-marker-logic104/1040.0%Honest misstask page ↗
tile-gradient-particles102/1020.0%Harness errortask page ↗
day-night-cycle-controller101/1010.0%Honest misstask page ↗
minimap-ui100/100.0%Honest misstask page ↗

Debuggingpreview

Real-world bug-fix tasks distilled from open-source issues (esbuild, klauspost/compress, rust-lang/semver), graded against each project's own tests.

3tasks
15trials
15/15resolved
100.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
evanw-esbuild-441755/5100.0%Legitimate solvetask page ↗
klauspost-compress-111555/5100.0%Legitimate solvetask page ↗
rust-lang-semver-30555/5100.0%Legitimate solvetask page ↗

ML Engineeringpreview

Train and tune models to a target metric on real engineering datasets, graded on held-out performance against a solved-reward threshold.

1task
10trials
4/10resolved
40.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
airfoil-self-noise104/1040.0%Honest misstask page ↗

Post Traininglive

Long-horizon, from-scratch numpy ML research: implement reverse-mode autodiff and a training method (quantization-aware, adversarial-robust, semi-supervised, worst-group), then train to a sealed held-out metric. Tasks published blinded.

6tasks
65trials
28/65resolved
43.1% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
fewlabel-ssl-fixmatch1010/10100.0%Legitimate solvetask page ↗
qat-int2-cifar159/1560.0%Legitimate solvetask page ↗
worst-group-spurious-dfr105/1050.0%Legitimate solvetask page ↗
adv-robust-pgd103/1030.0%Honest misstask page ↗
reverse-engineer-decoding101/1010.0%Honest misstask page ↗
reverse-engineer-objectives100/100.0%Honest misstask page ↗

Scientific MLlive

Physics-informed and surrogate modeling: PDE forecasting and CFD/FEA prediction, graded on quantitative predictive accuracy.

3tasks
30trials
10/30resolved
33.3% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
airfrans-high-reynolds-drag-extrapolation104/1040.0%Honest misstask page ↗
ks-equation-1d-forecast104/1040.0%Honest misstask page ↗
simjeb-bracket-fea-mass-prediction-real102/1020.0%Honest misstask page ↗

ML for Engineeringpreview

Machine-learning-for-engineering tasks (DeepCAD command-history canonicalization and equivalence), graded on a continuous partial-credit score against verifier-owned labels rather than all-or-nothing pass/fail.

1task
10trials
0.441mean score
scored 0–1, so no resolve rate applies
TaskRunsResolvedpass@1Dominant outcome
deepcad-canonical-equivalencegraded 0–110n/a0.441mean scoreHonest misstask page ↗

1 task here is graded on a continuous 0–1 score rather than pass/fail, so it has a mean score instead of a pass@1 — and no trial can register as resolved, which requires a reward of exactly 1.0.

Data Sciencelive

Modeling and statistical inference on real datasets (federated learning, event studies), graded against held-out ground truth.

2tasks
20trials
5/20resolved
25.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
product-recall-stock-price-event103/1030.0%Honest misstask page ↗
fedavg-federated-noniid-mnist102/1020.0%Honest misstask page ↗

Data Science: Robustnesslive

Bias-correction and outlier-robust analysis in R: recover trustworthy estimates from messy data, checked against reference results.

2tasks
20trials
8/20resolved
40.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
coffee-ratings-outliers104/1040.0%Honest misstask page ↗
lending-club-lgd-bias-correction-r104/1040.0%Honest misstask page ↗

Product Data Sciencepreview

Applied product analytics: causal impact and decision analysis on real product datasets (R).

1task
10trials
4/10resolved
40.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
ipl-toss-impact-analysis-r104/1040.0%Legitimate solvetask page ↗

Pharmacometricspreview

Population PK/PD modeling with nonlinear mixed-effects: fit drug-exposure models and recover the correct parameters.

1task
10trials
4/10resolved
40.0% of trials reached reward 1.0
TaskRunsResolvedpass@1Dominant outcome
neonatal-drug-exposure-nlme104/1040.0%Legitimate solvetask page ↗