Terminal-Bench 3 Frontier Set is a ten-task batch spanning five Terminal-Bench 3 categories — Security, Operations, Science, Software, and ML — each package hardened to be reference-solvable and bound to a delivered SHA-256 hash.
Every task is scored against GPT-5.6 Sol, the frontier reference, at pass@5 on the Terminus-2 scaffold, with Gemini 3.7 Flash and Hy3 (Preserve Thinking) run alongside as benchmarking arms at pass@5 and pass@10. Every task holds GPT-5.6 Sol at or below 40% full credit; the failure-mode analysis localises every miss to a named mechanism, and no reward-hacking behaviour was observed in the recorded executions.
- Hard but solvable. A package is included only when GPT-5.6 Sol, the frontier reference, earns full credit on at most 40% of pass@5 attempts; Sol clears 8 of 48 graded rollouts. Gemini 3.7 Flash and Hy3 (Preserve Thinking) are run alongside as benchmarking arms rather than admission criteria — Gemini clears 5 of 48 at pass@5 and Hy3 clears 4 of 100 at pass@10. Every packaged reference solution was then executed under its own verifier at the delivered hash and earned full credit, and a no-op candidate earned zero on every package. Ungraded slots count as zeros.
- Long-running rollouts. Median step counts vary widely by task: GPT-5.6 Sol’s per-task median ranges from 15 to 58 decisions (overall median 31), and Gemini 3.7 Flash’s ranges from 29 to 117 (overall median 69) — three tasks push Gemini past 100 decisions before a graded submission.
- Evidence bound to the hash. Every number across all 196 rollouts (48 Sol, 48 Gemini, 100 Hy3) is bound to the delivered package SHA-256, and each of the 48 GPT-5.6 Sol trials additionally carries a per-trial failure classification against its trajectory and verifier output.
Overview and category distribution
A hard, solvable 10-task set spanning five TB3 categories, every package reference-solvable. Security is the largest category at three tasks (30%).
Category distribution
Five categories across the ten tasks. Security leads with three tasks; the other four categories carry one or two each.
Difficulty and long horizon
Difficulty gate (GPT-5.6 Sol, the frontier reference):
- GPT-5.6 Sol full credit ≤ 40% at pass@5
Gemini 3.7 Flash and Hy3 (Preserve Thinking) are run alongside as benchmarking arms and are not part of the admission gate; their pass rates are reported for comparison throughout.
Long-running rollouts: GPT-5.6 Sol’s per-task median step count ranges from 15 to 58 decisions (overall median 31); Gemini 3.7 Flash’s ranges from 29 to 117 (overall median 69), with three tasks pushing Gemini past 100 decisions before a graded submission.
Every task in this set is solvable. Solvability is verified by executing the packaged reference solution under the packaged verifier at the delivered hash. Every package scored full credit and a no-op candidate scored zero.
Six of ten are never passed by GPT-5.6 Sol, and two — regalloc-spill-repair-v3-3RbQx and ntfs-timeline-submit — are passed by none of the three arms across all 39 of their graded rollouts. auth-sig-reconcile was passed by neither Sol nor Gemini but is cleared by Hy3 on 3 of its 10 pass@10 rollouts. Ungraded slots count as zeros.
Hy3 (Preserve Thinking), run at pass@10 on the same ten packages as a second benchmarking arm, is carried alongside Sol and Gemini in the difficulty and long-horizon views below; it clears 4 of 100 graded rollouts overall.
Difficulty: full-credit rate for GPT-5.6 Sol
Full-credit rate per task for GPT-5.6 Sol, the frontier reference, at pass@5, with Gemini 3.7 Flash and Hy3 (Preserve Thinking) charted alongside as benchmarking arms at pass@5 and pass@10 respectively. minic-bounds-checker only had 3 GPT-5.6 Sol rollouts because the task was flagged as a cybersecurity risk by the model provider, so 2 of 5 rollouts were cut off. ntfs-timeline-submit and stake-resets each have 4 graded Gemini 3.7 Flash rollouts. Each arm is measured at its own denominator and ungraded slots count as zeros.
Long horizon: median number of decisions per task
Median decisions per graded rollout, by arm and task. GPT-5.6 Sol’s overall median is 31 decisions, Gemini 3.7 Flash’s is 69, and Hy3’s is 40.
Failure mode analysis
This analysis covers all 48 graded GPT-5.6 Sol trials: 8 earned full credit and 40 did not. Each failure is classified from the agent trajectory and the verifier reward. Gemini 3.7 Flash and Hy3 (Preserve Thinking) are reported as benchmarking arms in the chart above; the per-trial failure classification below stays on the Sol arm.
Validity labels:
- GENUINE_DIFFICULTY
- INSTRUCTION_TEST_MISMATCH
- ENVIRONMENT_INSTABILITY
- MIXED
- INCONCLUSIVE
Failure pattern distribution
The 40 failing GPT-5.6 Sol trials are grouped by the earliest mechanism that blocked full credit. Mechanisms are read from the agent trajectory; the grade comes from the packaged verifier’s reward. The nine terminal-control breakdowns fall on two tasks and are reported separately rather than folded into the semantic misses.
| Task | GPT-5.6 Sol | Dominant mechanism | What the verifier caught | Secondary modes | Validity | Package integrity & closest trial |
|---|---|---|---|---|---|---|
| 01 · ballast-exchange-v4-schema-fi-W5r3QBallast-exchange compliance analysis | 0/5 full creditmedian 58 decisionsGemini 1/5, median 111 | terminal-control-breakdown · 4/5Each run wedged the pane on a heredoc or bracketed paste and spent the remaining budget on interrupt and paste-terminator sequences. The starter remained unchanged. |
|
| MIXED 5 · scaffold-bound on the Sol arm | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-01 completed the broadest environment inspection, but no slot began implementation. Gemini 3.7 Flash passed slot-05. |
| 02 · parallelism-strategy-v5-ywY87Distributed-training parallelism strategy | 0/5 full creditmedian 31 decisionsGemini 1/5, median 55 | terminal-control-breakdown · 5/5Every run lost the shell to a stuck foreground process and cycled through tmux prefixes, raw ETX bytes and C-c variants. No strategy artefact was written. |
| — | MIXED 5 · scaffold-bound on the Sol arm | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-05 ran furthest at 55 decisions without recovering the shell. Gemini 3.7 Flash passed slot-05. |
| 03 · polyploid-remapPolyploid genome coordinate remapping | 0/5 full creditmedian 32 decisionsGemini 2/5, median 68 | declared-complete-graded-wrong · 5/5Every run compiled the repaired sources, exercised both required invocation modes, ran its own digest and determinism checks, and submitted output that mismatched the withheld remapping fixtures. |
| — | GENUINE_DIFFICULTY 5 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-01 produced the most complete output at 37 decisions. Gemini 3.7 Flash passed slot-02 and slot-04. |
| 04 · regalloc-spill-repair-v3-3RbQxRegister-allocator spill/reload repair | 0/5 full creditmedian 44 decisionsGemini 0/5, median 117 | budget-exhausted-before-deliverables · 4/5Four runs were still deriving spill placement and reload ordering against the normative worked examples when the 10,800-second deadline arrived. |
|
| GENUINE_DIFFICULTY 5 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-05 completed and submitted a fully audited repair at 64 decisions. No arm passed; the packaged reference solution earns full credit at the delivered hash. |
| 05 · minic-bounds-checkerStatic array-bounds checker for a C subset | 2/3 full creditmedian 31 decisionsGemini 0/5, median 89 | budget-exhausted-before-deliverables · 1/1The single graded miss was still resolving i128-versus-mathematical-integer soundness on high-arity paths at the deadline. |
| — | INCONCLUSIVE 1 · 3 of 5 slots graded | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-03 and slot-05 earned full credit; slot-01 ran to 40 decisions without submitting. |
| 06 · auth-sig-reconcileAUTHMINI firmware signature reconciliation | 0/5 full creditmedian 16 decisionsGemini 0/5, median 51 | declared-complete-graded-wrong · 5/5Every run built the reconciler in Release mode, emitted all 29 manifests, verified them against the published schema and submitted. Grading is byte-exact against a withheld oracle. |
| — | GENUINE_DIFFICULTY 5 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-02 ran furthest at 31 decisions and produced all 29 manifests. The packaged reference solution earns full credit at the delivered hash. |
| 07 · ntfs-timeline-submitNTFS journal timeline compilation | 0/5 full creditmedian 23 decisionsGemini 0/4, median 43 | declared-complete-graded-wrong · 5/5Every run built through the official entrypoint, produced all 16 timeline documents and validated them against the output schema before submitting. |
| — | GENUINE_DIFFICULTY 5 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-02 ran furthest at 34 decisions and produced all 16 timelines. The packaged reference solution earns full credit at the delivered hash. |
| 08 · crew-deadheadCrew-pairing deadhead insertion optimizer | 2/5 full creditmedian 15 decisionsGemini 0/5, median 29 | declared-complete-graded-wrong · 3/3Each miss passed the bundled 13-stage selfcheck and its own randomized brute-force cover comparison, then diverged from the sealed schedules. |
| — | GENUINE_DIFFICULTY 3 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-03 and slot-04 earned full credit on the same sealed schedules the three misses failed. |
| 09 · stake-resetsMass balance across undocumented stake resets | 2/5 full creditmedian 33 decisionsGemini 0/4, median 113 | declared-complete-graded-wrong · 3/3Each miss compiled the Rust binary, produced a schema-valid balance.json and validated it against the published contract before submitting. |
| — | GENUINE_DIFFICULTY 3 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-01 and slot-04 earned full credit on the same held-out sites the three misses failed. |
| 10 · hls-playlist-reconcilerHLS live playlist reconciliation and audit | 2/5 full creditmedian 15 decisionsGemini 1/5, median 70 | declared-complete-graded-wrong · 3/3Each miss produced the public timeline and audit artefacts and passed its own CLI, canonicalisation, ordering and bounds checks before submitting. |
| — | GENUINE_DIFFICULTY 3 · no instruction/test mismatch | No package or verifier defect observed; all runs used the same task checksum.Reward hacking · No protected-path or reward-channel access observed.Closest Sol trial · slot-02 and slot-03 earned full credit; Gemini 3.7 Flash also passed slot-03. |
Quality assurance and acceptance
Every package was reviewed on five dimensions by human and LLM judges.
- instruction quality
- instruction-test alignment
- test and verifier quality
- environment reproducibility
- answer dependency
Cross-cutting findings
Misses are semantic, not budgetary. Twenty-five of the 40 failing Sol trials ended in a submission the run had validated itself: the deliverables were present, well-formed and schema-valid, and the sealed cases rejected them. On polyploid-remap, auth-sig-reconcile and ntfs-timeline-submit that accounts for every miss, five of five. The misses are wrong-but-complete outputs, not missing deliverables.
Arm coverage is complementary. Neither arm dominates: Sol leads on four tasks, Gemini on three, and three are ties at zero. Gemini spends far more decisions per attempt — a median of 69 against Sol’s 31 — and still lands lower overall, at 5 of 48 against 8 of 48. Gemini runs longer on every one of the ten tasks.
Two tasks fail on terminal control, not on the problem. All nine terminal-control breakdowns fall on ballast-exchange-v4-schema-fi-W5r3Q and parallelism-strategy-v5-ywY87, where the Sol runs wedged the pane on a heredoc or a bracketed paste and spent the remaining budget sending interrupts. Both are labelled MIXED rather than GENUINE_DIFFICULTY, and both are passed on the Gemini arm at 1/5, which establishes solvability independently of the Sol zeros.
Under-specified acceptance rules are the recurring open finding. Grading in this set is byte-exact and all-or-nothing, and several packages turn on a rule the specification states imprecisely. crew-deadhead gives the deadhead preference chain without a direction for any key, and that direction decides 30 of 32 sealed schedules. hls-playlist-reconciler fixes the diagnostic vocabulary but not the subject each reconciliation-phase code carries. auth-sig-reconcile requires a counter-signature chain ordering that appears in no agent-visible document. Each rejects correct work rather than granting unearned credit.
No task-package defect was established. Every package’s reference solution earns full credit under its own verifier at the delivered hash and every no-op earns zero, so no miss in this cohort is attributable to an unpassable package. No failure was classified as an instruction/test mismatch.
Both zero-across-all-arms tasks were independently re-verified and are genuinely hard, not broken. regalloc-spill-repair-v3-3RbQx (0/5 Sol, 0/5 Gemini, 0/10 Hy3) and ntfs-timeline-submit (0/5 Sol, 0/4 Gemini, 0/10 Hy3) were each targeted for a follow-up check: the reference solution was rebuilt from source and re-run against the full hidden grading set independently of the packaged verifier, reproducing an exact match on every case (regalloc’s 17 witnesses and 48 seeds; ntfs’s 16 volumes), while the shipped no-op scaffold independently failed as expected. Both verifiers grade byte-exact and all-or-nothing with no reward-hacking surface. Every one of the 39 failing rollouts across all three arms engaged with the real task — building through the official entrypoint, reading the real spec and fixtures, and submitting a plausible, self-validated answer that the sealed grading still rejected on at least one of dozens of independently-planted or independently-generated cases; regalloc-spill-repair-v3-3RbQx’s own instruction discloses that most of its planted defects cannot be locally confirmed before submission. The one exception is two of Gemini’s four ntfs-timeline-submit runs, which fabricated volume identifiers absent from the actual input and never grounded in the real fixtures — a model-side execution failure, not a task defect. Two latent documentation gaps were found in ntfs-timeline-submit’s spec (an unstated RENAME name-only match fallback and an unstated UNLINK not-found behavior, both present in the reference implementation but not in the prose) and confirmed to be exercised by none of the 16 graded fixtures, so neither caused any of the observed failures.
Hy3 trails GPT-5.6 Sol and Gemini 3.7 Flash alike, and breaks one task neither of them clears. Run at pass@10 on the same ten packages, Hy3 (Preserve Thinking) clears 4 of 100 graded rollouts overall — a lower full-credit rate than either Sol (8/48) or Gemini (5/48) even accounting for the larger sample. Its one substantive result is auth-sig-reconcile, which neither Sol nor Gemini ever passes: Hy3 clears it on 3 of 10 rollouts, and the passing runs share a specific trait — they read the stored signer-record digest bytes rather than recomputing them from the documented AUTHHASH formula, the same undocumented ordering gap recorded for this task above. Hy3 also reaches 1 of 10 on crew-deadhead, which Sol already clears at 2/5. On the remaining eight tasks Hy3 scores zero, including the two — regalloc-spill-repair-v3-3RbQx and ntfs-timeline-submit — still unpassed by any of the three arms. No reward-hacking behaviour was observed in the recorded Hy3 executions either: command scans across the failing trials on all ten tasks found no access to verifier-private paths, packaged solutions, or reward channels.
No reward-hacking behaviour was observed. A scan of the 96 recorded executions found no access to verifier-private paths, grader sources, reward channels, or packaged solutions. Five packages were additionally probed with candidates that forge every reachable reward and result channel, enumerate solution and test paths from the agent container, plant conftest.py and sitecustomize.py shadows, and plant an /app/pytest.py that hijacks the verifier’s own test collection; all scored zero. This is evidence about the recorded executions and the probes actually run, not a universal exploit-resistance claim.