DeepSWE Bench
Ten long-horizon coding tasks mined from live private production codebases, authored via an LLM + human-in-the-loop pipeline and calibrated on a fixed harness — the reference solution scores 1.0, the empty submission 0.0.
Seven environments across three domains — three scored RL gyms, three 3D environments (agentic Blender authoring, a 616 GB industrial-CAD pack, and video→3D room layout), and a verified robot failure-and-recovery engine. Every one built to the same standard: real tasks, machine verification, and full evidence.
Deccan builds verifiable environments and data across RL, 3D and Robotics — the three areas you're building in. This package brings our work in all three into one place: each a self-contained delivery with its own deep-dive report, every sample leading with the numbers.
Three verifiable-RL gyms with scored rollouts, verifier evidence and per-failure analysis.
Ten long-horizon coding tasks mined from live private production codebases, authored via an LLM + human-in-the-loop pipeline and calibrated on a fixed harness — the reference solution scores 1.0, the empty submission 0.0.
Ten cross-gym agent tasks across five simulated SaaS products — XSN, Xalendly, Xubspot, Xmail and Xira — each scored by atomic database-expectation verifiers after live browser execution.
Ten hard, reference-solvable tasks across five Terminal-Bench 3 categories — Security, Operations, Science, Software and ML — scored against GPT-5.6 Sol at pass@5, with Gemini 3.7 Flash and Hy3 run alongside as benchmarking arms.
Verifiable 3D authoring, and 3D at industrial scale.
A verifiable 3D-authoring environment: an agent writes Python (bpy) to build 3D scenes, and every episode is graded by an end-state verifier that re-opens the saved scene in a fresh Blender and renders it from the harness, not the agent.
A 6.8 TB structural-steel fabrication dataset — coordinated 3D models (STEP, IFC, Tekla), 2D shop drawings and CNC programs from live US fabrication work, all linked by piece mark: 3D model → drawing → machine code.

A verified failure-and-recovery data engine, with live demonstrations.
An engine that manufactures realistic robot-failure trajectories and provably-working recoveries in rich simulated scenes — then proves the data transfers. A standing supplier of on-demand failure-recovery data in LeRobot format, every recovery machine-verified to complete.
An early probe into video → 3D floor-plan reconstruction — now scored against real ground truth by an LLM judge.
Given an egocentric video walkthrough, the agent must reconstruct a globally consistent 2D spatial layout. Because ground-truth layouts exist, you get deterministic rewards for topology, geometry and connectivity, which enables scalable RL and evaluation. Agents that succeed demonstrate long-horizon spatial reasoning and memory — key to embodied AI and robotics.
Across RL, 3D and Robotics we build and operate these environments and datasets at scale — with the same standard of verifiable grading, full evidence, and per-failure analysis you see in each report here.
We'd be glad to talk through scaled operations, custom environment and dataset builds, and OTS access for any of these lines — including custom robot embodiments, tasks and failure taxonomies through the FailRec engine.