01Agent trajectories
Repository-level software work with executable verification
End-to-end traces across real codebases: planning, tool calls, multi-file edits, terminal execution, and tests. Built for the class of long-horizon work measured by SWE-bench and Terminal-Bench.
repo / payments-serviceenvironment live
01InspectCall graph mapped
02ReproduceFailing test added
03Patch3 files changed
04VerifyAll checks pass
02Model stumping
Expose the reasoning failures hidden by general benchmarks
Domain experts design adversarial tasks, evaluation rubrics, and failure analyses for consequential workflows in finance, healthcare, research, and operations.
Surface fluencyWorkflow depth
Failure modes surfaced
05Post-training
Expert preference data that captures why one response is better
SFT examples, preference pairs, critiques, and rubric traces for coding, STEM, multilingual, and professional reasoning.
A
Expert critique →
B · chosen
06Domain data
Ready-made and custom packs grounded in professional work
Finance, healthcare, law, technical, and multilingual datasets spanning reasoning, retrieval, evaluation, and multimodal tasks.
FinanceHealthcareLawTechnicalMultilingual