Data, built as a course of education.
Production-grade data across categories — from public benchmarks to reasoning, agents, software engineering, embodied interaction, and vertical industries. Buy off the shelf or commission to spec. Pick a category to see how we build it and exactly what ships.
Research-Grade Scientific Reasoning
Expert-authored STEM problems with verifiable answers, full derivations, and rigorous quality control.
02Terminal & System Operations
Hard terminal tasks with frozen environments, criterion rubrics, and deterministic offline verifiers.
03SWE Engineering Tasks
Multi-language tasks from real GitHub PRs with runnable repos, gold patches, and pass/fail test verification.
04Competition & Olympiad Problems
IMO/CMO and STEM olympiad problems with full solutions, final answers, and per-model run records.
05Vertical Industry Data
Expert-annotated legal, finance, healthcare, and manufacturing tasks with grounded worlds and scored rubrics.
06Medical & Clinical Reasoning
Physician-authored clinical and epidemiology problems with defensible answer keys and multi-model scoring.
07GUI Grounding & Interaction
Screenshot-grounded interface data — localized elements, verified action trajectories, and checked screen state — across web, desktop, and mobile.
08Agent Sessions & Tool Trajectories
Multi-tool trajectories from real dev sessions — tool calls, environment state, failure recovery, long horizons.
09World-Model Simulation Data
Engine-simulated multi-view video with exact camera parameters and complete native annotation, for spatially consistent world-model training.
10Off-the-Shelf Benchmark Packages
Ready-to-buy, standardized packages aligned to public benchmarks — procured, validated, and reproducible.
11Embodied & Multimodal Interaction
Perception–planning–action data with environment feedback, image/video understanding, and spatial tasks.
Don’t see your category? We also build to spec.
Get in Touch