Data, built as a course of education.

Production-grade data across categories — from public benchmarks to reasoning, agents, software engineering, embodied interaction, and vertical industries. Buy off the shelf or commission to spec. Pick a category to see how we build it and exactly what ships.

Research-Grade Scientific Reasoning

Expert-authored STEM problems with verifiable answers, full derivations, and rigorous quality control.

Terminal & System Operations

Hard terminal tasks with frozen environments, criterion rubrics, and deterministic offline verifiers.

SWE Engineering Tasks

Multi-language tasks from real GitHub PRs with runnable repos, gold patches, and pass/fail test verification.

Competition & Olympiad Problems

IMO/CMO and STEM olympiad problems with full solutions, final answers, and per-model run records.

Vertical Industry Data

Expert-annotated legal, finance, healthcare, and manufacturing tasks with grounded worlds and scored rubrics.

Medical & Clinical Reasoning

Physician-authored clinical and epidemiology problems with defensible answer keys and multi-model scoring.

GUI Grounding & Interaction

Screenshot-grounded interface data — localized elements, verified action trajectories, and checked screen state — across web, desktop, and mobile.

Agent Sessions & Tool Trajectories

Multi-tool trajectories from real dev sessions — tool calls, environment state, failure recovery, long horizons.

World-Model Simulation Data

Engine-simulated multi-view video with exact camera parameters and complete native annotation, for spatially consistent world-model training.

Off-the-Shelf Benchmark Packages

Ready-to-buy, standardized packages aligned to public benchmarks — procured, validated, and reproducible.

Embodied & Multimodal Interaction

Perception–planning–action data with environment feedback, image/video understanding, and spatial tasks.

Don’t see your category? We also build to spec.

Get in Touch