Data & Environment

Data and environments for work that can be verified.

We build tasks whose outcomes can be checked by a reference answer, an executable test suite, a graded work product, or held-out re-execution. The library spans Frontier STEM, coding, professional work, and recursive self-improvement.

Expert-defined objectives Runnable task environments Reproducible evaluation
The library

Four forms of verifiable work.

The unit of evaluation changes with the work: an answer, a code artifact, a professional deliverable, or an improved method.

Delivery

The task record and the machinery around it.

Each release is assembled around the requirements of its task rather than forced into one universal schema.

Data records

Task instructions, inputs, structured metadata, and—where the task permits—reference answers, completed artifacts, or research trajectories.

Execution environments

Container definitions, dependencies, fixtures, tools, starter artifacts, and run instructions needed to reproduce the task state.

Evaluation protocols

Task-specific tests, graders, or held-out re-execution procedures that turn a submitted answer or artifact into a reproducible result.