Commercial credit underwriting
Ratio baseline → cash-flow reconciliation → collateral eligibility → risk rating → credit memo
Download full sampleEnd-to-end environments for finance, engineering, operations, and other expert work. AI systems work across files, tools, and prior deliverables; each step produces a professional artifact checked by an executable grader.
The subject matter changes from one environment to the next. The underlying contract does not: the work is situated, stateful, consequential and verifiable.
A credit analyst receives applications, statements and lender policy. A grid operator receives network cases, telemetry and contingency rules. Each environment is built from the evidence, terminology and constraints of its own profession.
An analysis becomes the input to a model; the model informs a decision package. Earlier assumptions and errors remain in the workspace, so later tasks measure whether the AI can maintain coherence over a sustained engagement.
The task definition, container, mock services, workspaces and graders travel together. The same package can support training, regression testing or model evaluation without rebuilding the world around the task.
Each environment is different in content but consistent in construction, so a large library can be run and compared under the same interface.
A situated brief. The AI enters a specific organization, role and moment in an ongoing engagement, with a concrete request and the context needed to act.
Native evidence and tools. Spreadsheets, PDFs, scans, email, audio, logs, configuration and domain software appear in the forms used on the job—including material that may be incomplete, inconsistent or irrelevant.
Five dependent tasks. Each task advances the same engagement. Outputs remain in the workspace and become evidence for what follows, turning isolated capability into long-horizon execution.
Professional deliverables. The result is the actual artifact the work calls for: a workbook, model, memo, report, operating package, script or structured evidence bundle.
Executable grading. Graders recompute key quantities, inspect required evidence, enforce tolerances and test release conditions. Evaluation follows the deliverable rather than the AI's account of what it completed.
These downloads show the common environment format across very different kinds of work. Each contains the runnable environment, task inputs, executable checks, reference answers, and instructions for running it. Unrelated prior runs, provenance records, and identifying metadata are omitted.
Ratio baseline → cash-flow reconciliation → collateral eligibility → risk rating → credit memo
Download full sampleTrial-balance normalization → intercompany reconciliation → FX translation → eliminations → consolidated close package
Download full sampleNetwork-case intake → telemetry audit → corrective redispatch → N-1 screening → final release package
Download full sampleOperating point → discrete model → LQR design → commissioning run → heavy-line validation
Download full sampleClosed-ended research problems with determinate reference answers, from graduate-level material to questions that remain difficult for frontier models.
Repository-level tasks with a runnable environment and a concrete artifact contract. The submitted code is evaluated by an executable test suite.
Runnable research environments in which an AI system iterates on a working method. The submitted method is then re-executed on a sealed split.