Training data
The working image and the complete research trajectory it produces: every version, the idea behind it, the code, the visible metric, and whether it was kept or reverted. Failed ideas stay in the record instead of disappearing behind the final answer.
Research environmentInherited methodVersion-level trajectoryKept and reverted attempts
Evaluation data
The verifier image, hidden split, trusted scoring code and at least one reference solution measured end to end. Stronger references can be added where they exist. Training data and scoring data never cross the seal.
Verifier imageSealed hidden splitMeasured referenceHeld-out re-execution
Diagnosis
The visible trajectory, final held-out result and reference runs can be read together. They show whether the run stopped iterating, overfit the visible set, ranked candidates worse despite a better model of the problem, or spent its budget in the wrong place.
Visible-to-hidden gapCost by versionKept vs reverted ideasNext failure to target