Method
Evaluation and the memorisation control
How runs are scored, and why a solved problem sits in the problem set.
Runs are not scored on whether they solve anything. They are scored on whether the increment is checkable, whether the provenance labels survive inspection, and whether the identified obstruction matches the one the literature documents.
Why a solved problem is in the set
Ringel's Conjecture was settled by Montgomery, Pokrovskiy and Sudakov in 2020, and its proof is well represented in any recent training corpus. That makes it useless as a target and invaluable as a control: it is the one problem where a model can succeed entirely by recall, so it is the one place we can measure the gap between reconstruction and recitation.
A run is scored as reconstruction only if it rebuilds the absorption argument rather than quoting it, explains why a purely greedy edge-disjoint embedding stalls and where randomness rescues it, and identifies why the decomposition uses 2n+1 copies rather than 2n. Recitation is common; reconstruction is not. Everything else in the problem set is calibrated against that ratio.
What would count as signal
Not a proof. The realistic positive outcomes are narrower: an obstruction rediscovered without being prompted with it, a correct judgement about which of two dead ends is less dead, or a reformulation that a working mathematician finds worth pursuing. Each is a weak signal individually. The measurement is whether they occur more often than chance across a large number of runs.
Selection effects
Runs are never re-rolled. Publishing only the interesting transcripts would make the archive a selection artifact and destroy its value as a measurement, so the first pass is what gets published — including the large majority that restate the problem and stop.