Every public benchmark rots: it is published, it becomes the reference, it is scraped into the next model's training data, and from that moment a high score measures memorisation instead of ability. There are only two stable answers — keep the test set secret, or continuously replace it. This track does both, and below it does it in front of you.
Mint → seal → score → rotate → audit. Every panel below is a replay of one real run; nothing is illustrative.
The shape is published. The contents are not. The hash is written into the ledger before anything is scored, which is what makes “this set existed in final form first” checkable rather than assertable.
Run the same model on the sealed private set and on the public control arm. For an honest model the two agree to within sampling error. The comparison only exists because one harness runs both tracks — which is why the free public leaderboard runs forever: it is the control arm.
Generation 1 expired on its published date, was retired, and was republished in full as a dated, known-contaminated public benchmark. Generation 2 is fresh accrual the leaked model has never seen.
Rotation creates a measurement problem it also has to solve: a fresh generation is not the same difficulty as the one it replaces. A -case anchor sits inside both generations — the same physical cases — so the difference is measured, not assumed. The anchor never enters a ranked score, and a flagged model is excluded from the calibration pool so contamination cannot move the scale.
The auditor receives the seal manifest, the ledger, and a de-identified bundle — per case, a salted identifier, its published hash, the true label and the predicted probabilities. No image. No feature. No patient attribute. From that alone the published score is re-derived and the ordering of seal-before-score is proved. This is the artefact a regulator eventually asks for.