MedEval-1
LIVEA standard for measuring medical AI: what is measured, how it is measured, how it is reported, and what makes a result admissible. The leaderboard below is the deployed artifact itself, embedded rather than re-implemented, so nothing in this shell can move one of its numbers. Read the specification, or open the standalone page.
Failed evaluations
Published by name, on purpose.A model that cannot be evaluated on a task is published as a failure on that task, by name. It is never a zero, and never a silent omission.
Evaluations that did not complete
0 of the 196 records in this run failed to produce a score. The section stays regardless, because a table that only appears when it is empty of embarrassment is not a disclosure.
Designed controls that failed certification, as intended
Prognostic-model trial validity (PROCOVA). These are not models of ours that went wrong. They are inputs constructed in order to fail, so that a PASS elsewhere in the table is evidence of something rather than an assumption — the negative control in an assay. Each one produced numbers, and those numbers are published in full; what is withheld is only its place in the ranking, because the procedure is invalid by construction rather than merely worse than the others.
type-I error 0.236 vs nominal 0.050 -- 95% CI [0.218, 0.255] excludes nominal; treatment effect biased by -0.532 pts
Not a model -- a procedure. Every candidate is a legitimate pre-randomisation score; the selection rule is what breaks the trial. Invisible to a protocol review, visible to a type-I measurement.
type-I error 0.345 vs nominal 0.050 -- 95% CI [0.324, 0.366] excludes nominal; treatment effect biased by -0.133 pts
Not a model -- a procedure. Every candidate is a legitimate pre-randomisation score; the selection rule is what breaks the trial. Invisible to a protocol review, visible to a type-I measurement.
The leaderboard above is the whole run, not just the medical rows: it renders all 196 evaluations across 35 tasks, because the harness is one instrument and MedEval-1 is one of the standards it serves. The non-medical rows have their own pages — Benchmarks lists every one of them, with the same 9 dimensions read per domain.
Medicine
Where the question is hardest, so where the standard got built first.
MedEval-1 is the first full instantiation of the standard: a task registry, data adapters, and nine dimension evaluators, run over public medical imaging, tabular clinical and clinical text corpora. Every row carries its seed, its split sizes, and a fingerprint of the scoring code that produced it.
Medicine came first because medicine is where a model's output is hardest to act on: the regulation is heaviest, the failure is worst, and the public evidence is thinnest. A standard that survives here is not going to be embarrassed by an easier domain.
- Not a published or peer-reviewed benchmark. The models are open baselines, not clinical products, and nothing here is a claim about any deployed system.
- Dimension five falls back to worst-diagnostic-class on every imaging task, because no public imaging corpus in the suite ships subject demographics.
- Two of the nine dimensions are thin, and they still carry composite weight. Specification sensitivity is scored on 12 of the 196 records — all of them the same model — and cost on 45. Read the composite as seven broadly-measured dimensions plus two narrow ones, not as nine equally-evidenced columns.