MedEval-1
MedicineEvery dimension chance-corrected, scored per task and ranked within a modality class. Failures are published by name.
One instrument, run per domain. Every benchmark below is scored on the same 9 dimensions, from the same harness, with the same seed and the same code fingerprint — and each one opens on the models, the scores and the corpus behind it.
Every dimension chance-corrected, scored per task and ranked within a modality class. Failures are published by name.
Bean variety from grain photographs, and wine quality from an assay panel. The wine label is a sensory panel's median score, not ground truth, and the cut that turns it into two classes was made here rather than in the source — both facts travel with the number.
The same instrument on census income prediction — the strongest split provenance in the suite, and the corpus where dimension five is measured against real demographic subgroups rather than a class proxy.
The same dimensions, the same evaluators and the same seed, pointed at a consumer credit decision. Calibration and deferral are the two that need no translation at all.
Semiconductor fabrication pass/fail and steel plate surface faults. The fab corpus is split chronologically rather than at random, because a process-monitoring model is always asked to predict forward in time — and it has 590 sensor channels over 1,567 lots, so it is expected to be hard and the intervals say so.
The first non-medical corpus in the suite where dimension two is measured rather than nulled. The donors released one training period and two held-out test periods recorded under different physical conditions — a real re-collection of the same instrument, not a simulated perturbation.
One hundred long-tailed provision types over SEC contract text, on the LexGLUE published split. Plain accuracy flatters a frequency-matching baseline here, which is exactly why this corpus is useful to the calibration and deferral dimensions rather than to the accuracy one.
Statlog Landsat, four spectral bands over a 3x3 neighbourhood, on the donors' own released train/test split. The donors withdrew one class before release, so the raw codes are not contiguous — the gap is remapped and disclosed rather than quietly closed.
Anuran calls as MFCCs, split by recording and never by row. Syllables from one recording are near-duplicates, so a random split scores memorisation and returns about 99%. The grouped split is materially harder and is the only honest one.
Dimension five asks who a model fails. Answering it needs the subject attributes — sex, age, race — and almost no public medical imaging corpus ships them. Where they are missing the harness says so in the record and falls back to worst-diagnostic-class recall, which is a weaker question honestly labelled rather than a stronger one faked. Census-derived and credit corpora carry those attributes as first-class columns, so on the labour and financial tasks dimension five is measured against real demographic subgroups. Same instrument, different corpus. That is why the non-medical work is not a side project: it is where half of this standard gets exercised.
The gap is not simply medical-versus-not. One clinical tabular task in the suite carries age but is single-sex by construction, and it is in the registry as a worked example of a corpus that cannot support a sex-gap audit at all.
| Dimension | In medicine | Outside medicine |
|---|---|---|
| 1Task accuracy | Right on held-out patients? | Right on held-out applicants? |
| 2Cross-site robustness Not measured on either non-medical task, and deliberately so. The harness computes this dimension from genuine re-acquisition of the same subjects — the same abdomen scanned in a different plane. Neither the census nor the credit corpus supports that construction, so the records carry an explicit null and the column renders as an em-dash. There is a defensible cross-population analogue — train on one subgroup, test on another — but shipping it under dimension 2's published name would be quietly widening a definition to fill a table, so it is a spec change with a new dimension id, not a demo. | Survives being re-imaged elsewhere? | Survives a different population or period? |
| 3Calibration Identical in both readings — and it is the one underwriting actually needs. A credit model that is right often enough but confident at the wrong moments is not an underwritable model. | Are its confidences worth believing? | Are its confidences worth believing?· unchanged |
| 4Limited data | How much labelled data before it works? | How much labelled data before it works?· unchanged |
| 5Subgroup safety Fully measurable here, and that is the finding. See SUBGROUP_FINDING below. | Who does it fail? | Who does it fail?· unchanged |
| 6Corruption / OOD Scored on all three modalities, with a family per modality rather than one family stretched across them. Images get seven degradations; tables get field dropout, coding drift, unit change and entry error; documents get vocabulary drift, formatting damage, truncation, register change and OCR noise. Each runs at three severities against the model's own clean answers. | Survives a bad image? | Survives missing fields, entry errors, shifted distributions? |
| 7Uncertainty & deferral | Knows when it doesn't know? | Knows when to route to a human underwriter? |