NakedSignal OS · Documentation
Back to the consoleDocumentation

Benchmarks

LIVE

One instrument, run per domain. Every benchmark below is scored on the same 9 dimensions, from the same harness, with the same seed and the same code fingerprint — and each one opens on the models, the scores and the corpus behind it.

MedEval-1

Medicine
LIVE

Every dimension chance-corrected, scored per task and ranked within a modality class. Failures are published by name.

22
models scored
21
tasks
11
modalities
134
scored records
Tasks in the registry
Colorectal pathology (9-class tissue) HistopathologyDermatoscopy (7-class skin lesion) DermatoscopyPeripheral blood cells (8-class) MicroscopyRetinal OCT (4-class) OCTPaediatric chest X-ray (pneumonia) Chest X-rayFundus photography (DR grade) FundusBreast ultrasound (malignancy) UltrasoundKidney cortex cells (8-class) MicroscopyAbdominal CT organs — axial (A) CTAbdominal CT organs — coronal (C) CTAbdominal CT organs — sagittal (S) CTCoronary artery disease (Cleveland) Tabular / clinicalDiabetes onset (Pima) Tabular / clinicalBreast cytology (WDBC) Tabular / imaging-derivedMedical abstracts (5-class condition) Clinical textBreast ultrasound (malignancy) @224px UltrasoundFundus photography (DR grade) @224px FundusPaediatric chest X-ray (pneumonia) @224px Chest X-rayAbdominal CT organs — coronal (C) @224px CTDermatoscopy (7-class skin lesion) @224px DermatoscopyPeripheral blood cells (8-class) @224px Microscopy

Produce grading

Agriculture & foodrun, no separate spec
LIVE

Bean variety from grain photographs, and wine quality from an assay panel. The wine label is a sensory panel's median score, not ground truth, and the cut that turns it into two classes was made here rather than in the source — both facts travel with the number.

4
models scored
3
tasks
2
modalities
12
scored records
Tasks in the registry
Dry bean variety (7-class) Tabular / imaging-derivedWine quality, white (binarised) Tabular / physicochemicalWine quality, red (binarised) Tabular / physicochemical

Census income

Labour marketrun, no separate spec
LIVE

The same instrument on census income prediction — the strongest split provenance in the suite, and the corpus where dimension five is measured against real demographic subgroups rather than a class proxy.

11
models scored
2
tasks
2
modalities
11
scored records
Tasks in the registry
Census income (>$50K) Tabular / socioeconomicFraudulent job posting detection Recruitment text

Consumer credit risk

Financialrun, no separate spec
LIVE

The same dimensions, the same evaluators and the same seed, pointed at a consumer credit decision. Calibration and deferral are the two that need no translation at all.

4
models scored
2
tasks
2
modalities
8
scored records
Tasks in the registry
Consumer credit risk (Statlog) Tabular / financialCaravan insurance purchase (CoIL 2000) Tabular / insurance

Process and surface inspection

Industrial inspectionrun, no separate spec
LIVE

Semiconductor fabrication pass/fail and steel plate surface faults. The fab corpus is split chronologically rather than at random, because a process-monitoring model is always asked to predict forward in time — and it has 590 sensor channels over 1,567 lots, so it is expected to be hard and the intervals say so.

4
models scored
2
tasks
2
modalities
8
scored records
Tasks in the registry
Semiconductor fabrication pass/fail Tabular / process sensorsSteel plate surface faults (7-class) Tabular / imaging-derived

Room occupancy from sensors

Energy & buildingsrun, no separate spec
LIVE

The first non-medical corpus in the suite where dimension two is measured rather than nulled. The donors released one training period and two held-out test periods recorded under different physical conditions — a real re-collection of the same instrument, not a simulated perturbation.

4
models scored
2
tasks
1
modality
8
scored records
Tasks in the registry
Room occupancy from sensors (period 1) Tabular / environmental sensorsRoom occupancy from sensors (period 2) Tabular / environmental sensors

Contract provision type

Legalrun, no separate spec
LIVE

One hundred long-tailed provision types over SEC contract text, on the LexGLUE published split. Plain accuracy flatters a frequency-matching baseline here, which is exactly why this corpus is useful to the calibration and deferral dimensions rather than to the accuracy one.

Land cover from satellite

Earth observationrun, no separate spec
LIVE

Statlog Landsat, four spectral bands over a 3x3 neighbourhood, on the donors' own released train/test split. The donors withdrew one class before release, so the raw codes are not contiguous — the gap is remapped and disclosed rather than quietly closed.

4
models scored
1
task
1
modality
4
scored records
Tasks in the registry
Landsat satellite land cover (6-class) Tabular / multispectral

Species from field recordings

Ecology & bioacousticsrun, no separate spec
LIVE

Anuran calls as MFCCs, split by recording and never by row. Syllables from one recording are near-duplicates, so a random split scores memorisation and returns about 99%. The grouped split is materially harder and is the only honest one.

4
models scored
1
task
1
modality
4
scored records
Tasks in the registry
Anuran call species (10-class) Bioacoustic / MFCC

Dimension five is the one medicine cannot fully answer.

Dimension five asks who a model fails. Answering it needs the subject attributes — sex, age, race — and almost no public medical imaging corpus ships them. Where they are missing the harness says so in the record and falls back to worst-diagnostic-class recall, which is a weaker question honestly labelled rather than a stronger one faked. Census-derived and credit corpora carry those attributes as first-class columns, so on the labour and financial tasks dimension five is measured against real demographic subgroups. Same instrument, different corpus. That is why the non-medical work is not a side project: it is where half of this standard gets exercised.

The gap is not simply medical-versus-not. One clinical tabular task in the suite carries age but is single-sex by construction, and it is in the registry as a worked example of a corpus that cannot support a sex-gap audit at all.

The same 7 dimensions, read per domain

One instrument. The question does not change; what answers it does.
DimensionIn medicineOutside medicine
1Task accuracyRight on held-out patients?Right on held-out applicants?
2Cross-site robustness
Not measured on either non-medical task, and deliberately so. The harness computes this dimension from genuine re-acquisition of the same subjects — the same abdomen scanned in a different plane. Neither the census nor the credit corpus supports that construction, so the records carry an explicit null and the column renders as an em-dash. There is a defensible cross-population analogue — train on one subgroup, test on another — but shipping it under dimension 2's published name would be quietly widening a definition to fill a table, so it is a spec change with a new dimension id, not a demo.
Survives being re-imaged elsewhere?Survives a different population or period?
3Calibration
Identical in both readings — and it is the one underwriting actually needs. A credit model that is right often enough but confident at the wrong moments is not an underwritable model.
Are its confidences worth believing?Are its confidences worth believing?· unchanged
4Limited dataHow much labelled data before it works?How much labelled data before it works?· unchanged
5Subgroup safety
Fully measurable here, and that is the finding. See SUBGROUP_FINDING below.
Who does it fail?Who does it fail?· unchanged
6Corruption / OOD
Scored on all three modalities, with a family per modality rather than one family stretched across them. Images get seven degradations; tables get field dropout, coding drift, unit change and entry error; documents get vocabulary drift, formatting damage, truncation, register change and OCR noise. Each runs at three severities against the model's own clean answers.
Survives a bad image?Survives missing fields, entry errors, shifted distributions?
7Uncertainty & deferralKnows when it doesn't know?Knows when to route to a human underwriter?