BiomedCLIP ViT-B/16, zero-shot
LIVEThe signed evidence record for biomedclip_zeroshot. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.
biomedclip_zeroshotThe check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.
Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 17 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.
| Task | Domain | n test | accuracy | calibration | uncertainty | subgroup | corruption | acquisition | limited data | cost | spec sensitivity | Index |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bloodmnist_224 | medicine | 3421 | 1.323337 | 0 | 15.473877 | 0 | 69.97754 | 70.975895 | n/a | 100 | 11.215789 | 29.157445 |
| bloodmnist | medicine | 3421 | 0 | 0 | 0 | 0 | not scored | not scored | n/a | 100 | 0 | 4.918033 |
| breastmnist_224 | medicine | 156 | 20.050125 | 0 | 8.587353 | 0 | 79.732143 | 98.833333 | n/a | 100 | 2.325581 | 39.120227 |
| breastmnist | medicine | 156 | 6.015038 | 0 | 56.764976 | 0 | 79.861111 | 86.944444 | n/a | 100 | 0 | 36.0029 |
| dermamnist_224 | medicine | 2005 | 16.413154 | 0 | 15.345893 | 0 | 55.212509 | 78.320304 | n/a | 100 | 28.304676 | 33.35518 |
| dermamnist | medicine | 2005 | 0 | 0 | 0 | 0 | not scored | not scored | n/a | 100 | 0 | 4.918033 |
| organcmnist_224 | medicine | 8216 | 13.567459 | 0 | 26.303931 | 0 | 57.799225 | 85.676587 | n/a | 100 | 30.922684 | 35.25083 |
| organcmnist | medicine | 8216 | 3.519576 | 0 | 1.34789 | 0 | 50.303286 | 63.516083 | n/a | 100 | 0 | 23.567278 |
| pneumoniamnist_224 | medicine | 624 | 33.675214 | 0 | 35.923686 | 0 | 57.275804 | 75.228426 | n/a | 100 | 8.01105 | 36.579413 |
| pneumoniamnist | medicine | 624 | 23.760684 | 0 | 18.732307 | 0 | 56.509078 | 73.717026 | n/a | 100 | 25 | 34.355577 |
| retinamnist_224 | medicine | 400 | 0 | 3.912164 | 42.05505 | 0 | not scored | not scored | n/a | 100 | 0 | 9.198908 |
| retinamnist | medicine | 400 | 0 | 25.348975 | 22.846236 | 0 | not scored | not scored | n/a | 100 | 0 | 12.192916 |
Cross-site retention
not measured for this subject — it is only defined within a family of tasks that share a corpus and differ in acquisition, and this model was not run on one.Private held-out track
this subject was not submitted to a sealed generation.Known limitations
Audit chain
python src/heldout/audit.py 1f66c7719ab3943c6fcc17 source files, with hashes
| results/heldout/audit_report.json | df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2 |
| results/heldout/cycle.json | e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502 |
| results/heldout/manifests/gen1.seal.json | 91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81 |
| results/heldout/manifests/gen2.seal.json | f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547 |
| results/heldout/manifests/public.seal.json | 1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c |
| results/records/bloodmnist_224__biomedclip_zeroshot.json | 28d0f988885cab16dc30d8b4ffc350e9510752d4d372238a42be3cbc4cba804a |
| results/records/bloodmnist__biomedclip_zeroshot.json | 4437196da05ecf58b4d6f6f43c789ff677550a651d2f303f38e9a0fc81d373c6 |
| results/records/breastmnist_224__biomedclip_zeroshot.json | d8825e0fbf3c87d21f97dd6075bb8e28057de5d3c86f2145760d80c91bbec33b |
| results/records/breastmnist__biomedclip_zeroshot.json | f335b76d2529620a9ebb2d8241df0319a4d20c6ed395cee1e4ed7438b6ff52c9 |
| results/records/dermamnist_224__biomedclip_zeroshot.json | 77781b199642e67b1b7ab3a4711f69a9c9b2c6cde053837af744e6d27cc5f4dd |
| results/records/dermamnist__biomedclip_zeroshot.json | 5c04e18568422f094de0e013e1dcfe1ae2fa09a43ae043ab08d6cd82fd35a7f6 |
| results/records/organcmnist_224__biomedclip_zeroshot.json | d7c2ca5a1023bc5e600eb6749e3301b2a85f16694075259fba2472183a77612c |
| results/records/organcmnist__biomedclip_zeroshot.json | 6a445673b017f9518bb4965470eb87ad70ad665dfea783adf971f0344eb6a00c |
| results/records/pneumoniamnist_224__biomedclip_zeroshot.json | d4b8bc2215c65e86612ced45c4f036a2d22914cdfc8e4a050d44622f392e09f8 |
| results/records/pneumoniamnist__biomedclip_zeroshot.json | df1367e80e2f9d245d2a3ff5ed021883f80fcdf8274a52017f59a874d7433291 |
| results/records/retinamnist_224__biomedclip_zeroshot.json | 7d840843cb52f5667fd0a7c65acc6acf01d24d3cbf07b908d6d9cbec399131aa |
| results/records/retinamnist__biomedclip_zeroshot.json | 345a60f9f365f8b0871b31fcfc0add607bf81cc636e04d8ef136f10e2a6b134f |
Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.
No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.
Production history and drift — how the model has actually behaved since it was deployed, on real traffic.
Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.
Raw JSON — the whole signed document, 119,932 characters
{
"canonicalisation": {
"float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
"form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
"signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
},
"computed": {
"audit": {
"audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
"bundles": 48,
"ledger_entries": 100,
"ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
"ledger_intact": true,
"rederived": 48
},
"cross_site": null,
"evaluations": [
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 29.157445
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.197895,
"auc": 0.634089,
"balanced_accuracy": 0.136579,
"ci95": [
0.131291,
0.141767
],
"n_test": 3421,
"score": 1.323337
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.0168,
"mean_agreement": 0.839056,
"mean_kappa": 0.575316,
"mean_retention": 0.709759,
"n_cases": 1200,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.9133,
0.7158,
0.6092
],
"mean_kappa": 0.401075,
"mean_retention": 0.471296,
"retention_by_severity": [
0.8917,
0.5222,
0
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.8475,
0.7525,
0.6158
],
"mean_kappa": 0.397528,
"mean_retention": 0.896583,
"retention_by_severity": [
0.7074,
0.9824,
1
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.9558,
0.91,
0.8525
],
"mean_kappa": 0.729872,
"mean_retention": 0.508303,
"retention_by_severity": [
0.5512,
0.4366,
0.5371
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.9592,
0.9425,
0.9283
],
"mean_kappa": 0.834777,
"mean_retention": 0.834501,
"retention_by_severity": [
0.9053,
0.8236,
0.7746
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.915,
0.8567,
0.8117
],
"mean_kappa": 0.513329,
"mean_retention": 0.838112,
"retention_by_severity": [
0.8451,
0.7749,
0.8943
],
"scored": true
}
},
"score": 70.975895,
"worst_retention": 0
},
"calibration": {
"accuracy": 0.197895,
"brier": 1.161568,
"ece": 0.436072,
"mean_confidence": 0.633967,
"overconfidence": 0.436072,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.0168,
"mean_agreement": 0.569167,
"mean_kappa": 0.361946,
"mean_retention": 0.699775,
"n_cases": 1200,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.895,
0.8308,
0.7017
],
"mean_kappa": 0.406363,
"mean_retention": 0.824929,
"retention_by_severity": [
0.8779,
0.5969,
1
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.9283,
0.8675,
0.7783
],
"mean_kappa": 0.635892,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.9692,
0.9283,
0.8833
],
"mean_kappa": 0.78103,
"mean_retention": 0.781059,
"retention_by_severity": [
0.8709,
0.9069,
0.5654
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.7233,
0.5258,
0.0092
],
"mean_kappa": 0.144907,
"mean_retention": 0.296178,
"retention_by_severity": [
0.8885,
0,
0
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.0117,
0.0042,
0.0017
],
"mean_kappa": 0.002038,
"mean_retention": 0.901249,
"retention_by_severity": [
1,
0.7037,
1
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9483,
0.7975,
0.4208
],
"mean_kappa": 0.51251,
"mean_retention": 0.900599,
"retention_by_severity": [
0.7018,
1,
1
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.6308,
0.0933,
0.0033
],
"mean_kappa": 0.050882,
"mean_retention": 0.194414,
"retention_by_severity": [
0.5832,
0,
0
],
"scored": true
}
},
"score": 69.97754,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 47821,
"cases_per_second": 121.784,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.008211,
"seconds_total": 392.6718
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.060714,
"best_spec": "template_2",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.1397,
"chance_corrected": 0.0168,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.136323,
"chance_corrected": 0.012941,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.178125,
"chance_corrected": 0.060714,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.173234,
"chance_corrected": 0.055125,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.130958,
"chance_corrected": 0.00681,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.133846,
"chance_corrected": 0.01011,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.112158,
"score": 11.215789,
"scored": true,
"spread_balanced_accuracy": 0.047167,
"worst_chance_corrected": 0.00681,
"worst_spec": "template_4"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.197895,
"risk_coverage_auc": 0.260396,
"score": 15.473877,
"selective_acc_at_50": 0.294152,
"selective_acc_at_80": 0.229814
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 11959,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": true
},
"modality": "Microscopy",
"n_classes": 8,
"n_test": 3421,
"runtime_s": 546.9,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "bloodmnist_224",
"task_name": "Peripheral blood cells (8-class) @224px",
"timestamp": "2026-08-01T19:59:26+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.61,
"dimensions_scored": [
"accuracy",
"calibration",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 4.918033
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.126863,
"auc": 0.590911,
"balanced_accuracy": 0.114617,
"ci95": [
0.104209,
0.125475
],
"n_test": 3421,
"score": 0
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 1200,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.6025,
0.2975,
0.2367
],
"mean_kappa": 0.16666,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"field_of_view": {
"agreement_by_severity": [
0.6508,
0.5458,
0.4083
],
"mean_kappa": 0.345899,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gamma_shift": {
"agreement_by_severity": [
0.8108,
0.6667,
0.58
],
"mean_kappa": 0.515929,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"resolution": {
"agreement_by_severity": [
0.57,
0.4742,
0.4092
],
"mean_kappa": 0.277632,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"window_level": {
"agreement_by_severity": [
0.6733,
0.4892,
0.4417
],
"mean_kappa": 0.325299,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"calibration": {
"accuracy": 0.126863,
"brier": 1.127525,
"ece": 0.385935,
"mean_confidence": 0.512798,
"overconfidence": 0.385935,
"score": 0
},
"corruption": {
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 1200,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.5783,
0.4267,
0.41
],
"mean_kappa": 0.232991,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"contrast": {
"agreement_by_severity": [
0.8008,
0.6208,
0.3642
],
"mean_kappa": 0.399405,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"defocus_blur": {
"agreement_by_severity": [
0.6458,
0.4508,
0.2
],
"mean_kappa": 0.247683,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gaussian_noise": {
"agreement_by_severity": [
0.315,
0.2383,
0.2358
],
"mean_kappa": 0.031664,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"pixelate": {
"agreement_by_severity": [
0.4833,
0.0483,
0.1883
],
"mean_kappa": 0.118947,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"quantise": {
"agreement_by_severity": [
0.805,
0.5842,
0.3892
],
"mean_kappa": 0.435747,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"shot_noise": {
"agreement_by_severity": [
0.2383,
0.2358,
0.2342
],
"mean_kappa": 0.002387,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 47821,
"cases_per_second": 121.905,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.008203,
"seconds_total": 392.2794
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.087689,
"best_spec": "template_1",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.104782,
"chance_corrected": 0,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.201727,
"chance_corrected": 0.087689,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.133236,
"chance_corrected": 0.009412,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.115581,
"chance_corrected": 0,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.143801,
"chance_corrected": 0.021486,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.130824,
"chance_corrected": 0.006655,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.096946,
"worst_chance_corrected": 0,
"worst_spec": "template_0"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.126863,
"risk_coverage_auc": 0.121242,
"score": 0,
"selective_acc_at_50": 0.122807,
"selective_acc_at_80": 0.124954
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 11959,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Microscopy",
"n_classes": 8,
"n_test": 3421,
"runtime_s": 401.3,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "bloodmnist",
"task_name": "Peripheral blood cells (8-class)",
"timestamp": "2026-08-01T20:08:33+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 39.120227
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.448718,
"auc": 0.767753,
"balanced_accuracy": 0.600251,
"ci95": [
0.546294,
0.653504
],
"n_test": 156,
"score": 20.050125
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.200501,
"mean_agreement": 0.869658,
"mean_kappa": 0.684586,
"mean_retention": 0.988333,
"n_cases": 156,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.9423,
0.8077,
0.5962
],
"mean_kappa": 0.549982,
"mean_retention": 0.995833,
"retention_by_severity": [
1,
0.9875,
1
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.9103,
0.8718,
0.8782
],
"mean_kappa": 0.686845,
"mean_retention": 0.96875,
"retention_by_severity": [
1,
1,
0.9062
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.9679,
0.9423,
0.8974
],
"mean_kappa": 0.821871,
"mean_retention": 0.989583,
"retention_by_severity": [
1,
0.9688,
1
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.9423,
0.9038,
0.8654
],
"mean_kappa": 0.737433,
"mean_retention": 0.9875,
"retention_by_severity": [
0.9813,
0.9813,
1
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.9423,
0.8526,
0.7244
],
"mean_kappa": 0.626802,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
}
},
"score": 98.833333,
"worst_retention": 0.90625
},
"calibration": {
"accuracy": 0.448718,
"brier": 0.909597,
"ece": 0.446057,
"mean_confidence": 0.894775,
"overconfidence": 0.446057,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.200501,
"mean_agreement": 0.812882,
"mean_kappa": 0.516569,
"mean_retention": 0.797321,
"n_cases": 156,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9487,
0.8974,
0.8526
],
"mean_kappa": 0.748554,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.9808,
0.9359,
0.859
],
"mean_kappa": 0.787331,
"mean_retention": 0.747917,
"retention_by_severity": [
0.9562,
0.6,
0.6875
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.9487,
0.8077,
0.5449
],
"mean_kappa": 0.550654,
"mean_retention": 0.95,
"retention_by_severity": [
0.85,
1,
1
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.8526,
0.5705,
0.5321
],
"mean_kappa": 0.340724,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.7885,
0.7756,
0.7756
],
"mean_kappa": 0.006677,
"mean_retention": 0.014583,
"retention_by_severity": [
0.0438,
0,
0
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9744,
0.9359,
0.8526
],
"mean_kappa": 0.755937,
"mean_retention": 0.86875,
"retention_by_severity": [
1,
1,
0.6062
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.8141,
0.6795,
0.7436
],
"mean_kappa": 0.42611,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
}
},
"score": 79.732143,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 5772,
"cases_per_second": 116.604,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.008576,
"seconds_total": 49.5008
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.377193,
"best_spec": "template_5",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.600251,
"chance_corrected": 0.200501,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.567043,
"chance_corrected": 0.134085,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.530702,
"chance_corrected": 0.061404,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.504386,
"chance_corrected": 0.008772,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.530702,
"chance_corrected": 0.061404,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.688596,
"chance_corrected": 0.377193,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.023256,
"score": 2.325581,
"scored": true,
"spread_balanced_accuracy": 0.184211,
"worst_chance_corrected": 0.008772,
"worst_spec": "template_3"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.27193
},
"uncertainty": {
"full_accuracy": 0.448718,
"risk_coverage_auc": 0.542937,
"score": 8.587353,
"selective_acc_at_50": 0.525641,
"selective_acc_at_80": 0.456
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 546,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Ultrasound",
"n_classes": 2,
"n_test": 156,
"runtime_s": 57.8,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "breastmnist_224",
"task_name": "Breast ultrasound (malignancy) @224px",
"timestamp": "2026-08-01T19:45:20+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 36.0029
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.730769,
"auc": 0.599415,
"balanced_accuracy": 0.530075,
"ci95": [
0.490243,
0.581737
],
"n_test": 156,
"score": 6.015038
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.06015,
"mean_agreement": 0.75812,
"mean_kappa": 0.327646,
"mean_retention": 0.869444,
"n_cases": 156,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.7885,
0.4359,
0.1154
],
"mean_kappa": 0.101765,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.8269,
0.7821,
0.7244
],
"mean_kappa": 0.235987,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.9744,
0.9487,
0.9103
],
"mean_kappa": 0.570658,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.7115,
0.5833,
0.6987
],
"mean_kappa": 0.145787,
"mean_retention": 0.527778,
"retention_by_severity": [
1,
0.4792,
0.1042
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.9744,
0.9615,
0.9359
],
"mean_kappa": 0.584031,
"mean_retention": 0.819444,
"retention_by_severity": [
1,
1,
0.4583
],
"scored": true
}
},
"score": 86.944444,
"worst_retention": 0.104167
},
"calibration": {
"accuracy": 0.730769,
"brier": 0.463927,
"ece": 0.202305,
"mean_confidence": 0.922621,
"overconfidence": 0.191852,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.06015,
"mean_agreement": 0.618437,
"mean_kappa": 0.266298,
"mean_retention": 0.798611,
"n_cases": 156,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9808,
0.9744,
0.9231
],
"mean_kappa": 0.664735,
"mean_retention": 0.430556,
"retention_by_severity": [
1,
0.1667,
0.125
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.9679,
0.8974,
0.7051
],
"mean_kappa": 0.451446,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.8782,
0.6987,
0.8718
],
"mean_kappa": 0.225214,
"mean_retention": 0.777778,
"retention_by_severity": [
1,
1,
0.3333
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.5962,
0.1667,
0.0641
],
"mean_kappa": 0.045655,
"mean_retention": 0.763889,
"retention_by_severity": [
1,
1,
0.2917
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.7179,
0.2564,
0.3205
],
"mean_kappa": 0.078122,
"mean_retention": 0.902778,
"retention_by_severity": [
0.8333,
0.875,
1
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9487,
0.8846,
0.5833
],
"mean_kappa": 0.37938,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.3718,
0.1218,
0.0577
],
"mean_kappa": 0.019531,
"mean_retention": 0.715278,
"retention_by_severity": [
1,
1,
0.1458
],
"scored": true
}
},
"score": 79.861111,
"worst_retention": 0.125
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 5772,
"cases_per_second": 127.544,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.00784,
"seconds_total": 45.2551
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.154135,
"best_spec": "template_1",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.530075,
"chance_corrected": 0.06015,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.577068,
"chance_corrected": 0.154135,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.511905,
"chance_corrected": 0.02381,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.494987,
"chance_corrected": 0,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.541353,
"chance_corrected": 0.082707,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.495614,
"chance_corrected": 0,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.08208,
"worst_chance_corrected": 0,
"worst_spec": "template_3"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.095238
},
"uncertainty": {
"full_accuracy": 0.730769,
"risk_coverage_auc": 0.783825,
"score": 56.764976,
"selective_acc_at_50": 0.782051,
"selective_acc_at_80": 0.76
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 546,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Ultrasound",
"n_classes": 2,
"n_test": 156,
"runtime_s": 46.7,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "breastmnist",
"task_name": "Breast ultrasound (malignancy)",
"timestamp": "2026-08-01T19:46:18+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 33.35518
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.238903,
"auc": 0.690303,
"balanced_accuracy": 0.283541,
"ci95": [
0.24635,
0.32303
],
"n_test": 2005,
"score": 16.413154
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.169943,
"mean_agreement": 0.754611,
"mean_kappa": 0.614709,
"mean_retention": 0.783203,
"n_cases": 1200,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.8383,
0.6733,
0.5083
],
"mean_kappa": 0.488848,
"mean_retention": 0.698826,
"retention_by_severity": [
0.8172,
0.6722,
0.6071
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.8008,
0.7642,
0.6867
],
"mean_kappa": 0.619821,
"mean_retention": 0.713073,
"retention_by_severity": [
0.8542,
0.7365,
0.5485
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.8933,
0.8125,
0.7533
],
"mean_kappa": 0.723431,
"mean_retention": 0.94245,
"retention_by_severity": [
0.9281,
0.9507,
0.9485
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.8083,
0.7683,
0.7075
],
"mean_kappa": 0.629649,
"mean_retention": 0.761998,
"retention_by_severity": [
0.74,
0.8071,
0.739
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.8933,
0.7583,
0.6525
],
"mean_kappa": 0.611797,
"mean_retention": 0.799668,
"retention_by_severity": [
0.8564,
0.7803,
0.7623
],
"scored": true
}
},
"score": 78.320304,
"worst_retention": 0.548496
},
"calibration": {
"accuracy": 0.238903,
"brier": 0.983004,
"ece": 0.357605,
"mean_confidence": 0.594691,
"overconfidence": 0.355788,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.169943,
"mean_agreement": 0.539683,
"mean_kappa": 0.376942,
"mean_retention": 0.552125,
"n_cases": 1200,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.8433,
0.695,
0.575
],
"mean_kappa": 0.582437,
"mean_retention": 0.724649,
"retention_by_severity": [
1,
0.7884,
0.3855
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.6608,
0.3967,
0.2408
],
"mean_kappa": 0.318502,
"mean_retention": 0.607752,
"retention_by_severity": [
0.956,
0.6719,
0.1953
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.845,
0.7208,
0.5475
],
"mean_kappa": 0.568819,
"mean_retention": 0.670431,
"retention_by_severity": [
0.8218,
0.7007,
0.4888
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.6792,
0.5133,
0.2733
],
"mean_kappa": 0.285153,
"mean_retention": 0.640432,
"retention_by_severity": [
0.8007,
0.7812,
0.3394
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.3992,
0.3767,
0.2642
],
"mean_kappa": 0.12458,
"mean_retention": 0.064961,
"retention_by_severity": [
0.0315,
0.1634,
0
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.925,
0.72,
0.4917
],
"mean_kappa": 0.546417,
"mean_retention": 0.662379,
"retention_by_severity": [
0.9481,
0.7344,
0.3047
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.555,
0.3642,
0.2467
],
"mean_kappa": 0.212686,
"mean_retention": 0.494272,
"retention_by_severity": [
0.8083,
0.4029,
0.2717
],
"scored": true
}
},
"score": 55.212509,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 46405,
"cases_per_second": 141.256,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007079,
"seconds_total": 328.5159
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.169943,
"best_spec": "template_0",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.288523,
"chance_corrected": 0.169943,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.184087,
"chance_corrected": 0.048102,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.274595,
"chance_corrected": 0.153694,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.211492,
"chance_corrected": 0.080074,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.190909,
"chance_corrected": 0.056061,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.206552,
"chance_corrected": 0.074311,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.283047,
"score": 28.304676,
"scored": true,
"spread_balanced_accuracy": 0.104435,
"worst_chance_corrected": 0.048102,
"worst_spec": "template_1"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.137931
},
"uncertainty": {
"full_accuracy": 0.238903,
"risk_coverage_auc": 0.274393,
"score": 15.345893,
"selective_acc_at_50": 0.264471,
"selective_acc_at_80": 0.253741
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 7007,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": true
},
"modality": "Dermatoscopy",
"n_classes": 7,
"n_test": 2005,
"runtime_s": 455.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "dermamnist_224",
"task_name": "Dermatoscopy (7-class skin lesion) @224px",
"timestamp": "2026-08-01T12:39:38+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.61,
"dimensions_scored": [
"accuracy",
"calibration",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 4.918033
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.15212,
"auc": 0.556266,
"balanced_accuracy": 0.125952,
"ci95": [
0.107291,
0.150679
],
"n_test": 2005,
"score": 0
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 1200,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.4225,
0.22,
0.0608
],
"mean_kappa": 0.091706,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"field_of_view": {
"agreement_by_severity": [
0.8225,
0.7717,
0.6683
],
"mean_kappa": 0.556302,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gamma_shift": {
"agreement_by_severity": [
0.9,
0.8067,
0.7283
],
"mean_kappa": 0.662291,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"resolution": {
"agreement_by_severity": [
0.8033,
0.7725,
0.715
],
"mean_kappa": 0.55929,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"window_level": {
"agreement_by_severity": [
0.8792,
0.7658,
0.5917
],
"mean_kappa": 0.573185,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"calibration": {
"accuracy": 0.15212,
"brier": 1.336268,
"ece": 0.608845,
"mean_confidence": 0.760008,
"overconfidence": 0.607888,
"score": 0
},
"corruption": {
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 1200,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.8817,
0.7625,
0.5758
],
"mean_kappa": 0.544322,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"contrast": {
"agreement_by_severity": [
0.7925,
0.68,
0.6075
],
"mean_kappa": 0.39402,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"defocus_blur": {
"agreement_by_severity": [
0.8267,
0.7125,
0.6367
],
"mean_kappa": 0.466489,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gaussian_noise": {
"agreement_by_severity": [
0.2217,
0.0433,
0.0567
],
"mean_kappa": 0.031726,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"pixelate": {
"agreement_by_severity": [
0.6817,
0.37,
0.6192
],
"mean_kappa": 0.29218,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"quantise": {
"agreement_by_severity": [
0.6825,
0.3783,
0.2117
],
"mean_kappa": 0.249781,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"shot_noise": {
"agreement_by_severity": [
0.0525,
0.0425,
0.0833
],
"mean_kappa": 0.011388,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 46405,
"cases_per_second": 141.279,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007078,
"seconds_total": 328.4635
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.056309,
"best_spec": "template_5",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.131335,
"chance_corrected": 0,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.125855,
"chance_corrected": 0,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.160799,
"chance_corrected": 0.020932,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.114185,
"chance_corrected": 0,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.141908,
"chance_corrected": 0,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.191122,
"chance_corrected": 0.056309,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.076937,
"worst_chance_corrected": 0,
"worst_spec": "template_0"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.15212,
"risk_coverage_auc": 0.118703,
"score": 0,
"selective_acc_at_50": 0.10978,
"selective_acc_at_80": 0.135287
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 7007,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Dermatoscopy",
"n_classes": 7,
"n_test": 2005,
"runtime_s": 333.7,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "dermamnist",
"task_name": "Dermatoscopy (7-class skin lesion)",
"timestamp": "2026-08-01T12:47:13+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 35.25083
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.248174,
"auc": 0.710829,
"balanced_accuracy": 0.21425,
"ci95": [
0.205139,
0.223413
],
"n_test": 8216,
"score": 13.567459
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.141071,
"mean_agreement": 0.671167,
"mean_kappa": 0.606744,
"mean_retention": 0.856766,
"n_cases": 1200,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.5667,
0.3367,
0.2492
],
"mean_kappa": 0.287452,
"mean_retention": 0.37313,
"retention_by_severity": [
0.742,
0.3592,
0.0182
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.6725,
0.5467,
0.445
],
"mean_kappa": 0.456249,
"mean_retention": 0.936995,
"retention_by_severity": [
0.9776,
0.8923,
0.9411
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.8908,
0.7717,
0.6883
],
"mean_kappa": 0.739794,
"mean_retention": 0.975557,
"retention_by_severity": [
1,
1,
0.9267
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.8958,
0.8658,
0.83
],
"mean_kappa": 0.832764,
"mean_retention": 0.998147,
"retention_by_severity": [
0.9944,
1,
1
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.89,
0.7742,
0.6442
],
"mean_kappa": 0.717462,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
}
},
"score": 85.676587,
"worst_retention": 0.018176
},
"calibration": {
"accuracy": 0.248174,
"brier": 1.098689,
"ece": 0.434842,
"mean_confidence": 0.683016,
"overconfidence": 0.434842,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.141071,
"mean_agreement": 0.479087,
"mean_kappa": 0.392224,
"mean_retention": 0.577992,
"n_cases": 1200,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.8383,
0.6767,
0.525
],
"mean_kappa": 0.613603,
"mean_retention": 0.950286,
"retention_by_severity": [
1,
1,
0.8509
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.7867,
0.6692,
0.5383
],
"mean_kappa": 0.596316,
"mean_retention": 0.90793,
"retention_by_severity": [
1,
0.9618,
0.762
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.9267,
0.8333,
0.72
],
"mean_kappa": 0.786754,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.3425,
0.2792,
0.2492
],
"mean_kappa": 0.160556,
"mean_retention": 0.185986,
"retention_by_severity": [
0.3691,
0.1113,
0.0775
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.145,
0.1208,
0.0575
],
"mean_kappa": -0.005763,
"mean_retention": 0,
"retention_by_severity": [
0,
0,
0
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.8175,
0.5058,
0.3342
],
"mean_kappa": 0.46514,
"mean_retention": 0.726558,
"retention_by_severity": [
0.8513,
0.6758,
0.6526
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.2675,
0.2133,
0.2142
],
"mean_kappa": 0.128961,
"mean_retention": 0.275187,
"retention_by_severity": [
0.2991,
0.281,
0.2454
],
"scored": true
}
},
"score": 57.799225,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 52616,
"cases_per_second": 125.858,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007945,
"seconds_total": 418.057
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.149531,
"best_spec": "template_1",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.219156,
"chance_corrected": 0.141071,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.226846,
"chance_corrected": 0.149531,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.1951,
"chance_corrected": 0.11461,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.202518,
"chance_corrected": 0.12277,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.132945,
"chance_corrected": 0.046239,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.202736,
"chance_corrected": 0.123009,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.309227,
"score": 30.922684,
"scored": true,
"spread_balanced_accuracy": 0.093902,
"worst_chance_corrected": 0.046239,
"worst_spec": "template_4"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.021769
},
"uncertainty": {
"full_accuracy": 0.248174,
"risk_coverage_auc": 0.330036,
"score": 26.303931,
"selective_acc_at_50": 0.320107,
"selective_acc_at_80": 0.276586
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 12975,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": true
},
"modality": "CT",
"n_classes": 11,
"n_test": 8216,
"runtime_s": 474.4,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "organcmnist_224",
"task_name": "Abdominal CT organs — coronal (C) @224px",
"timestamp": "2026-08-01T11:34:55+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 23.567278
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.109421,
"auc": 0.593095,
"balanced_accuracy": 0.122905,
"ci95": [
0.115701,
0.129763
],
"n_test": 8216,
"score": 3.519576
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.053159,
"mean_agreement": 0.648389,
"mean_kappa": 0.49044,
"mean_retention": 0.635161,
"n_cases": 1200,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.8258,
0.64,
0.4908
],
"mean_kappa": 0.466594,
"mean_retention": 0.471996,
"retention_by_severity": [
0.5941,
0.3089,
0.5131
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.6025,
0.54,
0.4767
],
"mean_kappa": 0.341071,
"mean_retention": 0.562405,
"retention_by_severity": [
0.7088,
0.5567,
0.4218
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.87,
0.7725,
0.6642
],
"mean_kappa": 0.678969,
"mean_retention": 0.911262,
"retention_by_severity": [
0.9646,
0.821,
0.9482
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.5483,
0.5083,
0.4083
],
"mean_kappa": 0.274315,
"mean_retention": 0.303193,
"retention_by_severity": [
0.2622,
0.2654,
0.3819
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.9017,
0.7967,
0.68
],
"mean_kappa": 0.691251,
"mean_retention": 0.926948,
"retention_by_severity": [
0.9989,
0.8519,
0.93
],
"scored": true
}
},
"score": 63.516083,
"worst_retention": 0.262197
},
"calibration": {
"accuracy": 0.109421,
"brier": 1.342516,
"ece": 0.571727,
"mean_confidence": 0.681148,
"overconfidence": 0.571727,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.053159,
"mean_agreement": 0.58369,
"mean_kappa": 0.423127,
"mean_retention": 0.503033,
"n_cases": 1200,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.8208,
0.7125,
0.6075
],
"mean_kappa": 0.569731,
"mean_retention": 0.83857,
"retention_by_severity": [
0.9595,
0.7314,
0.8248
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.7908,
0.6875,
0.56
],
"mean_kappa": 0.54343,
"mean_retention": 0.634892,
"retention_by_severity": [
0.7388,
0.6601,
0.5058
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.6625,
0.4267,
0.2892
],
"mean_kappa": 0.264692,
"mean_retention": 0.263428,
"retention_by_severity": [
0.422,
0.2374,
0.131
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.7167,
0.59,
0.4083
],
"mean_kappa": 0.365347,
"mean_retention": 0.363153,
"retention_by_severity": [
0.5435,
0.5459,
0
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.4933,
0.1608,
0.2392
],
"mean_kappa": 0.144424,
"mean_retention": 0.308526,
"retention_by_severity": [
0.8442,
0.0813,
0
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9258,
0.8508,
0.6725
],
"mean_kappa": 0.734627,
"mean_retention": 0.855595,
"retention_by_severity": [
0.8532,
0.8045,
0.9091
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.6642,
0.5392,
0.4392
],
"mean_kappa": 0.339636,
"mean_retention": 0.257067,
"retention_by_severity": [
0.3793,
0.3919,
0
],
"scored": true
}
},
"score": 50.303286,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 52616,
"cases_per_second": 129.257,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007736,
"seconds_total": 407.0636
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.059233,
"best_spec": "template_1",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.139235,
"chance_corrected": 0.053159,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.144758,
"chance_corrected": 0.059233,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.125562,
"chance_corrected": 0.038118,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.136821,
"chance_corrected": 0.050503,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.089504,
"chance_corrected": 0,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.121262,
"chance_corrected": 0.033388,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.055254,
"worst_chance_corrected": 0,
"worst_spec": "template_4"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.001361
},
"uncertainty": {
"full_accuracy": 0.109421,
"risk_coverage_auc": 0.103163,
"score": 1.34789,
"selective_acc_at_50": 0.106378,
"selective_acc_at_80": 0.109843
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 12975,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": true
},
"modality": "CT",
"n_classes": 11,
"n_test": 8216,
"runtime_s": 413.3,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "organcmnist",
"task_name": "Abdominal CT organs — coronal (C)",
"timestamp": "2026-08-01T11:42:49+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 36.579413
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.586538,
"auc": 0.854076,
"balanced_accuracy": 0.668376,
"ci95": [
0.643017,
0.690429
],
"n_test": 624,
"score": 33.675214
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.336752,
"mean_agreement": 0.890278,
"mean_kappa": 0.695544,
"mean_retention": 0.752284,
"n_cases": 624,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.9135,
0.8141,
0.5208
],
"mean_kappa": 0.455318,
"mean_retention": 0.610829,
"retention_by_severity": [
0.8934,
0.9391,
0
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.9135,
0.8686,
0.8413
],
"mean_kappa": 0.54853,
"mean_retention": 0.559222,
"retention_by_severity": [
0.7259,
0.533,
0.4188
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.9696,
0.9503,
0.9327
],
"mean_kappa": 0.840049,
"mean_retention": 0.771574,
"retention_by_severity": [
0.8553,
0.764,
0.6954
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.9535,
0.9375,
0.9151
],
"mean_kappa": 0.79387,
"mean_retention": 0.819797,
"retention_by_severity": [
0.8858,
0.8401,
0.7335
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.9776,
0.9471,
0.899
],
"mean_kappa": 0.839953,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
}
},
"score": 75.228426,
"worst_retention": 0
},
"calibration": {
"accuracy": 0.586538,
"brier": 0.73701,
"ece": 0.353211,
"mean_confidence": 0.938758,
"overconfidence": 0.352219,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.336752,
"mean_agreement": 0.776938,
"mean_kappa": 0.467415,
"mean_retention": 0.572758,
"n_cases": 624,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9712,
0.9407,
0.9119
],
"mean_kappa": 0.838116,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.9728,
0.9423,
0.8942
],
"mean_kappa": 0.802186,
"mean_retention": 0.914552,
"retention_by_severity": [
0.9924,
1,
0.7513
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.9663,
0.9183,
0.8718
],
"mean_kappa": 0.722424,
"mean_retention": 0.703892,
"retention_by_severity": [
0.901,
0.7183,
0.4924
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.8333,
0.5481,
0.484
],
"mean_kappa": 0.212179,
"mean_retention": 0.28934,
"retention_by_severity": [
0.868,
0,
0
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.4631,
0.6442,
0.7532
],
"mean_kappa": 0.057084,
"mean_retention": 0.203046,
"retention_by_severity": [
0,
0.4213,
0.1878
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9583,
0.8654,
0.351
],
"mean_kappa": 0.519998,
"mean_retention": 0.836717,
"retention_by_severity": [
0.9797,
0.9188,
0.6117
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.7676,
0.6346,
0.6234
],
"mean_kappa": 0.119921,
"mean_retention": 0.06176,
"retention_by_severity": [
0.1853,
0,
0
],
"scored": true
}
},
"score": 57.275804,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 23088,
"cases_per_second": 119.316,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.008381,
"seconds_total": 193.5038
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.618803,
"best_spec": "template_4",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.668376,
"chance_corrected": 0.336752,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.725214,
"chance_corrected": 0.450427,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.725641,
"chance_corrected": 0.451282,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.551282,
"chance_corrected": 0.102564,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.809402,
"chance_corrected": 0.618803,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.524786,
"chance_corrected": 0.049573,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.08011,
"score": 8.01105,
"scored": true,
"spread_balanced_accuracy": 0.284615,
"worst_chance_corrected": 0.049573,
"worst_spec": "template_5"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.341026
},
"uncertainty": {
"full_accuracy": 0.586538,
"risk_coverage_auc": 0.679618,
"score": 35.923686,
"selective_acc_at_50": 0.695513,
"selective_acc_at_80": 0.617234
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 4708,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": true
},
"modality": "Chest X-ray",
"n_classes": 2,
"n_test": 624,
"runtime_s": 236.4,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "pneumoniamnist_224",
"task_name": "Paediatric chest X-ray (pneumonia) @224px",
"timestamp": "2026-08-01T11:04:31+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.92,
"dimensions_scored": [
"accuracy",
"acquisition",
"calibration",
"corruption",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 34.355577
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.551282,
"auc": 0.716656,
"balanced_accuracy": 0.618803,
"ci95": [
0.589015,
0.648389
],
"n_test": 624,
"score": 23.760684
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0.237607,
"mean_agreement": 0.790278,
"mean_kappa": 0.513554,
"mean_retention": 0.73717,
"n_cases": 624,
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.8301,
0.7244,
0.6186
],
"mean_kappa": 0.222791,
"mean_retention": 0.107914,
"retention_by_severity": [
0.3237,
0,
0
],
"scored": true
},
"field_of_view": {
"agreement_by_severity": [
0.8846,
0.8429,
0.7228
],
"mean_kappa": 0.55515,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"gamma_shift": {
"agreement_by_severity": [
0.9471,
0.9103,
0.8606
],
"mean_kappa": 0.719528,
"mean_retention": 0.681055,
"retention_by_severity": [
0.8094,
0.6763,
0.5576
],
"scored": true
},
"resolution": {
"agreement_by_severity": [
0.7901,
0.6154,
0.484
],
"mean_kappa": 0.343175,
"mean_retention": 0.932854,
"retention_by_severity": [
1,
1,
0.7986
],
"scored": true
},
"window_level": {
"agreement_by_severity": [
0.9535,
0.8974,
0.7724
],
"mean_kappa": 0.727125,
"mean_retention": 0.964029,
"retention_by_severity": [
1,
1,
0.8921
],
"scored": true
}
},
"score": 73.717026,
"worst_retention": 0
},
"calibration": {
"accuracy": 0.551282,
"brier": 0.767054,
"ece": 0.362413,
"mean_confidence": 0.913695,
"overconfidence": 0.362413,
"score": 0
},
"corruption": {
"clean_reference_cc": 0.237607,
"mean_agreement": 0.677121,
"mean_kappa": 0.375029,
"mean_retention": 0.565091,
"n_cases": 624,
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9327,
0.8381,
0.6859
],
"mean_kappa": 0.627531,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"contrast": {
"agreement_by_severity": [
0.9551,
0.9054,
0.7981
],
"mean_kappa": 0.710744,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"defocus_blur": {
"agreement_by_severity": [
0.899,
0.6971,
0.641
],
"mean_kappa": 0.462557,
"mean_retention": 0.666667,
"retention_by_severity": [
1,
1,
0
],
"scored": true
},
"gaussian_noise": {
"agreement_by_severity": [
0.7147,
0.5112,
0.258
],
"mean_kappa": 0.062926,
"mean_retention": 0.020384,
"retention_by_severity": [
0,
0,
0.0612
],
"scored": true
},
"pixelate": {
"agreement_by_severity": [
0.7644,
0.3974,
0.5433
],
"mean_kappa": 0.213795,
"mean_retention": 0.659472,
"retention_by_severity": [
1,
0.9784,
0
],
"scored": true
},
"quantise": {
"agreement_by_severity": [
0.9279,
0.8462,
0.5897
],
"mean_kappa": 0.508689,
"mean_retention": 0.593525,
"retention_by_severity": [
0.8885,
0.8921,
0
],
"scored": true
},
"shot_noise": {
"agreement_by_severity": [
0.6506,
0.399,
0.2644
],
"mean_kappa": 0.038961,
"mean_retention": 0.015588,
"retention_by_severity": [
0,
0,
0.0468
],
"scored": true
}
},
"score": 56.509078,
"worst_retention": 0
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 23088,
"cases_per_second": 85.46,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.011701,
"seconds_total": 270.1603
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.478632,
"best_spec": "template_2",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.618803,
"chance_corrected": 0.237607,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.647009,
"chance_corrected": 0.294017,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.739316,
"chance_corrected": 0.478632,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.601709,
"chance_corrected": 0.203419,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.599145,
"chance_corrected": 0.198291,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.559829,
"chance_corrected": 0.119658,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0.25,
"score": 25,
"scored": true,
"spread_balanced_accuracy": 0.179487,
"worst_chance_corrected": 0.119658,
"worst_spec": "template_5"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.348718
},
"uncertainty": {
"full_accuracy": 0.551282,
"risk_coverage_auc": 0.593662,
"score": 18.732307,
"selective_acc_at_50": 0.576923,
"selective_acc_at_80": 0.57515
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 4708,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Chest X-ray",
"n_classes": 2,
"n_test": 624,
"runtime_s": 275,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "pneumoniamnist",
"task_name": "Paediatric chest X-ray (pneumonia)",
"timestamp": "2026-08-01T11:08:28+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.61,
"dimensions_scored": [
"accuracy",
"calibration",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 9.198908
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.435,
"auc": 0.611249,
"balanced_accuracy": 0.2,
"ci95": [
0.2,
0.2
],
"n_test": 400,
"score": 0
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 400,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
1,
1,
0.9975
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"field_of_view": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gamma_shift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"resolution": {
"agreement_by_severity": [
1,
0.9975,
0.995
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"window_level": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"calibration": {
"accuracy": 0.435,
"brier": 0.758073,
"ece": 0.192176,
"mean_confidence": 0.627176,
"overconfidence": 0.192176,
"score": 3.912164
},
"corruption": {
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 400,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9975,
1,
1
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"contrast": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"defocus_blur": {
"agreement_by_severity": [
1,
0.9975,
0.9975
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gaussian_noise": {
"agreement_by_severity": [
1,
1,
0.995
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"pixelate": {
"agreement_by_severity": [
1,
0.995,
0.985
],
"mean_kappa": 0,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"quantise": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"shot_noise": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": null,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 14800,
"cases_per_second": 128.633,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007774,
"seconds_total": 115.0558
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.068552,
"best_spec": "template_3",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.2,
"chance_corrected": 0,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.18026,
"chance_corrected": 0,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.183333,
"chance_corrected": 0,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.254842,
"chance_corrected": 0.068552,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.238266,
"chance_corrected": 0.047833,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.202941,
"chance_corrected": 0.003676,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.074582,
"worst_chance_corrected": 0,
"worst_spec": "template_0"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.435,
"risk_coverage_auc": 0.53644,
"score": 42.05505,
"selective_acc_at_50": 0.545,
"selective_acc_at_80": 0.496875
}
},
"domain": "medicine",
"kind": "image224",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 1080,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Fundus",
"n_classes": 5,
"n_test": 400,
"runtime_s": 159.4,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "retinamnist_224",
"task_name": "Fundus photography (DR grade) @224px",
"timestamp": "2026-08-01T19:51:41+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.61,
"dimensions_scored": [
"accuracy",
"calibration",
"cost",
"spec_sensitivity",
"subgroup",
"uncertainty"
],
"index": 12.192916
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.41,
"auc": 0.472783,
"balanced_accuracy": 0.194903,
"ci95": [
0.18227,
0.211552
],
"n_test": 400,
"score": 0
},
"acquisition": {
"basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 400,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"detector_noise": {
"agreement_by_severity": [
0.82,
0.6475,
0.8075
],
"mean_kappa": 0.24409,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"field_of_view": {
"agreement_by_severity": [
0.95,
0.94,
0.9375
],
"mean_kappa": 0.231241,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gamma_shift": {
"agreement_by_severity": [
0.9475,
0.945,
0.94
],
"mean_kappa": 0.216361,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"resolution": {
"agreement_by_severity": [
0.9275,
0.9325,
0.9375
],
"mean_kappa": 0.205313,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"window_level": {
"agreement_by_severity": [
0.96,
0.945,
0.9225
],
"mean_kappa": 0.491876,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"calibration": {
"accuracy": 0.41,
"brier": 0.760994,
"ece": 0.149302,
"mean_confidence": 0.459839,
"overconfidence": 0.049839,
"score": 25.348975
},
"corruption": {
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 400,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"brightness": {
"agreement_by_severity": [
0.9625,
0.9375,
0.9075
],
"mean_kappa": 0.52981,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"contrast": {
"agreement_by_severity": [
0.9575,
0.9425,
0.935
],
"mean_kappa": 0.267389,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"defocus_blur": {
"agreement_by_severity": [
0.9475,
0.9375,
0.9375
],
"mean_kappa": 0.144762,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"gaussian_noise": {
"agreement_by_severity": [
0.6425,
0.7875,
0.9025
],
"mean_kappa": 0.17752,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"pixelate": {
"agreement_by_severity": [
0.9225,
0.93,
0.93
],
"mean_kappa": 0.228827,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"quantise": {
"agreement_by_severity": [
0.94,
0.76,
0.7075
],
"mean_kappa": 0.359965,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"shot_noise": {
"agreement_by_severity": [
0.57,
0.675,
0.8325
],
"mean_kappa": 0.181886,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"cost": {
"basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
"cases": 14800,
"cases_per_second": 128.674,
"caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
"device": "mps",
"latency_reference_seconds": 0.02,
"parameters": 195902721,
"parameters_millions": 195.9,
"score": 100,
"scored": true,
"seconds_per_case": 0.007772,
"seconds_total": 115.0195
},
"limited_data": null,
"spec_sensitivity": {
"basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
"best_chance_corrected": 0.010544,
"best_spec": "template_4",
"n_specs": 6,
"per_spec": {
"template_0": {
"balanced_accuracy": 0.194903,
"chance_corrected": 0,
"spec": "this is a photo of {c}"
},
"template_1": {
"balanced_accuracy": 0.180213,
"chance_corrected": 0,
"spec": "an image of {c}"
},
"template_2": {
"balanced_accuracy": 0.206347,
"chance_corrected": 0.007934,
"spec": "a medical image showing {c}"
},
"template_3": {
"balanced_accuracy": 0.187549,
"chance_corrected": 0,
"spec": "{c}"
},
"template_4": {
"balanced_accuracy": 0.208435,
"chance_corrected": 0.010544,
"spec": "histology or clinical image, category: {c}"
},
"template_5": {
"balanced_accuracy": 0.2,
"chance_corrected": 0,
"spec": "the correct label for this image is {c}"
}
},
"retention": 0,
"score": 0,
"scored": true,
"spread_balanced_accuracy": 0.028223,
"worst_chance_corrected": 0,
"worst_spec": "template_0"
},
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.41,
"risk_coverage_auc": 0.38277,
"score": 22.846236,
"selective_acc_at_50": 0.35,
"selective_acc_at_80": 0.421875
}
},
"domain": "medicine",
"kind": "image",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 1080,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Fundus",
"n_classes": 5,
"n_test": 400,
"runtime_s": 118.6,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split",
"task": "retinamnist",
"task_name": "Fundus photography (DR grade)",
"timestamp": "2026-08-01T19:54:21+00:00"
}
],
"heldout": null,
"sources": [
{
"path": "results/heldout/audit_report.json",
"sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
},
{
"path": "results/heldout/cycle.json",
"sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
},
{
"path": "results/heldout/manifests/gen1.seal.json",
"sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
},
{
"path": "results/heldout/manifests/gen2.seal.json",
"sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
},
{
"path": "results/heldout/manifests/public.seal.json",
"sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
},
{
"path": "results/records/bloodmnist_224__biomedclip_zeroshot.json",
"sha256": "28d0f988885cab16dc30d8b4ffc350e9510752d4d372238a42be3cbc4cba804a"
},
{
"path": "results/records/bloodmnist__biomedclip_zeroshot.json",
"sha256": "4437196da05ecf58b4d6f6f43c789ff677550a651d2f303f38e9a0fc81d373c6"
},
{
"path": "results/records/breastmnist_224__biomedclip_zeroshot.json",
"sha256": "d8825e0fbf3c87d21f97dd6075bb8e28057de5d3c86f2145760d80c91bbec33b"
},
{
"path": "results/records/breastmnist__biomedclip_zeroshot.json",
"sha256": "f335b76d2529620a9ebb2d8241df0319a4d20c6ed395cee1e4ed7438b6ff52c9"
},
{
"path": "results/records/dermamnist_224__biomedclip_zeroshot.json",
"sha256": "77781b199642e67b1b7ab3a4711f69a9c9b2c6cde053837af744e6d27cc5f4dd"
},
{
"path": "results/records/dermamnist__biomedclip_zeroshot.json",
"sha256": "5c04e18568422f094de0e013e1dcfe1ae2fa09a43ae043ab08d6cd82fd35a7f6"
},
{
"path": "results/records/organcmnist_224__biomedclip_zeroshot.json",
"sha256": "d7c2ca5a1023bc5e600eb6749e3301b2a85f16694075259fba2472183a77612c"
},
{
"path": "results/records/organcmnist__biomedclip_zeroshot.json",
"sha256": "6a445673b017f9518bb4965470eb87ad70ad665dfea783adf971f0344eb6a00c"
},
{
"path": "results/records/pneumoniamnist_224__biomedclip_zeroshot.json",
"sha256": "d4b8bc2215c65e86612ced45c4f036a2d22914cdfc8e4a050d44622f392e09f8"
},
{
"path": "results/records/pneumoniamnist__biomedclip_zeroshot.json",
"sha256": "df1367e80e2f9d245d2a3ff5ed021883f80fcdf8274a52017f59a874d7433291"
},
{
"path": "results/records/retinamnist_224__biomedclip_zeroshot.json",
"sha256": "7d840843cb52f5667fd0a7c65acc6acf01d24d3cbf07b908d6d9cbec399131aa"
},
{
"path": "results/records/retinamnist__biomedclip_zeroshot.json",
"sha256": "345a60f9f365f8b0871b31fcfc0add607bf81cc636e04d8ef136f10e2a6b134f"
}
]
},
"declared": null,
"declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
"expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
"expiry_utc": "2026-09-20T00:00:00Z",
"issued_utc": "2026-08-03T16:03:41Z",
"issuer": {
"algo": "Ed25519",
"key_id": "ns-passport-2026-07",
"name": "NakedSignal",
"public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
},
"observed": null,
"observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
"passport_version": "0.1",
"signature": "CFcqMgJl8WaqJ1dkUniu6+b13Djb1SWG3tyyRJ7UsOEOTCJ3s5rWG0jhWCwmjgIcaKxyNfvd2PyHB0fCjS6BCA==",
"spec_version": "MedEval-1 v0.1",
"subject": {
"behaviour_version": "1.0.0",
"code_fingerprint": "b444c02112cc",
"code_fingerprints": [
"b444c02112cc"
],
"family": "foundation",
"model_id": "biomedclip_zeroshot",
"model_name": "BiomedCLIP ViT-B/16, zero-shot",
"params": "195.9M frozen; classified by text prompt, no labels used"
}
}Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.