NakedSignal OS · Documentation
Back to the consoleDocumentation

BiomedCLIP ViT-B/16, zero-shot

LIVE

The signed evidence record for biomedclip_zeroshot. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.

SIGNATURE VALIDexit 0signature valid, issued by NakedSignal under key ns-passport-2026-07, not expired
SubjectBiomedCLIP ViT-B/16, zero-shot biomedclip_zeroshot
Configuration195.9M frozen; classified by text prompt, no labels used
Code fingerprintb444c02112cc
Spec versionMedEval-1 v0.1
Issued2026-08-03T16:03:41Z by NakedSignal
Expires2026-09-20T00:00:00Zthe sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first
key ns-passport-2026-07 · Ed25519

The check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.

Measured by NakedSignal

Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 17 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.

TaskDomainn testaccuracycalibrationuncertaintysubgroupcorruptionacquisitionlimited datacostspec sensitivityIndex
bloodmnist_224medicine34211.323337015.473877069.9775470.975895n/a10011.21578929.157445
bloodmnistmedicine34210000not scorednot scoredn/a10004.918033
breastmnist_224medicine15620.05012508.587353079.73214398.833333n/a1002.32558139.120227
breastmnistmedicine1566.015038056.764976079.86111186.944444n/a100036.0029
dermamnist_224medicine200516.413154015.345893055.21250978.320304n/a10028.30467633.35518
dermamnistmedicine20050000not scorednot scoredn/a10004.918033
organcmnist_224medicine821613.567459026.303931057.79922585.676587n/a10030.92268435.25083
organcmnistmedicine82163.51957601.34789050.30328663.516083n/a100023.567278
pneumoniamnist_224medicine62433.675214035.923686057.27580475.228426n/a1008.0110536.579413
pneumoniamnistmedicine62423.760684018.732307056.50907873.717026n/a1002534.355577
retinamnist_224medicine40003.91216442.055050not scorednot scoredn/a10009.198908
retinamnistmedicine400025.34897522.8462360not scorednot scoredn/a100012.192916

Cross-site retention

not measured for this subject — it is only defined within a family of tasks that share a corpus and differ in acquisition, and this model was not run on one.

Private held-out track

this subject was not submitted to a sealed generation.

Known limitations

Subgroup metadataabsent on 12 of 12 corpora, so the subgroup score falls back to the worst diagnostic class. That absence is a disclosure, not a pass.
Training set cappedyes on at least one task — see the raw JSON for which
Observed in deploymentnothing — see the third panel

Audit chain

Bundles re-derived48 of 48
Ledger100 entries, head 6b7b371ff650a8ed, chain intact
Re-run it yourselfpython src/heldout/audit.py 1f66c7719ab3943c6fcc
17 source files, with hashes
results/heldout/audit_report.jsondf460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2
results/heldout/cycle.jsone012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502
results/heldout/manifests/gen1.seal.json91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81
results/heldout/manifests/gen2.seal.jsonf61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547
results/heldout/manifests/public.seal.json1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c
results/records/bloodmnist_224__biomedclip_zeroshot.json28d0f988885cab16dc30d8b4ffc350e9510752d4d372238a42be3cbc4cba804a
results/records/bloodmnist__biomedclip_zeroshot.json4437196da05ecf58b4d6f6f43c789ff677550a651d2f303f38e9a0fc81d373c6
results/records/breastmnist_224__biomedclip_zeroshot.jsond8825e0fbf3c87d21f97dd6075bb8e28057de5d3c86f2145760d80c91bbec33b
results/records/breastmnist__biomedclip_zeroshot.jsonf335b76d2529620a9ebb2d8241df0319a4d20c6ed395cee1e4ed7438b6ff52c9
results/records/dermamnist_224__biomedclip_zeroshot.json77781b199642e67b1b7ab3a4711f69a9c9b2c6cde053837af744e6d27cc5f4dd
results/records/dermamnist__biomedclip_zeroshot.json5c04e18568422f094de0e013e1dcfe1ae2fa09a43ae043ab08d6cd82fd35a7f6
results/records/organcmnist_224__biomedclip_zeroshot.jsond7c2ca5a1023bc5e600eb6749e3301b2a85f16694075259fba2472183a77612c
results/records/organcmnist__biomedclip_zeroshot.json6a445673b017f9518bb4965470eb87ad70ad665dfea783adf971f0344eb6a00c
results/records/pneumoniamnist_224__biomedclip_zeroshot.jsond4b8bc2215c65e86612ced45c4f036a2d22914cdfc8e4a050d44622f392e09f8
results/records/pneumoniamnist__biomedclip_zeroshot.jsondf1367e80e2f9d245d2a3ff5ed021883f80fcdf8274a52017f59a874d7433291
results/records/retinamnist_224__biomedclip_zeroshot.json7d840843cb52f5667fd0a7c65acc6acf01d24d3cbf07b908d6d9cbec399131aa
results/records/retinamnist__biomedclip_zeroshot.json345a60f9f365f8b0871b31fcfc0add607bf81cc636e04d8ef136f10e2a6b134f
Declared by vendor

Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.

No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.

Observed in deployment

Production history and drift — how the model has actually behaved since it was deployed, on real traffic.

Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.

Raw JSON — the whole signed document, 119,932 characters
{
 "canonicalisation": {
  "float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
  "form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
  "signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
 },
 "computed": {
  "audit": {
   "audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
   "bundles": 48,
   "ledger_entries": 100,
   "ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
   "ledger_intact": true,
   "rederived": 48
  },
  "cross_site": null,
  "evaluations": [
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 29.157445
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.197895,
      "auc": 0.634089,
      "balanced_accuracy": 0.136579,
      "ci95": [
       0.131291,
       0.141767
      ],
      "n_test": 3421,
      "score": 1.323337
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.0168,
      "mean_agreement": 0.839056,
      "mean_kappa": 0.575316,
      "mean_retention": 0.709759,
      "n_cases": 1200,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.9133,
         0.7158,
         0.6092
        ],
        "mean_kappa": 0.401075,
        "mean_retention": 0.471296,
        "retention_by_severity": [
         0.8917,
         0.5222,
         0
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.8475,
         0.7525,
         0.6158
        ],
        "mean_kappa": 0.397528,
        "mean_retention": 0.896583,
        "retention_by_severity": [
         0.7074,
         0.9824,
         1
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9558,
         0.91,
         0.8525
        ],
        "mean_kappa": 0.729872,
        "mean_retention": 0.508303,
        "retention_by_severity": [
         0.5512,
         0.4366,
         0.5371
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.9592,
         0.9425,
         0.9283
        ],
        "mean_kappa": 0.834777,
        "mean_retention": 0.834501,
        "retention_by_severity": [
         0.9053,
         0.8236,
         0.7746
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.915,
         0.8567,
         0.8117
        ],
        "mean_kappa": 0.513329,
        "mean_retention": 0.838112,
        "retention_by_severity": [
         0.8451,
         0.7749,
         0.8943
        ],
        "scored": true
       }
      },
      "score": 70.975895,
      "worst_retention": 0
     },
     "calibration": {
      "accuracy": 0.197895,
      "brier": 1.161568,
      "ece": 0.436072,
      "mean_confidence": 0.633967,
      "overconfidence": 0.436072,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.0168,
      "mean_agreement": 0.569167,
      "mean_kappa": 0.361946,
      "mean_retention": 0.699775,
      "n_cases": 1200,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.895,
         0.8308,
         0.7017
        ],
        "mean_kappa": 0.406363,
        "mean_retention": 0.824929,
        "retention_by_severity": [
         0.8779,
         0.5969,
         1
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.9283,
         0.8675,
         0.7783
        ],
        "mean_kappa": 0.635892,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.9692,
         0.9283,
         0.8833
        ],
        "mean_kappa": 0.78103,
        "mean_retention": 0.781059,
        "retention_by_severity": [
         0.8709,
         0.9069,
         0.5654
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.7233,
         0.5258,
         0.0092
        ],
        "mean_kappa": 0.144907,
        "mean_retention": 0.296178,
        "retention_by_severity": [
         0.8885,
         0,
         0
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.0117,
         0.0042,
         0.0017
        ],
        "mean_kappa": 0.002038,
        "mean_retention": 0.901249,
        "retention_by_severity": [
         1,
         0.7037,
         1
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9483,
         0.7975,
         0.4208
        ],
        "mean_kappa": 0.51251,
        "mean_retention": 0.900599,
        "retention_by_severity": [
         0.7018,
         1,
         1
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.6308,
         0.0933,
         0.0033
        ],
        "mean_kappa": 0.050882,
        "mean_retention": 0.194414,
        "retention_by_severity": [
         0.5832,
         0,
         0
        ],
        "scored": true
       }
      },
      "score": 69.97754,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 47821,
      "cases_per_second": 121.784,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.008211,
      "seconds_total": 392.6718
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.060714,
      "best_spec": "template_2",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.1397,
        "chance_corrected": 0.0168,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.136323,
        "chance_corrected": 0.012941,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.178125,
        "chance_corrected": 0.060714,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.173234,
        "chance_corrected": 0.055125,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.130958,
        "chance_corrected": 0.00681,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.133846,
        "chance_corrected": 0.01011,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.112158,
      "score": 11.215789,
      "scored": true,
      "spread_balanced_accuracy": 0.047167,
      "worst_chance_corrected": 0.00681,
      "worst_spec": "template_4"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.197895,
      "risk_coverage_auc": 0.260396,
      "score": 15.473877,
      "selective_acc_at_50": 0.294152,
      "selective_acc_at_80": 0.229814
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 11959,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": true
    },
    "modality": "Microscopy",
    "n_classes": 8,
    "n_test": 3421,
    "runtime_s": 546.9,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "bloodmnist_224",
    "task_name": "Peripheral blood cells (8-class) @224px",
    "timestamp": "2026-08-01T19:59:26+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.61,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 4.918033
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.126863,
      "auc": 0.590911,
      "balanced_accuracy": 0.114617,
      "ci95": [
       0.104209,
       0.125475
      ],
      "n_test": 3421,
      "score": 0
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 1200,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.6025,
         0.2975,
         0.2367
        ],
        "mean_kappa": 0.16666,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.6508,
         0.5458,
         0.4083
        ],
        "mean_kappa": 0.345899,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.8108,
         0.6667,
         0.58
        ],
        "mean_kappa": 0.515929,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "resolution": {
        "agreement_by_severity": [
         0.57,
         0.4742,
         0.4092
        ],
        "mean_kappa": 0.277632,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "window_level": {
        "agreement_by_severity": [
         0.6733,
         0.4892,
         0.4417
        ],
        "mean_kappa": 0.325299,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "calibration": {
      "accuracy": 0.126863,
      "brier": 1.127525,
      "ece": 0.385935,
      "mean_confidence": 0.512798,
      "overconfidence": 0.385935,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 1200,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.5783,
         0.4267,
         0.41
        ],
        "mean_kappa": 0.232991,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "contrast": {
        "agreement_by_severity": [
         0.8008,
         0.6208,
         0.3642
        ],
        "mean_kappa": 0.399405,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.6458,
         0.4508,
         0.2
        ],
        "mean_kappa": 0.247683,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.315,
         0.2383,
         0.2358
        ],
        "mean_kappa": 0.031664,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "pixelate": {
        "agreement_by_severity": [
         0.4833,
         0.0483,
         0.1883
        ],
        "mean_kappa": 0.118947,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "quantise": {
        "agreement_by_severity": [
         0.805,
         0.5842,
         0.3892
        ],
        "mean_kappa": 0.435747,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.2383,
         0.2358,
         0.2342
        ],
        "mean_kappa": 0.002387,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 47821,
      "cases_per_second": 121.905,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.008203,
      "seconds_total": 392.2794
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.087689,
      "best_spec": "template_1",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.104782,
        "chance_corrected": 0,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.201727,
        "chance_corrected": 0.087689,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.133236,
        "chance_corrected": 0.009412,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.115581,
        "chance_corrected": 0,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.143801,
        "chance_corrected": 0.021486,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.130824,
        "chance_corrected": 0.006655,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.096946,
      "worst_chance_corrected": 0,
      "worst_spec": "template_0"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.126863,
      "risk_coverage_auc": 0.121242,
      "score": 0,
      "selective_acc_at_50": 0.122807,
      "selective_acc_at_80": 0.124954
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 11959,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Microscopy",
    "n_classes": 8,
    "n_test": 3421,
    "runtime_s": 401.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "bloodmnist",
    "task_name": "Peripheral blood cells (8-class)",
    "timestamp": "2026-08-01T20:08:33+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 39.120227
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.448718,
      "auc": 0.767753,
      "balanced_accuracy": 0.600251,
      "ci95": [
       0.546294,
       0.653504
      ],
      "n_test": 156,
      "score": 20.050125
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.200501,
      "mean_agreement": 0.869658,
      "mean_kappa": 0.684586,
      "mean_retention": 0.988333,
      "n_cases": 156,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.9423,
         0.8077,
         0.5962
        ],
        "mean_kappa": 0.549982,
        "mean_retention": 0.995833,
        "retention_by_severity": [
         1,
         0.9875,
         1
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.9103,
         0.8718,
         0.8782
        ],
        "mean_kappa": 0.686845,
        "mean_retention": 0.96875,
        "retention_by_severity": [
         1,
         1,
         0.9062
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9679,
         0.9423,
         0.8974
        ],
        "mean_kappa": 0.821871,
        "mean_retention": 0.989583,
        "retention_by_severity": [
         1,
         0.9688,
         1
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.9423,
         0.9038,
         0.8654
        ],
        "mean_kappa": 0.737433,
        "mean_retention": 0.9875,
        "retention_by_severity": [
         0.9813,
         0.9813,
         1
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.9423,
         0.8526,
         0.7244
        ],
        "mean_kappa": 0.626802,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 98.833333,
      "worst_retention": 0.90625
     },
     "calibration": {
      "accuracy": 0.448718,
      "brier": 0.909597,
      "ece": 0.446057,
      "mean_confidence": 0.894775,
      "overconfidence": 0.446057,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.200501,
      "mean_agreement": 0.812882,
      "mean_kappa": 0.516569,
      "mean_retention": 0.797321,
      "n_cases": 156,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9487,
         0.8974,
         0.8526
        ],
        "mean_kappa": 0.748554,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.9808,
         0.9359,
         0.859
        ],
        "mean_kappa": 0.787331,
        "mean_retention": 0.747917,
        "retention_by_severity": [
         0.9562,
         0.6,
         0.6875
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.9487,
         0.8077,
         0.5449
        ],
        "mean_kappa": 0.550654,
        "mean_retention": 0.95,
        "retention_by_severity": [
         0.85,
         1,
         1
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.8526,
         0.5705,
         0.5321
        ],
        "mean_kappa": 0.340724,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.7885,
         0.7756,
         0.7756
        ],
        "mean_kappa": 0.006677,
        "mean_retention": 0.014583,
        "retention_by_severity": [
         0.0438,
         0,
         0
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9744,
         0.9359,
         0.8526
        ],
        "mean_kappa": 0.755937,
        "mean_retention": 0.86875,
        "retention_by_severity": [
         1,
         1,
         0.6062
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.8141,
         0.6795,
         0.7436
        ],
        "mean_kappa": 0.42611,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 79.732143,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 5772,
      "cases_per_second": 116.604,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.008576,
      "seconds_total": 49.5008
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.377193,
      "best_spec": "template_5",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.600251,
        "chance_corrected": 0.200501,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.567043,
        "chance_corrected": 0.134085,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.530702,
        "chance_corrected": 0.061404,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.504386,
        "chance_corrected": 0.008772,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.530702,
        "chance_corrected": 0.061404,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.688596,
        "chance_corrected": 0.377193,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.023256,
      "score": 2.325581,
      "scored": true,
      "spread_balanced_accuracy": 0.184211,
      "worst_chance_corrected": 0.008772,
      "worst_spec": "template_3"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.27193
     },
     "uncertainty": {
      "full_accuracy": 0.448718,
      "risk_coverage_auc": 0.542937,
      "score": 8.587353,
      "selective_acc_at_50": 0.525641,
      "selective_acc_at_80": 0.456
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 546,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Ultrasound",
    "n_classes": 2,
    "n_test": 156,
    "runtime_s": 57.8,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "breastmnist_224",
    "task_name": "Breast ultrasound (malignancy) @224px",
    "timestamp": "2026-08-01T19:45:20+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 36.0029
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.730769,
      "auc": 0.599415,
      "balanced_accuracy": 0.530075,
      "ci95": [
       0.490243,
       0.581737
      ],
      "n_test": 156,
      "score": 6.015038
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.06015,
      "mean_agreement": 0.75812,
      "mean_kappa": 0.327646,
      "mean_retention": 0.869444,
      "n_cases": 156,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.7885,
         0.4359,
         0.1154
        ],
        "mean_kappa": 0.101765,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.8269,
         0.7821,
         0.7244
        ],
        "mean_kappa": 0.235987,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9744,
         0.9487,
         0.9103
        ],
        "mean_kappa": 0.570658,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.7115,
         0.5833,
         0.6987
        ],
        "mean_kappa": 0.145787,
        "mean_retention": 0.527778,
        "retention_by_severity": [
         1,
         0.4792,
         0.1042
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.9744,
         0.9615,
         0.9359
        ],
        "mean_kappa": 0.584031,
        "mean_retention": 0.819444,
        "retention_by_severity": [
         1,
         1,
         0.4583
        ],
        "scored": true
       }
      },
      "score": 86.944444,
      "worst_retention": 0.104167
     },
     "calibration": {
      "accuracy": 0.730769,
      "brier": 0.463927,
      "ece": 0.202305,
      "mean_confidence": 0.922621,
      "overconfidence": 0.191852,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.06015,
      "mean_agreement": 0.618437,
      "mean_kappa": 0.266298,
      "mean_retention": 0.798611,
      "n_cases": 156,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9808,
         0.9744,
         0.9231
        ],
        "mean_kappa": 0.664735,
        "mean_retention": 0.430556,
        "retention_by_severity": [
         1,
         0.1667,
         0.125
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.9679,
         0.8974,
         0.7051
        ],
        "mean_kappa": 0.451446,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.8782,
         0.6987,
         0.8718
        ],
        "mean_kappa": 0.225214,
        "mean_retention": 0.777778,
        "retention_by_severity": [
         1,
         1,
         0.3333
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.5962,
         0.1667,
         0.0641
        ],
        "mean_kappa": 0.045655,
        "mean_retention": 0.763889,
        "retention_by_severity": [
         1,
         1,
         0.2917
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.7179,
         0.2564,
         0.3205
        ],
        "mean_kappa": 0.078122,
        "mean_retention": 0.902778,
        "retention_by_severity": [
         0.8333,
         0.875,
         1
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9487,
         0.8846,
         0.5833
        ],
        "mean_kappa": 0.37938,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.3718,
         0.1218,
         0.0577
        ],
        "mean_kappa": 0.019531,
        "mean_retention": 0.715278,
        "retention_by_severity": [
         1,
         1,
         0.1458
        ],
        "scored": true
       }
      },
      "score": 79.861111,
      "worst_retention": 0.125
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 5772,
      "cases_per_second": 127.544,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.00784,
      "seconds_total": 45.2551
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.154135,
      "best_spec": "template_1",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.530075,
        "chance_corrected": 0.06015,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.577068,
        "chance_corrected": 0.154135,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.511905,
        "chance_corrected": 0.02381,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.494987,
        "chance_corrected": 0,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.541353,
        "chance_corrected": 0.082707,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.495614,
        "chance_corrected": 0,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.08208,
      "worst_chance_corrected": 0,
      "worst_spec": "template_3"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.095238
     },
     "uncertainty": {
      "full_accuracy": 0.730769,
      "risk_coverage_auc": 0.783825,
      "score": 56.764976,
      "selective_acc_at_50": 0.782051,
      "selective_acc_at_80": 0.76
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 546,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Ultrasound",
    "n_classes": 2,
    "n_test": 156,
    "runtime_s": 46.7,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "breastmnist",
    "task_name": "Breast ultrasound (malignancy)",
    "timestamp": "2026-08-01T19:46:18+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 33.35518
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.238903,
      "auc": 0.690303,
      "balanced_accuracy": 0.283541,
      "ci95": [
       0.24635,
       0.32303
      ],
      "n_test": 2005,
      "score": 16.413154
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.169943,
      "mean_agreement": 0.754611,
      "mean_kappa": 0.614709,
      "mean_retention": 0.783203,
      "n_cases": 1200,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.8383,
         0.6733,
         0.5083
        ],
        "mean_kappa": 0.488848,
        "mean_retention": 0.698826,
        "retention_by_severity": [
         0.8172,
         0.6722,
         0.6071
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.8008,
         0.7642,
         0.6867
        ],
        "mean_kappa": 0.619821,
        "mean_retention": 0.713073,
        "retention_by_severity": [
         0.8542,
         0.7365,
         0.5485
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.8933,
         0.8125,
         0.7533
        ],
        "mean_kappa": 0.723431,
        "mean_retention": 0.94245,
        "retention_by_severity": [
         0.9281,
         0.9507,
         0.9485
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.8083,
         0.7683,
         0.7075
        ],
        "mean_kappa": 0.629649,
        "mean_retention": 0.761998,
        "retention_by_severity": [
         0.74,
         0.8071,
         0.739
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.8933,
         0.7583,
         0.6525
        ],
        "mean_kappa": 0.611797,
        "mean_retention": 0.799668,
        "retention_by_severity": [
         0.8564,
         0.7803,
         0.7623
        ],
        "scored": true
       }
      },
      "score": 78.320304,
      "worst_retention": 0.548496
     },
     "calibration": {
      "accuracy": 0.238903,
      "brier": 0.983004,
      "ece": 0.357605,
      "mean_confidence": 0.594691,
      "overconfidence": 0.355788,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.169943,
      "mean_agreement": 0.539683,
      "mean_kappa": 0.376942,
      "mean_retention": 0.552125,
      "n_cases": 1200,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.8433,
         0.695,
         0.575
        ],
        "mean_kappa": 0.582437,
        "mean_retention": 0.724649,
        "retention_by_severity": [
         1,
         0.7884,
         0.3855
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.6608,
         0.3967,
         0.2408
        ],
        "mean_kappa": 0.318502,
        "mean_retention": 0.607752,
        "retention_by_severity": [
         0.956,
         0.6719,
         0.1953
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.845,
         0.7208,
         0.5475
        ],
        "mean_kappa": 0.568819,
        "mean_retention": 0.670431,
        "retention_by_severity": [
         0.8218,
         0.7007,
         0.4888
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.6792,
         0.5133,
         0.2733
        ],
        "mean_kappa": 0.285153,
        "mean_retention": 0.640432,
        "retention_by_severity": [
         0.8007,
         0.7812,
         0.3394
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.3992,
         0.3767,
         0.2642
        ],
        "mean_kappa": 0.12458,
        "mean_retention": 0.064961,
        "retention_by_severity": [
         0.0315,
         0.1634,
         0
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.925,
         0.72,
         0.4917
        ],
        "mean_kappa": 0.546417,
        "mean_retention": 0.662379,
        "retention_by_severity": [
         0.9481,
         0.7344,
         0.3047
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.555,
         0.3642,
         0.2467
        ],
        "mean_kappa": 0.212686,
        "mean_retention": 0.494272,
        "retention_by_severity": [
         0.8083,
         0.4029,
         0.2717
        ],
        "scored": true
       }
      },
      "score": 55.212509,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 46405,
      "cases_per_second": 141.256,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007079,
      "seconds_total": 328.5159
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.169943,
      "best_spec": "template_0",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.288523,
        "chance_corrected": 0.169943,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.184087,
        "chance_corrected": 0.048102,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.274595,
        "chance_corrected": 0.153694,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.211492,
        "chance_corrected": 0.080074,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.190909,
        "chance_corrected": 0.056061,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.206552,
        "chance_corrected": 0.074311,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.283047,
      "score": 28.304676,
      "scored": true,
      "spread_balanced_accuracy": 0.104435,
      "worst_chance_corrected": 0.048102,
      "worst_spec": "template_1"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.137931
     },
     "uncertainty": {
      "full_accuracy": 0.238903,
      "risk_coverage_auc": 0.274393,
      "score": 15.345893,
      "selective_acc_at_50": 0.264471,
      "selective_acc_at_80": 0.253741
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 7007,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": true
    },
    "modality": "Dermatoscopy",
    "n_classes": 7,
    "n_test": 2005,
    "runtime_s": 455.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "dermamnist_224",
    "task_name": "Dermatoscopy (7-class skin lesion) @224px",
    "timestamp": "2026-08-01T12:39:38+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.61,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 4.918033
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.15212,
      "auc": 0.556266,
      "balanced_accuracy": 0.125952,
      "ci95": [
       0.107291,
       0.150679
      ],
      "n_test": 2005,
      "score": 0
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 1200,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.4225,
         0.22,
         0.0608
        ],
        "mean_kappa": 0.091706,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.8225,
         0.7717,
         0.6683
        ],
        "mean_kappa": 0.556302,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9,
         0.8067,
         0.7283
        ],
        "mean_kappa": 0.662291,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "resolution": {
        "agreement_by_severity": [
         0.8033,
         0.7725,
         0.715
        ],
        "mean_kappa": 0.55929,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "window_level": {
        "agreement_by_severity": [
         0.8792,
         0.7658,
         0.5917
        ],
        "mean_kappa": 0.573185,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "calibration": {
      "accuracy": 0.15212,
      "brier": 1.336268,
      "ece": 0.608845,
      "mean_confidence": 0.760008,
      "overconfidence": 0.607888,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 1200,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.8817,
         0.7625,
         0.5758
        ],
        "mean_kappa": 0.544322,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "contrast": {
        "agreement_by_severity": [
         0.7925,
         0.68,
         0.6075
        ],
        "mean_kappa": 0.39402,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.8267,
         0.7125,
         0.6367
        ],
        "mean_kappa": 0.466489,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.2217,
         0.0433,
         0.0567
        ],
        "mean_kappa": 0.031726,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "pixelate": {
        "agreement_by_severity": [
         0.6817,
         0.37,
         0.6192
        ],
        "mean_kappa": 0.29218,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "quantise": {
        "agreement_by_severity": [
         0.6825,
         0.3783,
         0.2117
        ],
        "mean_kappa": 0.249781,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.0525,
         0.0425,
         0.0833
        ],
        "mean_kappa": 0.011388,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 46405,
      "cases_per_second": 141.279,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007078,
      "seconds_total": 328.4635
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.056309,
      "best_spec": "template_5",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.131335,
        "chance_corrected": 0,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.125855,
        "chance_corrected": 0,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.160799,
        "chance_corrected": 0.020932,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.114185,
        "chance_corrected": 0,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.141908,
        "chance_corrected": 0,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.191122,
        "chance_corrected": 0.056309,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.076937,
      "worst_chance_corrected": 0,
      "worst_spec": "template_0"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.15212,
      "risk_coverage_auc": 0.118703,
      "score": 0,
      "selective_acc_at_50": 0.10978,
      "selective_acc_at_80": 0.135287
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 7007,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Dermatoscopy",
    "n_classes": 7,
    "n_test": 2005,
    "runtime_s": 333.7,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "dermamnist",
    "task_name": "Dermatoscopy (7-class skin lesion)",
    "timestamp": "2026-08-01T12:47:13+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 35.25083
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.248174,
      "auc": 0.710829,
      "balanced_accuracy": 0.21425,
      "ci95": [
       0.205139,
       0.223413
      ],
      "n_test": 8216,
      "score": 13.567459
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.141071,
      "mean_agreement": 0.671167,
      "mean_kappa": 0.606744,
      "mean_retention": 0.856766,
      "n_cases": 1200,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.5667,
         0.3367,
         0.2492
        ],
        "mean_kappa": 0.287452,
        "mean_retention": 0.37313,
        "retention_by_severity": [
         0.742,
         0.3592,
         0.0182
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.6725,
         0.5467,
         0.445
        ],
        "mean_kappa": 0.456249,
        "mean_retention": 0.936995,
        "retention_by_severity": [
         0.9776,
         0.8923,
         0.9411
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.8908,
         0.7717,
         0.6883
        ],
        "mean_kappa": 0.739794,
        "mean_retention": 0.975557,
        "retention_by_severity": [
         1,
         1,
         0.9267
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.8958,
         0.8658,
         0.83
        ],
        "mean_kappa": 0.832764,
        "mean_retention": 0.998147,
        "retention_by_severity": [
         0.9944,
         1,
         1
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.89,
         0.7742,
         0.6442
        ],
        "mean_kappa": 0.717462,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 85.676587,
      "worst_retention": 0.018176
     },
     "calibration": {
      "accuracy": 0.248174,
      "brier": 1.098689,
      "ece": 0.434842,
      "mean_confidence": 0.683016,
      "overconfidence": 0.434842,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.141071,
      "mean_agreement": 0.479087,
      "mean_kappa": 0.392224,
      "mean_retention": 0.577992,
      "n_cases": 1200,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.8383,
         0.6767,
         0.525
        ],
        "mean_kappa": 0.613603,
        "mean_retention": 0.950286,
        "retention_by_severity": [
         1,
         1,
         0.8509
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.7867,
         0.6692,
         0.5383
        ],
        "mean_kappa": 0.596316,
        "mean_retention": 0.90793,
        "retention_by_severity": [
         1,
         0.9618,
         0.762
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.9267,
         0.8333,
         0.72
        ],
        "mean_kappa": 0.786754,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.3425,
         0.2792,
         0.2492
        ],
        "mean_kappa": 0.160556,
        "mean_retention": 0.185986,
        "retention_by_severity": [
         0.3691,
         0.1113,
         0.0775
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.145,
         0.1208,
         0.0575
        ],
        "mean_kappa": -0.005763,
        "mean_retention": 0,
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.8175,
         0.5058,
         0.3342
        ],
        "mean_kappa": 0.46514,
        "mean_retention": 0.726558,
        "retention_by_severity": [
         0.8513,
         0.6758,
         0.6526
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.2675,
         0.2133,
         0.2142
        ],
        "mean_kappa": 0.128961,
        "mean_retention": 0.275187,
        "retention_by_severity": [
         0.2991,
         0.281,
         0.2454
        ],
        "scored": true
       }
      },
      "score": 57.799225,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 52616,
      "cases_per_second": 125.858,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007945,
      "seconds_total": 418.057
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.149531,
      "best_spec": "template_1",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.219156,
        "chance_corrected": 0.141071,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.226846,
        "chance_corrected": 0.149531,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.1951,
        "chance_corrected": 0.11461,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.202518,
        "chance_corrected": 0.12277,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.132945,
        "chance_corrected": 0.046239,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.202736,
        "chance_corrected": 0.123009,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.309227,
      "score": 30.922684,
      "scored": true,
      "spread_balanced_accuracy": 0.093902,
      "worst_chance_corrected": 0.046239,
      "worst_spec": "template_4"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.021769
     },
     "uncertainty": {
      "full_accuracy": 0.248174,
      "risk_coverage_auc": 0.330036,
      "score": 26.303931,
      "selective_acc_at_50": 0.320107,
      "selective_acc_at_80": 0.276586
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 12975,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": true
    },
    "modality": "CT",
    "n_classes": 11,
    "n_test": 8216,
    "runtime_s": 474.4,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "organcmnist_224",
    "task_name": "Abdominal CT organs — coronal (C) @224px",
    "timestamp": "2026-08-01T11:34:55+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 23.567278
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.109421,
      "auc": 0.593095,
      "balanced_accuracy": 0.122905,
      "ci95": [
       0.115701,
       0.129763
      ],
      "n_test": 8216,
      "score": 3.519576
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.053159,
      "mean_agreement": 0.648389,
      "mean_kappa": 0.49044,
      "mean_retention": 0.635161,
      "n_cases": 1200,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.8258,
         0.64,
         0.4908
        ],
        "mean_kappa": 0.466594,
        "mean_retention": 0.471996,
        "retention_by_severity": [
         0.5941,
         0.3089,
         0.5131
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.6025,
         0.54,
         0.4767
        ],
        "mean_kappa": 0.341071,
        "mean_retention": 0.562405,
        "retention_by_severity": [
         0.7088,
         0.5567,
         0.4218
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.87,
         0.7725,
         0.6642
        ],
        "mean_kappa": 0.678969,
        "mean_retention": 0.911262,
        "retention_by_severity": [
         0.9646,
         0.821,
         0.9482
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.5483,
         0.5083,
         0.4083
        ],
        "mean_kappa": 0.274315,
        "mean_retention": 0.303193,
        "retention_by_severity": [
         0.2622,
         0.2654,
         0.3819
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.9017,
         0.7967,
         0.68
        ],
        "mean_kappa": 0.691251,
        "mean_retention": 0.926948,
        "retention_by_severity": [
         0.9989,
         0.8519,
         0.93
        ],
        "scored": true
       }
      },
      "score": 63.516083,
      "worst_retention": 0.262197
     },
     "calibration": {
      "accuracy": 0.109421,
      "brier": 1.342516,
      "ece": 0.571727,
      "mean_confidence": 0.681148,
      "overconfidence": 0.571727,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.053159,
      "mean_agreement": 0.58369,
      "mean_kappa": 0.423127,
      "mean_retention": 0.503033,
      "n_cases": 1200,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.8208,
         0.7125,
         0.6075
        ],
        "mean_kappa": 0.569731,
        "mean_retention": 0.83857,
        "retention_by_severity": [
         0.9595,
         0.7314,
         0.8248
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.7908,
         0.6875,
         0.56
        ],
        "mean_kappa": 0.54343,
        "mean_retention": 0.634892,
        "retention_by_severity": [
         0.7388,
         0.6601,
         0.5058
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.6625,
         0.4267,
         0.2892
        ],
        "mean_kappa": 0.264692,
        "mean_retention": 0.263428,
        "retention_by_severity": [
         0.422,
         0.2374,
         0.131
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.7167,
         0.59,
         0.4083
        ],
        "mean_kappa": 0.365347,
        "mean_retention": 0.363153,
        "retention_by_severity": [
         0.5435,
         0.5459,
         0
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.4933,
         0.1608,
         0.2392
        ],
        "mean_kappa": 0.144424,
        "mean_retention": 0.308526,
        "retention_by_severity": [
         0.8442,
         0.0813,
         0
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9258,
         0.8508,
         0.6725
        ],
        "mean_kappa": 0.734627,
        "mean_retention": 0.855595,
        "retention_by_severity": [
         0.8532,
         0.8045,
         0.9091
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.6642,
         0.5392,
         0.4392
        ],
        "mean_kappa": 0.339636,
        "mean_retention": 0.257067,
        "retention_by_severity": [
         0.3793,
         0.3919,
         0
        ],
        "scored": true
       }
      },
      "score": 50.303286,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 52616,
      "cases_per_second": 129.257,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007736,
      "seconds_total": 407.0636
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.059233,
      "best_spec": "template_1",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.139235,
        "chance_corrected": 0.053159,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.144758,
        "chance_corrected": 0.059233,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.125562,
        "chance_corrected": 0.038118,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.136821,
        "chance_corrected": 0.050503,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.089504,
        "chance_corrected": 0,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.121262,
        "chance_corrected": 0.033388,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.055254,
      "worst_chance_corrected": 0,
      "worst_spec": "template_4"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.001361
     },
     "uncertainty": {
      "full_accuracy": 0.109421,
      "risk_coverage_auc": 0.103163,
      "score": 1.34789,
      "selective_acc_at_50": 0.106378,
      "selective_acc_at_80": 0.109843
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 12975,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": true
    },
    "modality": "CT",
    "n_classes": 11,
    "n_test": 8216,
    "runtime_s": 413.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "organcmnist",
    "task_name": "Abdominal CT organs — coronal (C)",
    "timestamp": "2026-08-01T11:42:49+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 36.579413
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.586538,
      "auc": 0.854076,
      "balanced_accuracy": 0.668376,
      "ci95": [
       0.643017,
       0.690429
      ],
      "n_test": 624,
      "score": 33.675214
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.336752,
      "mean_agreement": 0.890278,
      "mean_kappa": 0.695544,
      "mean_retention": 0.752284,
      "n_cases": 624,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.9135,
         0.8141,
         0.5208
        ],
        "mean_kappa": 0.455318,
        "mean_retention": 0.610829,
        "retention_by_severity": [
         0.8934,
         0.9391,
         0
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.9135,
         0.8686,
         0.8413
        ],
        "mean_kappa": 0.54853,
        "mean_retention": 0.559222,
        "retention_by_severity": [
         0.7259,
         0.533,
         0.4188
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9696,
         0.9503,
         0.9327
        ],
        "mean_kappa": 0.840049,
        "mean_retention": 0.771574,
        "retention_by_severity": [
         0.8553,
         0.764,
         0.6954
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.9535,
         0.9375,
         0.9151
        ],
        "mean_kappa": 0.79387,
        "mean_retention": 0.819797,
        "retention_by_severity": [
         0.8858,
         0.8401,
         0.7335
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.9776,
         0.9471,
         0.899
        ],
        "mean_kappa": 0.839953,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 75.228426,
      "worst_retention": 0
     },
     "calibration": {
      "accuracy": 0.586538,
      "brier": 0.73701,
      "ece": 0.353211,
      "mean_confidence": 0.938758,
      "overconfidence": 0.352219,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.336752,
      "mean_agreement": 0.776938,
      "mean_kappa": 0.467415,
      "mean_retention": 0.572758,
      "n_cases": 624,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9712,
         0.9407,
         0.9119
        ],
        "mean_kappa": 0.838116,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.9728,
         0.9423,
         0.8942
        ],
        "mean_kappa": 0.802186,
        "mean_retention": 0.914552,
        "retention_by_severity": [
         0.9924,
         1,
         0.7513
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.9663,
         0.9183,
         0.8718
        ],
        "mean_kappa": 0.722424,
        "mean_retention": 0.703892,
        "retention_by_severity": [
         0.901,
         0.7183,
         0.4924
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.8333,
         0.5481,
         0.484
        ],
        "mean_kappa": 0.212179,
        "mean_retention": 0.28934,
        "retention_by_severity": [
         0.868,
         0,
         0
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.4631,
         0.6442,
         0.7532
        ],
        "mean_kappa": 0.057084,
        "mean_retention": 0.203046,
        "retention_by_severity": [
         0,
         0.4213,
         0.1878
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9583,
         0.8654,
         0.351
        ],
        "mean_kappa": 0.519998,
        "mean_retention": 0.836717,
        "retention_by_severity": [
         0.9797,
         0.9188,
         0.6117
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.7676,
         0.6346,
         0.6234
        ],
        "mean_kappa": 0.119921,
        "mean_retention": 0.06176,
        "retention_by_severity": [
         0.1853,
         0,
         0
        ],
        "scored": true
       }
      },
      "score": 57.275804,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 23088,
      "cases_per_second": 119.316,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.008381,
      "seconds_total": 193.5038
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.618803,
      "best_spec": "template_4",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.668376,
        "chance_corrected": 0.336752,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.725214,
        "chance_corrected": 0.450427,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.725641,
        "chance_corrected": 0.451282,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.551282,
        "chance_corrected": 0.102564,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.809402,
        "chance_corrected": 0.618803,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.524786,
        "chance_corrected": 0.049573,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.08011,
      "score": 8.01105,
      "scored": true,
      "spread_balanced_accuracy": 0.284615,
      "worst_chance_corrected": 0.049573,
      "worst_spec": "template_5"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.341026
     },
     "uncertainty": {
      "full_accuracy": 0.586538,
      "risk_coverage_auc": 0.679618,
      "score": 35.923686,
      "selective_acc_at_50": 0.695513,
      "selective_acc_at_80": 0.617234
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 4708,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": true
    },
    "modality": "Chest X-ray",
    "n_classes": 2,
    "n_test": 624,
    "runtime_s": 236.4,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "pneumoniamnist_224",
    "task_name": "Paediatric chest X-ray (pneumonia) @224px",
    "timestamp": "2026-08-01T11:04:31+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.92,
     "dimensions_scored": [
      "accuracy",
      "acquisition",
      "calibration",
      "corruption",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 34.355577
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.551282,
      "auc": 0.716656,
      "balanced_accuracy": 0.618803,
      "ci95": [
       0.589015,
       0.648389
      ],
      "n_test": 624,
      "score": 23.760684
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0.237607,
      "mean_agreement": 0.790278,
      "mean_kappa": 0.513554,
      "mean_retention": 0.73717,
      "n_cases": 624,
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.8301,
         0.7244,
         0.6186
        ],
        "mean_kappa": 0.222791,
        "mean_retention": 0.107914,
        "retention_by_severity": [
         0.3237,
         0,
         0
        ],
        "scored": true
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.8846,
         0.8429,
         0.7228
        ],
        "mean_kappa": 0.55515,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9471,
         0.9103,
         0.8606
        ],
        "mean_kappa": 0.719528,
        "mean_retention": 0.681055,
        "retention_by_severity": [
         0.8094,
         0.6763,
         0.5576
        ],
        "scored": true
       },
       "resolution": {
        "agreement_by_severity": [
         0.7901,
         0.6154,
         0.484
        ],
        "mean_kappa": 0.343175,
        "mean_retention": 0.932854,
        "retention_by_severity": [
         1,
         1,
         0.7986
        ],
        "scored": true
       },
       "window_level": {
        "agreement_by_severity": [
         0.9535,
         0.8974,
         0.7724
        ],
        "mean_kappa": 0.727125,
        "mean_retention": 0.964029,
        "retention_by_severity": [
         1,
         1,
         0.8921
        ],
        "scored": true
       }
      },
      "score": 73.717026,
      "worst_retention": 0
     },
     "calibration": {
      "accuracy": 0.551282,
      "brier": 0.767054,
      "ece": 0.362413,
      "mean_confidence": 0.913695,
      "overconfidence": 0.362413,
      "score": 0
     },
     "corruption": {
      "clean_reference_cc": 0.237607,
      "mean_agreement": 0.677121,
      "mean_kappa": 0.375029,
      "mean_retention": 0.565091,
      "n_cases": 624,
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9327,
         0.8381,
         0.6859
        ],
        "mean_kappa": 0.627531,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "contrast": {
        "agreement_by_severity": [
         0.9551,
         0.9054,
         0.7981
        ],
        "mean_kappa": 0.710744,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.899,
         0.6971,
         0.641
        ],
        "mean_kappa": 0.462557,
        "mean_retention": 0.666667,
        "retention_by_severity": [
         1,
         1,
         0
        ],
        "scored": true
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.7147,
         0.5112,
         0.258
        ],
        "mean_kappa": 0.062926,
        "mean_retention": 0.020384,
        "retention_by_severity": [
         0,
         0,
         0.0612
        ],
        "scored": true
       },
       "pixelate": {
        "agreement_by_severity": [
         0.7644,
         0.3974,
         0.5433
        ],
        "mean_kappa": 0.213795,
        "mean_retention": 0.659472,
        "retention_by_severity": [
         1,
         0.9784,
         0
        ],
        "scored": true
       },
       "quantise": {
        "agreement_by_severity": [
         0.9279,
         0.8462,
         0.5897
        ],
        "mean_kappa": 0.508689,
        "mean_retention": 0.593525,
        "retention_by_severity": [
         0.8885,
         0.8921,
         0
        ],
        "scored": true
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.6506,
         0.399,
         0.2644
        ],
        "mean_kappa": 0.038961,
        "mean_retention": 0.015588,
        "retention_by_severity": [
         0,
         0,
         0.0468
        ],
        "scored": true
       }
      },
      "score": 56.509078,
      "worst_retention": 0
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 23088,
      "cases_per_second": 85.46,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.011701,
      "seconds_total": 270.1603
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.478632,
      "best_spec": "template_2",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.618803,
        "chance_corrected": 0.237607,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.647009,
        "chance_corrected": 0.294017,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.739316,
        "chance_corrected": 0.478632,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.601709,
        "chance_corrected": 0.203419,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.599145,
        "chance_corrected": 0.198291,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.559829,
        "chance_corrected": 0.119658,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0.25,
      "score": 25,
      "scored": true,
      "spread_balanced_accuracy": 0.179487,
      "worst_chance_corrected": 0.119658,
      "worst_spec": "template_5"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.348718
     },
     "uncertainty": {
      "full_accuracy": 0.551282,
      "risk_coverage_auc": 0.593662,
      "score": 18.732307,
      "selective_acc_at_50": 0.576923,
      "selective_acc_at_80": 0.57515
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 4708,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Chest X-ray",
    "n_classes": 2,
    "n_test": 624,
    "runtime_s": 275,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "pneumoniamnist",
    "task_name": "Paediatric chest X-ray (pneumonia)",
    "timestamp": "2026-08-01T11:08:28+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.61,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 9.198908
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.435,
      "auc": 0.611249,
      "balanced_accuracy": 0.2,
      "ci95": [
       0.2,
       0.2
      ],
      "n_test": 400,
      "score": 0
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 400,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         1,
         1,
         0.9975
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_of_view": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gamma_shift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "resolution": {
        "agreement_by_severity": [
         1,
         0.9975,
         0.995
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "window_level": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "calibration": {
      "accuracy": 0.435,
      "brier": 0.758073,
      "ece": 0.192176,
      "mean_confidence": 0.627176,
      "overconfidence": 0.192176,
      "score": 3.912164
     },
     "corruption": {
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 400,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9975,
         1,
         1
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "contrast": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "defocus_blur": {
        "agreement_by_severity": [
         1,
         0.9975,
         0.9975
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         1,
         1,
         0.995
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "pixelate": {
        "agreement_by_severity": [
         1,
         0.995,
         0.985
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "quantise": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "shot_noise": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": null,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 14800,
      "cases_per_second": 128.633,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007774,
      "seconds_total": 115.0558
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.068552,
      "best_spec": "template_3",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.2,
        "chance_corrected": 0,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.18026,
        "chance_corrected": 0,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.183333,
        "chance_corrected": 0,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.254842,
        "chance_corrected": 0.068552,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.238266,
        "chance_corrected": 0.047833,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.202941,
        "chance_corrected": 0.003676,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.074582,
      "worst_chance_corrected": 0,
      "worst_spec": "template_0"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.435,
      "risk_coverage_auc": 0.53644,
      "score": 42.05505,
      "selective_acc_at_50": 0.545,
      "selective_acc_at_80": 0.496875
     }
    },
    "domain": "medicine",
    "kind": "image224",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 1080,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Fundus",
    "n_classes": 5,
    "n_test": 400,
    "runtime_s": 159.4,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "retinamnist_224",
    "task_name": "Fundus photography (DR grade) @224px",
    "timestamp": "2026-08-01T19:51:41+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.61,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "cost",
      "spec_sensitivity",
      "subgroup",
      "uncertainty"
     ],
     "index": 12.192916
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.41,
      "auc": 0.472783,
      "balanced_accuracy": 0.194903,
      "ci95": [
       0.18227,
       0.211552
      ],
      "n_test": 400,
      "score": 0
     },
     "acquisition": {
      "basis": "simulated re-acquisition (gamma, window/level, resolution, FOV, detector noise)",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 400,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "detector_noise": {
        "agreement_by_severity": [
         0.82,
         0.6475,
         0.8075
        ],
        "mean_kappa": 0.24409,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_of_view": {
        "agreement_by_severity": [
         0.95,
         0.94,
         0.9375
        ],
        "mean_kappa": 0.231241,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gamma_shift": {
        "agreement_by_severity": [
         0.9475,
         0.945,
         0.94
        ],
        "mean_kappa": 0.216361,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "resolution": {
        "agreement_by_severity": [
         0.9275,
         0.9325,
         0.9375
        ],
        "mean_kappa": 0.205313,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "window_level": {
        "agreement_by_severity": [
         0.96,
         0.945,
         0.9225
        ],
        "mean_kappa": 0.491876,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "calibration": {
      "accuracy": 0.41,
      "brier": 0.760994,
      "ece": 0.149302,
      "mean_confidence": 0.459839,
      "overconfidence": 0.049839,
      "score": 25.348975
     },
     "corruption": {
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 400,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "brightness": {
        "agreement_by_severity": [
         0.9625,
         0.9375,
         0.9075
        ],
        "mean_kappa": 0.52981,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "contrast": {
        "agreement_by_severity": [
         0.9575,
         0.9425,
         0.935
        ],
        "mean_kappa": 0.267389,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "defocus_blur": {
        "agreement_by_severity": [
         0.9475,
         0.9375,
         0.9375
        ],
        "mean_kappa": 0.144762,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "gaussian_noise": {
        "agreement_by_severity": [
         0.6425,
         0.7875,
         0.9025
        ],
        "mean_kappa": 0.17752,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "pixelate": {
        "agreement_by_severity": [
         0.9225,
         0.93,
         0.93
        ],
        "mean_kappa": 0.228827,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "quantise": {
        "agreement_by_severity": [
         0.94,
         0.76,
         0.7075
        ],
        "mean_kappa": 0.359965,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "shot_noise": {
        "agreement_by_severity": [
         0.57,
         0.675,
         0.8325
        ],
        "mean_kappa": 0.181886,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": {
      "basis": "measured wall-clock inference over the evaluated cases on mps; score = 2*ref/(ref+measured), capped at 100",
      "cases": 14800,
      "cases_per_second": 128.674,
      "caveat": "hardware-dependent. This number must never be quoted without the device it was measured on.",
      "device": "mps",
      "latency_reference_seconds": 0.02,
      "parameters": 195902721,
      "parameters_millions": 195.9,
      "score": 100,
      "scored": true,
      "seconds_per_case": 0.007772,
      "seconds_total": 115.0195
     },
     "limited_data": null,
     "spec_sensitivity": {
      "basis": "worst / best chance-corrected balanced accuracy across faithful paraphrases of the task specification",
      "best_chance_corrected": 0.010544,
      "best_spec": "template_4",
      "n_specs": 6,
      "per_spec": {
       "template_0": {
        "balanced_accuracy": 0.194903,
        "chance_corrected": 0,
        "spec": "this is a photo of {c}"
       },
       "template_1": {
        "balanced_accuracy": 0.180213,
        "chance_corrected": 0,
        "spec": "an image of {c}"
       },
       "template_2": {
        "balanced_accuracy": 0.206347,
        "chance_corrected": 0.007934,
        "spec": "a medical image showing {c}"
       },
       "template_3": {
        "balanced_accuracy": 0.187549,
        "chance_corrected": 0,
        "spec": "{c}"
       },
       "template_4": {
        "balanced_accuracy": 0.208435,
        "chance_corrected": 0.010544,
        "spec": "histology or clinical image, category: {c}"
       },
       "template_5": {
        "balanced_accuracy": 0.2,
        "chance_corrected": 0,
        "spec": "the correct label for this image is {c}"
       }
      },
      "retention": 0,
      "score": 0,
      "scored": true,
      "spread_balanced_accuracy": 0.028223,
      "worst_chance_corrected": 0,
      "worst_spec": "template_0"
     },
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.41,
      "risk_coverage_auc": 0.38277,
      "score": 22.846236,
      "selective_acc_at_50": 0.35,
      "selective_acc_at_80": 0.421875
     }
    },
    "domain": "medicine",
    "kind": "image",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 1080,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Fundus",
    "n_classes": 5,
    "n_test": 400,
    "runtime_s": 118.6,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split",
    "task": "retinamnist",
    "task_name": "Fundus photography (DR grade)",
    "timestamp": "2026-08-01T19:54:21+00:00"
   }
  ],
  "heldout": null,
  "sources": [
   {
    "path": "results/heldout/audit_report.json",
    "sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
   },
   {
    "path": "results/heldout/cycle.json",
    "sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
   },
   {
    "path": "results/heldout/manifests/gen1.seal.json",
    "sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
   },
   {
    "path": "results/heldout/manifests/gen2.seal.json",
    "sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
   },
   {
    "path": "results/heldout/manifests/public.seal.json",
    "sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
   },
   {
    "path": "results/records/bloodmnist_224__biomedclip_zeroshot.json",
    "sha256": "28d0f988885cab16dc30d8b4ffc350e9510752d4d372238a42be3cbc4cba804a"
   },
   {
    "path": "results/records/bloodmnist__biomedclip_zeroshot.json",
    "sha256": "4437196da05ecf58b4d6f6f43c789ff677550a651d2f303f38e9a0fc81d373c6"
   },
   {
    "path": "results/records/breastmnist_224__biomedclip_zeroshot.json",
    "sha256": "d8825e0fbf3c87d21f97dd6075bb8e28057de5d3c86f2145760d80c91bbec33b"
   },
   {
    "path": "results/records/breastmnist__biomedclip_zeroshot.json",
    "sha256": "f335b76d2529620a9ebb2d8241df0319a4d20c6ed395cee1e4ed7438b6ff52c9"
   },
   {
    "path": "results/records/dermamnist_224__biomedclip_zeroshot.json",
    "sha256": "77781b199642e67b1b7ab3a4711f69a9c9b2c6cde053837af744e6d27cc5f4dd"
   },
   {
    "path": "results/records/dermamnist__biomedclip_zeroshot.json",
    "sha256": "5c04e18568422f094de0e013e1dcfe1ae2fa09a43ae043ab08d6cd82fd35a7f6"
   },
   {
    "path": "results/records/organcmnist_224__biomedclip_zeroshot.json",
    "sha256": "d7c2ca5a1023bc5e600eb6749e3301b2a85f16694075259fba2472183a77612c"
   },
   {
    "path": "results/records/organcmnist__biomedclip_zeroshot.json",
    "sha256": "6a445673b017f9518bb4965470eb87ad70ad665dfea783adf971f0344eb6a00c"
   },
   {
    "path": "results/records/pneumoniamnist_224__biomedclip_zeroshot.json",
    "sha256": "d4b8bc2215c65e86612ced45c4f036a2d22914cdfc8e4a050d44622f392e09f8"
   },
   {
    "path": "results/records/pneumoniamnist__biomedclip_zeroshot.json",
    "sha256": "df1367e80e2f9d245d2a3ff5ed021883f80fcdf8274a52017f59a874d7433291"
   },
   {
    "path": "results/records/retinamnist_224__biomedclip_zeroshot.json",
    "sha256": "7d840843cb52f5667fd0a7c65acc6acf01d24d3cbf07b908d6d9cbec399131aa"
   },
   {
    "path": "results/records/retinamnist__biomedclip_zeroshot.json",
    "sha256": "345a60f9f365f8b0871b31fcfc0add607bf81cc636e04d8ef136f10e2a6b134f"
   }
  ]
 },
 "declared": null,
 "declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
 "expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
 "expiry_utc": "2026-09-20T00:00:00Z",
 "issued_utc": "2026-08-03T16:03:41Z",
 "issuer": {
  "algo": "Ed25519",
  "key_id": "ns-passport-2026-07",
  "name": "NakedSignal",
  "public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
 },
 "observed": null,
 "observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
 "passport_version": "0.1",
 "signature": "CFcqMgJl8WaqJ1dkUniu6+b13Djb1SWG3tyyRJ7UsOEOTCJ3s5rWG0jhWCwmjgIcaKxyNfvd2PyHB0fCjS6BCA==",
 "spec_version": "MedEval-1 v0.1",
 "subject": {
  "behaviour_version": "1.0.0",
  "code_fingerprint": "b444c02112cc",
  "code_fingerprints": [
   "b444c02112cc"
  ],
  "family": "foundation",
  "model_id": "biomedclip_zeroshot",
  "model_name": "BiomedCLIP ViT-B/16, zero-shot",
  "params": "195.9M frozen; classified by text prompt, no labels used"
 }
}

Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.