NakedSignal OS · Documentation
Back to the consoleDocumentation

MLP (128, 64)

LIVE

The signed evidence record for tab_mlp. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.

SIGNATURE VALIDexit 0signature valid, issued by NakedSignal under key ns-passport-2026-07, not expired
SubjectMLP (128, 64) tab_mlp
Configuration2 hidden layers, adam, early stopping when n>=200
Code fingerprintb444c02112cc, de32bca72d56
Spec versionMedEval-1 v0.1
Issued2026-08-03T16:03:42Z by NakedSignal
Expires2026-09-20T00:00:00Zthe sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first
key ns-passport-2026-07 · Ed25519

The check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.

Measured by NakedSignal

Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 21 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.

TaskDomainn testaccuracycalibrationuncertaintysubgroupcorruptionacquisitionlimited datacostspec sensitivityIndex
adult_incomelabour1628150.79159892.28124990.65191524.88721889.681639n/a87.690003n/an/a69.114807
anuran_mfccecology99576.90221811.67689369.89762582.62711365.852553n/a84.968828n/an/a64.181861
breast_wdbcmedicine14389.454927098.48191481.13207585.425618n/a97.015858n/an/a72.830892
coil2000_caravanfinancial40000.68084789.17093493.521817076.932367n/a100n/an/a48.873097
diabetes_pimamedicine19338.36768565.12619869.9444421.81818292.123769n/a97.993208n/an/a60.171001
dry_beanagriculture340493.00789893.27798498.85258588.66970278.097657n/a94.413882n/an/a90.117024
german_creditfinancial2503.04761962.41703455.65103087.5n/a100n/an/a44.190614
heart_clevelandmedicine7554.6428577.49345777.74609344.41379390.994916n/a94.989107n/an/a57.864832
landsat_statlogearth200086.10956289.53842397.62866367.01421878.443291n/a84.820301n/an/a83.242553
occupancy_detection_2energy975275.14942764.09254196.43752474.13372481.23729n/a70.532463n/an/a75.16107
occupancy_detectionenergy266595.14321983.68466698.63432890.90372183.740308n/a76.671821n/an/a88.458322
secomindustrial392024.60150883.9769990not scoredn/a0n/an/a12.753528
steel_plate_faultsindustrial48669.1455774.52196685.86989930.58155868.183834n/a91.400901n/an/a68.207417
wine_quality_redagriculture40052.07014473.49821267.88867143.92523474.640548n/a54.337547n/an/a60.546995
wine_quality_whiteagriculture122545.66811391.72752480.26915712.19512260.334034n/a61.17737n/an/a56.313177

Cross-site retention · occupancy

retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column

Sitesoccupancy_detection, occupancy_detection_2
Mean retention0.894928
Worst retention0.789856

Private held-out track

this subject was not submitted to a sealed generation.

Known limitations

Subgroup metadataabsent on 9 of 15 corpora, so the subgroup score falls back to the worst diagnostic class. That absence is a disclosure, not a pass.
Training set cappedyes on at least one task — see the raw JSON for which
Observed in deploymentnothing — see the third panel

Audit chain

Bundles re-derived48 of 48
Ledger100 entries, head 6b7b371ff650a8ed, chain intact
Re-run it yourselfpython src/heldout/audit.py 1f66c7719ab3943c6fcc
21 source files, with hashes
results/cross_site/occupancy__tab_mlp.json96185a3fc9e50e26176f58216d6d522c5e40aeab923d6ad774380e4c9f9ae0b4
results/heldout/audit_report.jsondf460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2
results/heldout/cycle.jsone012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502
results/heldout/manifests/gen1.seal.json91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81
results/heldout/manifests/gen2.seal.jsonf61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547
results/heldout/manifests/public.seal.json1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c
results/records/adult_income__tab_mlp.json49049d8cc6e2d1511d0385ccb566ad799ad906a00a22d961d04563d6a089eb2e
results/records/anuran_mfcc__tab_mlp.jsonc10bdf3bcd87fd9e4ab7431167b8442d5c53313bb1e06b3b22994604aa89e1b3
results/records/breast_wdbc__tab_mlp.json7bc5146466b300da8ebd3a3e208e1ac0b75c0f39c6575f172e17e7670bf3cfbc
results/records/coil2000_caravan__tab_mlp.json7dafdcc9eb9a9b07415bd6c683b526722fc1ac0044f69e998c963c012888ef3d
results/records/diabetes_pima__tab_mlp.jsonb2936c4291f7fc6111c6d2d58cf696b02438981314fbb8fdff3bc65bfa8161ca
results/records/dry_bean__tab_mlp.json62a6714e9e0ad9464a8fd2f4eb464569af90b093536fa102a5a2de15595241e6
results/records/german_credit__tab_mlp.json8b03a5007c340ee097c252031d573b153b2455f39a79695fb26160d28bdbbeca
results/records/heart_cleveland__tab_mlp.jsonf90552415bce174bc8f36a2cdad2cea905172b9519dbdc6f0ba2587719a17769
results/records/landsat_statlog__tab_mlp.jsone670f180d3cf429e24816c312e75eb254daa225344297c41c4ab2885b618e9ce
results/records/occupancy_detection_2__tab_mlp.jsonf7cbc3ac657e97ef2b920b9c17df8fdc128e06a7d45a46a9ef0d58e5469c2466
results/records/occupancy_detection__tab_mlp.json05e8a020e37e594519c3a6143e63cc85d2a07aff8f411c22288addd891e8b8e9
results/records/secom__tab_mlp.jsonbea110d82ce415d7813b396932e2cea6856de9aca96fc0f0a5f3063e430f5637
results/records/steel_plate_faults__tab_mlp.json51d1d0ee9be444d388aa1d69a1ff2e2345c7d52fd83120214425e9961b53d0f9
results/records/wine_quality_red__tab_mlp.json71bc7b14871de8292e51da89ce19fb9197099a0f24b71c766008c2e12744547b
results/records/wine_quality_white__tab_mlp.json02a436ca6e8edabe3ac34b39a230eb63b12a450238e64ec9597c55caf58c9ddc
Declared by vendor

Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.

No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.

Observed in deployment

Production history and drift — how the model has actually behaved since it was deployed, on real traffic.

Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.

Raw JSON — the whole signed document, 76,566 characters
{
 "canonicalisation": {
  "float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
  "form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
  "signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
 },
 "computed": {
  "audit": {
   "audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
   "bundles": 48,
   "ledger_entries": 100,
   "ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
   "ledger_intact": true,
   "rederived": 48
  },
  "cross_site": {
   "code_fingerprint": "b444c02112cc",
   "device": "mps",
   "family": "occupancy",
   "matrix": {
    "occupancy_detection": {
     "occupancy_detection": {
      "balanced_accuracy": 0.975716,
      "chance_corrected": 0.951432,
      "ece": 0.032631
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.875747,
      "chance_corrected": 0.751494,
      "ece": 0.071815,
      "retention": 0.789856
     }
    },
    "occupancy_detection_2": {
     "occupancy_detection": {
      "balanced_accuracy": 0.975716,
      "chance_corrected": 0.951432,
      "ece": 0.032631,
      "retention": 1
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.875747,
      "chance_corrected": 0.751494,
      "ece": 0.071815
     }
    }
   },
   "mean_retention": 0.894928,
   "note": "retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column",
   "score": 89.492792,
   "site_labels": {
    "occupancy_detection": "test period 1",
    "occupancy_detection_2": "test period 2"
   },
   "sites": [
    "occupancy_detection",
    "occupancy_detection_2"
   ],
   "spec_version": "MedEval-1 v0.1",
   "worst_retention": 0.789856
  },
  "evaluations": [
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 69.114807
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.846508,
      "auc": 0.897352,
      "balanced_accuracy": 0.753958,
      "ci95": [
       0.747389,
       0.763205
      ],
      "n_test": 16281,
      "score": 50.791598
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.846508,
      "brier": 0.210854,
      "ece": 0.015438,
      "mean_confidence": 0.860154,
      "overconfidence": 0.013646,
      "score": 92.281249
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.524625,
      "mean_agreement": 0.95287,
      "mean_kappa": 0.831963,
      "mean_retention": 0.896816,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9983,
         0.99,
         0.975
        ],
        "mean_kappa": 0.96203,
        "mean_retention": 0.990257,
        "retention_by_severity": [
         0.9913,
         1,
         0.9795
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9617,
         0.9242,
         0.8392
        ],
        "mean_kappa": 0.634163,
        "mean_retention": 0.700192,
        "retention_by_severity": [
         0.9423,
         0.8454,
         0.3129
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9975,
         0.9942,
         0.8958
        ],
        "mean_kappa": 0.899697,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 89.681639,
      "worst_retention": 0.312909
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.507916,
      "mean_retention": 0.8769,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.531275,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.458308,
        "n_labels": 40,
        "retention": 0.902331
       },
       "n5": {
        "chance_corrected": 0.369951,
        "n_labels": 10,
        "retention": 0.72837
       }
      },
      "score": 87.690003
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.018759,
        "per_group": {
         "45-59": 0.754722,
         "60+": 0.735963,
         "<45": 0.742274
        },
        "worst": 0.735963
       },
       "race": {
        "gap": 0.131625,
        "per_group": {
         "Amer-Indian-Eskimo": 0.624436,
         "Asian-Pac-Islander": 0.71473,
         "Black": 0.726215,
         "Other": 0.646364,
         "White": 0.756061
        },
        "worst": 0.624436
       },
       "sex": {
        "gap": 0.03144,
        "per_group": {
         "Female": 0.717946,
         "Male": 0.749386
        },
        "worst": 0.717946
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 24.887218,
      "worst_class_recall": 0.578523
     },
     "uncertainty": {
      "full_accuracy": 0.846508,
      "risk_coverage_auc": 0.95326,
      "score": 90.651915,
      "selective_acc_at_50": 0.976044,
      "selective_acc_at_80": 0.911631
     }
    },
    "domain": "labour",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 27676,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": true
    },
    "modality": "Tabular / socioeconomic",
    "n_classes": 2,
    "n_test": 16281,
    "runtime_s": 0.6,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (UCI adult.test scored whole); validation carved from adult.data only (stratified 15%, seed 20260727)",
    "task": "adult_income",
    "task_name": "Census income (>$50K)",
    "timestamp": "2026-08-01T19:16:13+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 64.181861
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.663317,
      "auc": 0.954618,
      "balanced_accuracy": 0.79212,
      "ci95": [
       0.750681,
       0.817312
      ],
      "n_test": 995,
      "score": 76.902218
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.663317,
      "brier": 0.528144,
      "ece": 0.176646,
      "mean_confidence": 0.839836,
      "overconfidence": 0.176519,
      "score": 11.676893
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.769022,
      "mean_agreement": 0.625796,
      "mean_kappa": 0.583308,
      "mean_retention": 0.658526,
      "n_cases": 995,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9628,
         0.8683,
         0.7276
        ],
        "mean_kappa": 0.831591,
        "mean_retention": 0.897583,
        "retention_by_severity": [
         0.9831,
         0.9435,
         0.7661
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.6794,
         0.4503,
         0.3457
        ],
        "mean_kappa": 0.427991,
        "mean_retention": 0.532515,
        "retention_by_severity": [
         0.8741,
         0.4562,
         0.2673
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.8744,
         0.5628,
         0.1608
        ],
        "mean_kappa": 0.490342,
        "mean_retention": 0.545478,
        "retention_by_severity": [
         0.9642,
         0.5315,
         0.1407
        ],
        "scored": true
       }
      },
      "score": 65.852553,
      "worst_retention": 0.140685
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.769022,
      "mean_retention": 0.849688,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.82135,
        "n_labels": 794,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.767379,
        "n_labels": 200,
        "retention": 0.997863
       },
       "n5": {
        "chance_corrected": 0.423887,
        "n_labels": 50,
        "retention": 0.551202
       }
      },
      "score": 84.968828
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "family": {
        "gap": 0.039283,
        "per_group": {
         "Hylidae": 0.843644,
         "Leptodactylidae": 0.882927
        },
        "worst": 0.843644
       },
       "genus": {
        "gap": null,
        "per_group": {
         "Hypsiboas": 0.914138
        },
        "worst": null
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 82.627113,
      "worst_class_recall": 0.29653
     },
     "uncertainty": {
      "full_accuracy": 0.663317,
      "risk_coverage_auc": 0.729079,
      "score": 69.897625,
      "selective_acc_at_50": 0.795181,
      "selective_acc_at_80": 0.721106
     }
    },
    "domain": "ecology",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 5073,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Bioacoustic / MFCC",
    "n_classes": 10,
    "n_test": 995,
    "runtime_s": 0.5,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "no official split; minted at the standard seed (20260727) by partitioning the 60 RecordID recording groups 60/15/15/25 and taking whole recordings, so no recording spans two splits. A row-level random split would place near-duplicate syllables from one recording on both sides and score memorisation (~99%); class balance is therefore uncontrolled across splits and is not stratified",
    "task": "anuran_mfcc",
    "task_name": "Anuran call species (10-class)",
    "timestamp": "2026-08-01T19:15:16+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 72.830892
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.958042,
      "auc": 0.98218,
      "balanced_accuracy": 0.947275,
      "ci95": [
       0.904427,
       0.980862
      ],
      "n_test": 143,
      "score": 89.454927
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.958042,
      "brier": 0.191103,
      "ece": 0.234659,
      "mean_confidence": 0.723383,
      "overconfidence": -0.234659,
      "score": 0
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.894549,
      "mean_agreement": 0.887335,
      "mean_kappa": 0.7976,
      "mean_retention": 0.854256,
      "n_cases": 143,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.965,
         0.8322
        ],
        "mean_kappa": 0.86261,
        "mean_retention": 0.942426,
        "retention_by_severity": [
         1,
         1,
         0.8273
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.993,
         0.986,
         0.9301
        ],
        "mean_kappa": 0.929948,
        "mean_retention": 0.937974,
        "retention_by_severity": [
         1,
         0.9913,
         0.8226
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.986,
         0.8811,
         0.4126
        ],
        "mean_kappa": 0.600241,
        "mean_retention": 0.682369,
        "retention_by_severity": [
         1,
         0.9229,
         0.1242
        ],
        "scored": true
       }
      },
      "score": 85.425618,
      "worst_retention": 0.124209
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.894549,
      "mean_retention": 0.970159,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.902306,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.814465,
        "n_labels": 40,
        "retention": 0.910476
       },
       "n5": {
        "chance_corrected": 0.95891,
        "n_labels": 10,
        "retention": 1
       }
      },
      "score": 97.015858
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 81.132075,
      "worst_class_recall": 0.90566
     },
     "uncertainty": {
      "full_accuracy": 0.958042,
      "risk_coverage_auc": 0.99241,
      "score": 98.481914,
      "selective_acc_at_50": 0.986111,
      "selective_acc_at_80": 0.991228
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 341,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 2,
    "n_test": 143,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "breast_wdbc",
    "task_name": "Breast cytology (WDBC)",
    "timestamp": "2026-08-01T19:13:39+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 48.873097
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.9395,
      "auc": 0.692535,
      "balanced_accuracy": 0.503404,
      "ci95": [
       0.498931,
       0.510293
      ],
      "n_test": 4000,
      "score": 0.680847
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.9395,
      "brier": 0.109898,
      "ece": 0.021658,
      "mean_confidence": 0.944567,
      "overconfidence": 0.005067,
      "score": 89.170934
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.023262,
      "mean_agreement": 0.996574,
      "mean_kappa": 0.700986,
      "mean_retention": 0.769324,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.9967,
         0.9958
        ],
        "mean_kappa": 0.836971,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9967,
         0.9925,
         0.9942
        ],
        "mean_kappa": 0.424505,
        "mean_retention": 0.333333,
        "retention_by_severity": [
         1,
         0,
         0
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9983,
         0.9983,
         0.9967
        ],
        "mean_kappa": 0.841481,
        "mean_retention": 0.974638,
        "retention_by_severity": [
         1,
         0.9239,
         1
        ],
        "scored": true
       }
      },
      "score": 76.932367,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.006808,
      "mean_retention": 1,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.281638,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.02087,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.134798,
        "n_labels": 10,
        "retention": 1
       }
      },
      "score": 100
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.008403
     },
     "uncertainty": {
      "full_accuracy": 0.9395,
      "risk_coverage_auc": 0.967609,
      "score": 93.521817,
      "selective_acc_at_50": 0.9695,
      "selective_acc_at_80": 0.95625
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 4948,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / insurance",
    "n_classes": 2,
    "n_test": 4000,
    "runtime_s": 0.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (CoIL Challenge 2000: ticdata2000.txt 5822 train / ticeval2000.txt 4000 test). The test labels were released separately, in tictgts2000.txt, only after the challenge closed -- they were not available to anyone tuning on this benchmark at the time. Test scored whole; validation carved from ticdata2000.txt only (stratified 15%, seed 20260727).",
    "task": "coil2000_caravan",
    "task_name": "Caravan insurance purchase (CoIL 2000)",
    "timestamp": "2026-08-01T19:15:26+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 60.171001
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.725389,
      "auc": 0.787965,
      "balanced_accuracy": 0.691838,
      "ci95": [
       0.606644,
       0.76225
      ],
      "n_test": 193,
      "score": 38.367685
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.725389,
      "brier": 0.351749,
      "ece": 0.069748,
      "mean_confidence": 0.745541,
      "overconfidence": 0.020153,
      "score": 65.126198
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.383677,
      "mean_agreement": 0.913644,
      "mean_kappa": 0.831919,
      "mean_retention": 0.921238,
      "n_cases": 193,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9948,
         0.9741,
         0.9326
        ],
        "mean_kappa": 0.926233,
        "mean_retention": 0.933519,
        "retention_by_severity": [
         0.9611,
         0.9197,
         0.9197
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9482,
         0.943,
         0.9223
        ],
        "mean_kappa": 0.858832,
        "mean_retention": 0.941237,
        "retention_by_severity": [
         0.8444,
         1,
         0.9793
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9896,
         0.9482,
         0.5699
        ],
        "mean_kappa": 0.710693,
        "mean_retention": 0.888957,
        "retention_by_severity": [
         1,
         0.9537,
         0.7132
        ],
        "scored": true
       }
      },
      "score": 92.123769,
      "worst_retention": 0.713183
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.383677,
      "mean_retention": 0.979932,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.362118,
        "n_labels": 200,
        "retention": 0.94381
       },
       "n20": {
        "chance_corrected": 0.392916,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.382137,
        "n_labels": 10,
        "retention": 0.995986
       }
      },
      "score": 97.993208
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.081818,
        "per_group": {
         "45-59": 0.609091,
         "<45": 0.690909
        },
        "worst": 0.609091
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 21.818182,
      "worst_class_recall": 0.58209
     },
     "uncertainty": {
      "full_accuracy": 0.725389,
      "risk_coverage_auc": 0.849722,
      "score": 69.94444,
      "selective_acc_at_50": 0.875,
      "selective_acc_at_80": 0.766234
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 460,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 193,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "diabetes_pima",
    "task_name": "Diabetes onset (Pima)",
    "timestamp": "2026-08-01T19:13:37+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 90.117024
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.929495,
      "auc": 0.995736,
      "balanced_accuracy": 0.940068,
      "ci95": [
       0.933663,
       0.947357
      ],
      "n_test": 3404,
      "score": 93.007898
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.929495,
      "brier": 0.10343,
      "ece": 0.013444,
      "mean_confidence": 0.917904,
      "overconfidence": -0.011591,
      "score": 93.277984
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.925667,
      "mean_agreement": 0.790093,
      "mean_kappa": 0.749546,
      "mean_retention": 0.780977,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9825,
         0.915,
         0.7858
        ],
        "mean_kappa": 0.872274,
        "mean_retention": 0.906527,
        "retention_by_severity": [
         0.9848,
         0.9315,
         0.8033
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9392,
         0.845,
         0.5783
        ],
        "mean_kappa": 0.738591,
        "mean_retention": 0.741521,
        "retention_by_severity": [
         0.9496,
         0.8687,
         0.4063
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9808,
         0.8333,
         0.2508
        ],
        "mean_kappa": 0.637772,
        "mean_retention": 0.694882,
        "retention_by_severity": [
         1,
         0.91,
         0.1747
        ],
        "scored": true
       }
      },
      "score": 78.097657,
      "worst_retention": 0.174653
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.930079,
      "mean_retention": 0.944139,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.889257,
        "n_labels": 700,
        "retention": 0.956109
       },
       "n20": {
        "chance_corrected": 0.90131,
        "n_labels": 140,
        "retention": 0.969069
       },
       "n5": {
        "chance_corrected": 0.843804,
        "n_labels": 35,
        "retention": 0.907239
       }
      },
      "score": 94.413882
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 88.669702,
      "worst_class_recall": 0.902883
     },
     "uncertainty": {
      "full_accuracy": 0.929495,
      "risk_coverage_auc": 0.990165,
      "score": 98.852585,
      "selective_acc_at_50": 0.998237,
      "selective_acc_at_80": 0.984576
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 8166,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 3404,
    "runtime_s": 0.5,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. Parsed from the ARFF in the donor's zip, whose nominal class order is the class order used here.",
    "task": "dry_bean",
    "task_name": "Dry bean variety (7-class)",
    "timestamp": "2026-08-01T19:16:08+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 44.190614
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.7,
      "auc": 0.647086,
      "balanced_accuracy": 0.515238,
      "ci95": [
       0.491047,
       0.557287
      ],
      "n_test": 250,
      "score": 3.047619
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.7,
      "brier": 0.398872,
      "ece": 0.075166,
      "mean_confidence": 0.692566,
      "overconfidence": -0.007434,
      "score": 62.417034
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.030476,
      "mean_agreement": 0.983111,
      "mean_kappa": 0.724422,
      "mean_retention": 0.875,
      "n_cases": 250,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         1,
         0.988
        ],
        "mean_kappa": 0.945343,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.98,
         0.956,
         0.96
        ],
        "mean_kappa": 0.417408,
        "mean_retention": 0.770833,
        "retention_by_severity": [
         0.5625,
         1,
         0.75
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.996,
         0.988,
         0.98
        ],
        "mean_kappa": 0.810516,
        "mean_retention": 0.854167,
        "retention_by_severity": [
         1,
         1,
         0.5625
        ],
        "scored": true
       }
      },
      "score": 87.5,
      "worst_retention": 0.5625
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.030476,
      "mean_retention": 1,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.373333,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.24381,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.060952,
        "n_labels": 10,
        "retention": 1
       }
      },
      "score": 100
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.008021,
        "per_group": {
         "45-59": 0.5,
         "<45": 0.508021
        },
        "worst": 0.5
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 0,
      "worst_class_recall": 0.053333
     },
     "uncertainty": {
      "full_accuracy": 0.7,
      "risk_coverage_auc": 0.778255,
      "score": 55.65103,
      "selective_acc_at_50": 0.816,
      "selective_acc_at_80": 0.725
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 600,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / financial",
    "n_classes": 2,
    "n_test": 250,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "german_credit",
    "task_name": "Consumer credit risk (Statlog)",
    "timestamp": "2026-08-01T19:13:41+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 57.864832
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.773333,
      "auc": 0.875,
      "balanced_accuracy": 0.773214,
      "ci95": [
       0.666652,
       0.855136
      ],
      "n_test": 75,
      "score": 54.642857
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.773333,
      "brier": 0.356081,
      "ece": 0.185013,
      "mean_confidence": 0.953685,
      "overconfidence": 0.180352,
      "score": 7.493457
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.546429,
      "mean_agreement": 0.862222,
      "mean_kappa": 0.720948,
      "mean_retention": 0.909949,
      "n_cases": 75,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9867,
         0.96,
         0.92
        ],
        "mean_kappa": 0.911121,
        "mean_retention": 0.969499,
        "retention_by_severity": [
         1,
         1,
         0.9085
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9067,
         0.8267,
         0.64
        ],
        "mean_kappa": 0.576133,
        "mean_retention": 0.877996,
        "retention_by_severity": [
         1,
         0.7451,
         0.8889
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9733,
         0.7867,
         0.76
        ],
        "mean_kappa": 0.67559,
        "mean_retention": 0.882353,
        "retention_by_severity": [
         0.9935,
         0.6536,
         1
        ],
        "scored": true
       }
      },
      "score": 90.994916,
      "worst_retention": 0.653595
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.546429,
      "mean_retention": 0.949891,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.546429,
        "n_labels": 178,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.528571,
        "n_labels": 40,
        "retention": 0.96732
       },
       "n5": {
        "chance_corrected": 0.482143,
        "n_labels": 10,
        "retention": 0.882353
       }
      },
      "score": 94.989107
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": null,
        "per_group": {
         "45-59": 0.775
        },
        "worst": null
       },
       "sex": {
        "gap": 0.182476,
        "per_group": {
         "female": 0.904545,
         "male": 0.722069
        },
        "worst": 0.722069
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 44.413793,
      "worst_class_recall": 0.771429
     },
     "uncertainty": {
      "full_accuracy": 0.773333,
      "risk_coverage_auc": 0.88873,
      "score": 77.746093,
      "selective_acc_at_50": 0.894737,
      "selective_acc_at_80": 0.9
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 178,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 75,
    "runtime_s": 0.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "heart_cleveland",
    "task_name": "Coronary artery disease (Cleveland)",
    "timestamp": "2026-08-01T19:13:35+00:00"
   },
   {
    "code_fingerprint": "de32bca72d56",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 83.242553
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.8925,
      "auc": 0.987593,
      "balanced_accuracy": 0.884246,
      "ci95": [
       0.871449,
       0.896673
      ],
      "n_test": 2000,
      "score": 86.109562
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.8925,
      "brier": 0.154096,
      "ece": 0.020923,
      "mean_confidence": 0.890577,
      "overconfidence": -0.001923,
      "score": 89.538423
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.856597,
      "mean_agreement": 0.792778,
      "mean_kappa": 0.746817,
      "mean_retention": 0.784433,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9775,
         0.9208,
         0.8417
        ],
        "mean_kappa": 0.893837,
        "mean_retention": 0.931594,
        "retention_by_severity": [
         0.9775,
         0.9464,
         0.8709
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.8983,
         0.8425,
         0.4708
        ],
        "mean_kappa": 0.683308,
        "mean_retention": 0.709405,
        "retention_by_severity": [
         0.9107,
         0.8527,
         0.3648
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9567,
         0.765,
         0.4617
        ],
        "mean_kappa": 0.663307,
        "mean_retention": 0.712299,
        "retention_by_severity": [
         0.9956,
         0.7609,
         0.3804
        ],
        "scored": true
       }
      },
      "score": 78.443291,
      "worst_retention": 0.364805
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.861096,
      "mean_retention": 0.848203,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.758046,
        "n_labels": 600,
        "retention": 0.880328
       },
       "n20": {
        "chance_corrected": 0.759698,
        "n_labels": 120,
        "retention": 0.882246
       },
       "n5": {
        "chance_corrected": 0.673407,
        "n_labels": 30,
        "retention": 0.782035
       }
      },
      "score": 84.820301
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 67.014218,
      "worst_class_recall": 0.725118
     },
     "uncertainty": {
      "full_accuracy": 0.8925,
      "risk_coverage_auc": 0.980239,
      "score": 97.628663,
      "selective_acc_at_50": 0.995,
      "selective_acc_at_80": 0.964375
     }
    },
    "domain": "earth",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 3769,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / multispectral",
    "n_classes": 6,
    "n_test": 2000,
    "runtime_s": 0.5,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (Statlog sat.trn 4435 / sat.tst 2000), scored whole, with validation carved from sat.trn only (stratified 15%, seed 20260727). LEAKS, AND THE ACCURACIES HERE ARE OPTIMISTIC BECAUSE OF IT: each row is a 3x3 ground window, and 1,939 of the 2,000 test windows (97%) have an immediate neighbour in the training file sharing six of their nine ground pixels -- the two files are interleaved windows of one 82x100 scene, not separated regions. No 36-vector appears in both files, so byte-level deduplication passes while the test set overlaps training data on the ground; byte disjointness is not independence. Measured in variation/modalities/multispectral.md section 8. Raw class codes are 1,2,3,4,5,7 -- code 6 was withdrawn by the donors -- and are mapped positionally onto 0..5, not by subtracting one.",
    "task": "landsat_statlog",
    "task_name": "Landsat satellite land cover (6-class)",
    "timestamp": "2026-08-03T16:02:53+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 75.16107
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.878692,
      "auc": 0.966828,
      "balanced_accuracy": 0.875747,
      "ci95": [
       0.868973,
       0.88311
      ],
      "n_test": 9752,
      "score": 75.149427
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.878692,
      "brier": 0.152525,
      "ece": 0.071815,
      "mean_confidence": 0.903722,
      "overconfidence": 0.02503,
      "score": 64.092541
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.759081,
      "mean_agreement": 0.93787,
      "mean_kappa": 0.784962,
      "mean_retention": 0.812373,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9958,
         0.9883,
         0.97
        ],
        "mean_kappa": 0.959605,
        "mean_retention": 0.954662,
        "retention_by_severity": [
         0.9852,
         0.9693,
         0.9095
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         1,
         0.9867,
         0.8133
        ],
        "mean_kappa": 0.775717,
        "mean_retention": 0.81579,
        "retention_by_severity": [
         1,
         1,
         0.4474
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.99,
         0.9533,
         0.7433
        ],
        "mean_kappa": 0.619565,
        "mean_retention": 0.666667,
        "retention_by_severity": [
         1,
         1,
         0
        ],
        "scored": true
       }
      },
      "score": 81.23729,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.751494,
      "mean_retention": 0.705325,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.159243,
        "n_labels": 200,
        "retention": 0.211901
       },
       "n20": {
        "chance_corrected": 0.805951,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.679405,
        "n_labels": 10,
        "retention": 0.904072
       }
      },
      "score": 70.532463
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 74.133724,
      "worst_class_recall": 0.870669
     },
     "uncertainty": {
      "full_accuracy": 0.878692,
      "risk_coverage_auc": 0.982188,
      "score": 96.437524,
      "selective_acc_at_50": 0.998769,
      "selective_acc_at_80": 0.966163
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 9752,
    "runtime_s": 0.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 2 (datatest2.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection_2",
    "task_name": "Room occupancy from sensors (period 2)",
    "timestamp": "2026-08-01T19:15:22+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 88.458322
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.969981,
      "auc": 0.992196,
      "balanced_accuracy": 0.975716,
      "ci95": [
       0.971323,
       0.980114
      ],
      "n_test": 2665,
      "score": 95.143219
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.969981,
      "brier": 0.050988,
      "ece": 0.032631,
      "mean_confidence": 0.939775,
      "overconfidence": -0.030206,
      "score": 83.684666
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.943902,
      "mean_agreement": 0.933704,
      "mean_kappa": 0.835065,
      "mean_retention": 0.837403,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.9933,
         0.975
        ],
        "mean_kappa": 0.977544,
        "mean_retention": 0.982317,
        "retention_by_severity": [
         1,
         0.9938,
         0.9531
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         1,
         1,
         0.8333
        ],
        "mean_kappa": 0.871762,
        "mean_retention": 0.868416,
        "retention_by_severity": [
         1,
         1,
         0.6052
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9992,
         0.9842,
         0.6183
        ],
        "mean_kappa": 0.65589,
        "mean_retention": 0.661477,
        "retention_by_severity": [
         1,
         0.9819,
         0.0025
        ],
        "scored": true
       }
      },
      "score": 83.740308,
      "worst_retention": 0.002516
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.951432,
      "mean_retention": 0.766718,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.60361,
        "n_labels": 200,
        "retention": 0.634422
       },
       "n20": {
        "chance_corrected": 0.869615,
        "n_labels": 40,
        "retention": 0.914006
       },
       "n5": {
        "chance_corrected": 0.715216,
        "n_labels": 10,
        "retention": 0.751726
       }
      },
      "score": 76.671821
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 90.903721,
      "worst_class_recall": 0.954519
     },
     "uncertainty": {
      "full_accuracy": 0.969981,
      "risk_coverage_auc": 0.993172,
      "score": 98.634328,
      "selective_acc_at_50": 0.994745,
      "selective_acc_at_80": 0.994371
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 2665,
    "runtime_s": 0.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 1 (datatest.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection",
    "task_name": "Room occupancy from sensors (period 1)",
    "timestamp": "2026-08-01T19:15:19+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.58,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 12.753528
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.92602,
      "auc": 0.381907,
      "balanced_accuracy": 0.493207,
      "ci95": [
       0.487303,
       0.497326
      ],
      "n_test": 392,
      "score": 0
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.92602,
      "brier": 0.194624,
      "ece": 0.150797,
      "mean_confidence": 0.784734,
      "overconfidence": -0.141286,
      "score": 24.601508
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 392,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 0,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9847,
         0.977,
         0.9337
        ],
        "mean_kappa": 0.272715,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9898,
         0.9898,
         0.9872
        ],
        "mean_kappa": 0.275219,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9668,
         0.8342,
         0.9872
        ],
        "mean_kappa": 0.085471,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0,
      "mean_retention": 0,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.024457,
        "n_labels": 176,
        "retention": 0
       },
       "n20": {
        "chance_corrected": 0,
        "n_labels": 40,
        "retention": 0
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 0
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.92602,
      "risk_coverage_auc": 0.919885,
      "score": 83.976999,
      "selective_acc_at_50": 0.903061,
      "selective_acc_at_80": 0.933121
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 940,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / process sensors",
    "n_classes": 2,
    "n_test": 392,
    "runtime_s": 0.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "chronological split minted by MedEval-1 on the production timestamp shipped in secom_labels.data: earliest 60% train, next 15% validation, latest 25% test. A random split is not defensible on a process-monitoring corpus -- tools drift and are recalibrated, so shuffling lets the model interpolate a tool's state from wafers measured minutes on either side, while production always extrapolates forward. Missing values imputed with the TRAIN median only; channels all-missing or constant on the training period dropped (468 of 590 channels kept).",
    "task": "secom",
    "task_name": "Semiconductor fabrication pass/fail",
    "timestamp": "2026-08-01T19:14:08+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 68.207417
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.716049,
      "auc": 0.934156,
      "balanced_accuracy": 0.735533,
      "ci95": [
       0.683119,
       0.788691
      ],
      "n_test": 486,
      "score": 69.14557
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.716049,
      "brier": 0.382533,
      "ece": 0.050956,
      "mean_confidence": 0.697842,
      "overconfidence": -0.018207,
      "score": 74.521966
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.691456,
      "mean_agreement": 0.779835,
      "mean_kappa": 0.697477,
      "mean_retention": 0.681838,
      "n_cases": 486,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9753,
         0.9218,
         0.823
        ],
        "mean_kappa": 0.878718,
        "mean_retention": 0.850371,
        "retention_by_severity": [
         0.9389,
         0.854,
         0.7582
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.7572,
         0.6481,
         0.6214
        ],
        "mean_kappa": 0.560715,
        "mean_retention": 0.556781,
        "retention_by_severity": [
         0.7287,
         0.5074,
         0.4343
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9877,
         0.8416,
         0.4424
        ],
        "mean_kappa": 0.652999,
        "mean_retention": 0.638364,
        "retention_by_severity": [
         0.999,
         0.8073,
         0.1088
        ],
        "scored": true
       }
      },
      "score": 68.183834,
      "worst_retention": 0.108786
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.691456,
      "mean_retention": 0.914009,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.719862,
        "n_labels": 571,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.618915,
        "n_labels": 140,
        "retention": 0.89509
       },
       "n5": {
        "chance_corrected": 0.585619,
        "n_labels": 35,
        "retention": 0.846937
       }
      },
      "score": 91.400901
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "steel_type": {
        "gap": 0.218065,
        "per_group": {
         "A300": 0.404985,
         "A400": 0.62305
        },
        "worst": 0.404985
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 30.581558,
      "worst_class_recall": 0.554455
     },
     "uncertainty": {
      "full_accuracy": 0.716049,
      "risk_coverage_auc": 0.878885,
      "score": 85.869899,
      "selective_acc_at_50": 0.893004,
      "selective_acc_at_80": 0.784062
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 1164,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 486,
    "runtime_s": 0.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. The label is the argmax of the seven trailing one-hot fault columns, which are dropped from the features.",
    "task": "steel_plate_faults",
    "task_name": "Steel plate surface faults (7-class)",
    "timestamp": "2026-08-01T19:13:53+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 60.546995
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.7575,
      "auc": 0.817958,
      "balanced_accuracy": 0.760351,
      "ci95": [
       0.720465,
       0.799296
      ],
      "n_test": 400,
      "score": 52.070144
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.7575,
      "brier": 0.345518,
      "ece": 0.053004,
      "mean_confidence": 0.757054,
      "overconfidence": -0.000446,
      "score": 73.498212
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.520701,
      "mean_agreement": 0.835556,
      "mean_kappa": 0.66599,
      "mean_retention": 0.746405,
      "n_cases": 400,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.975,
         0.9475,
         0.8875
        ],
        "mean_kappa": 0.872575,
        "mean_retention": 0.944064,
        "retention_by_severity": [
         1,
         1,
         0.8322
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.895,
         0.7125,
         0.6975
        ],
        "mean_kappa": 0.536864,
        "mean_retention": 0.633439,
        "retention_by_severity": [
         0.9201,
         0.4838,
         0.4964
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9575,
         0.925,
         0.5225
        ],
        "mean_kappa": 0.58853,
        "mean_retention": 0.661713,
        "retention_by_severity": [
         1,
         0.9851,
         0
        ],
        "scored": true
       }
      },
      "score": 74.640548,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.520701,
      "mean_retention": 0.543375,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.51231,
        "n_labels": 200,
        "retention": 0.983885
       },
       "n20": {
        "chance_corrected": 0.336499,
        "n_labels": 40,
        "retention": 0.646241
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 54.337547
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 43.925234,
      "worst_class_recall": 0.719626
     },
     "uncertainty": {
      "full_accuracy": 0.7575,
      "risk_coverage_auc": 0.839443,
      "score": 67.888671,
      "selective_acc_at_50": 0.855,
      "selective_acc_at_80": 0.8
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 959,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 400,
    "runtime_s": 0.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (red wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_red",
    "task_name": "Wine quality, red (binarised)",
    "timestamp": "2026-08-01T19:13:43+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 56.313177
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.783673,
      "auc": 0.856765,
      "balanced_accuracy": 0.728341,
      "ci95": [
       0.703696,
       0.752881
      ],
      "n_test": 1225,
      "score": 45.668113
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.783673,
      "brier": 0.289985,
      "ece": 0.016545,
      "mean_confidence": 0.796485,
      "overconfidence": 0.012812,
      "score": 91.727524
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.453878,
      "mean_agreement": 0.800185,
      "mean_kappa": 0.535321,
      "mean_retention": 0.60334,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9875,
         0.9425,
         0.8633
        ],
        "mean_kappa": 0.830075,
        "mean_retention": 0.9397,
        "retention_by_severity": [
         1,
         0.9553,
         0.8638
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.8992,
         0.785,
         0.7467
        ],
        "mean_kappa": 0.333708,
        "mean_retention": 0.338513,
        "retention_by_severity": [
         0.8038,
         0.1819,
         0.0299
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9658,
         0.7533,
         0.2583
        ],
        "mean_kappa": 0.442181,
        "mean_retention": 0.531808,
        "retention_by_severity": [
         0.9079,
         0.6875,
         0
        ],
        "scored": true
       }
      },
      "score": 60.334034,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.456681,
      "mean_retention": 0.611774,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.437214,
        "n_labels": 200,
        "retention": 0.957372
       },
       "n20": {
        "chance_corrected": 0.400943,
        "n_labels": 40,
        "retention": 0.877949
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 61.17737
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 12.195122,
      "worst_class_recall": 0.560976
     },
     "uncertainty": {
      "full_accuracy": 0.783673,
      "risk_coverage_auc": 0.901346,
      "score": 80.269157,
      "selective_acc_at_50": 0.924837,
      "selective_acc_at_80": 0.84898
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 2938,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 1225,
    "runtime_s": 0.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (white wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_white",
    "task_name": "Wine quality, white (binarised)",
    "timestamp": "2026-08-01T19:14:11+00:00"
   }
  ],
  "heldout": null,
  "sources": [
   {
    "path": "results/cross_site/occupancy__tab_mlp.json",
    "sha256": "96185a3fc9e50e26176f58216d6d522c5e40aeab923d6ad774380e4c9f9ae0b4"
   },
   {
    "path": "results/heldout/audit_report.json",
    "sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
   },
   {
    "path": "results/heldout/cycle.json",
    "sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
   },
   {
    "path": "results/heldout/manifests/gen1.seal.json",
    "sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
   },
   {
    "path": "results/heldout/manifests/gen2.seal.json",
    "sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
   },
   {
    "path": "results/heldout/manifests/public.seal.json",
    "sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
   },
   {
    "path": "results/records/adult_income__tab_mlp.json",
    "sha256": "49049d8cc6e2d1511d0385ccb566ad799ad906a00a22d961d04563d6a089eb2e"
   },
   {
    "path": "results/records/anuran_mfcc__tab_mlp.json",
    "sha256": "c10bdf3bcd87fd9e4ab7431167b8442d5c53313bb1e06b3b22994604aa89e1b3"
   },
   {
    "path": "results/records/breast_wdbc__tab_mlp.json",
    "sha256": "7bc5146466b300da8ebd3a3e208e1ac0b75c0f39c6575f172e17e7670bf3cfbc"
   },
   {
    "path": "results/records/coil2000_caravan__tab_mlp.json",
    "sha256": "7dafdcc9eb9a9b07415bd6c683b526722fc1ac0044f69e998c963c012888ef3d"
   },
   {
    "path": "results/records/diabetes_pima__tab_mlp.json",
    "sha256": "b2936c4291f7fc6111c6d2d58cf696b02438981314fbb8fdff3bc65bfa8161ca"
   },
   {
    "path": "results/records/dry_bean__tab_mlp.json",
    "sha256": "62a6714e9e0ad9464a8fd2f4eb464569af90b093536fa102a5a2de15595241e6"
   },
   {
    "path": "results/records/german_credit__tab_mlp.json",
    "sha256": "8b03a5007c340ee097c252031d573b153b2455f39a79695fb26160d28bdbbeca"
   },
   {
    "path": "results/records/heart_cleveland__tab_mlp.json",
    "sha256": "f90552415bce174bc8f36a2cdad2cea905172b9519dbdc6f0ba2587719a17769"
   },
   {
    "path": "results/records/landsat_statlog__tab_mlp.json",
    "sha256": "e670f180d3cf429e24816c312e75eb254daa225344297c41c4ab2885b618e9ce"
   },
   {
    "path": "results/records/occupancy_detection_2__tab_mlp.json",
    "sha256": "f7cbc3ac657e97ef2b920b9c17df8fdc128e06a7d45a46a9ef0d58e5469c2466"
   },
   {
    "path": "results/records/occupancy_detection__tab_mlp.json",
    "sha256": "05e8a020e37e594519c3a6143e63cc85d2a07aff8f411c22288addd891e8b8e9"
   },
   {
    "path": "results/records/secom__tab_mlp.json",
    "sha256": "bea110d82ce415d7813b396932e2cea6856de9aca96fc0f0a5f3063e430f5637"
   },
   {
    "path": "results/records/steel_plate_faults__tab_mlp.json",
    "sha256": "51d1d0ee9be444d388aa1d69a1ff2e2345c7d52fd83120214425e9961b53d0f9"
   },
   {
    "path": "results/records/wine_quality_red__tab_mlp.json",
    "sha256": "71bc7b14871de8292e51da89ce19fb9197099a0f24b71c766008c2e12744547b"
   },
   {
    "path": "results/records/wine_quality_white__tab_mlp.json",
    "sha256": "02a436ca6e8edabe3ac34b39a230eb63b12a450238e64ec9597c55caf58c9ddc"
   }
  ]
 },
 "declared": null,
 "declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
 "expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
 "expiry_utc": "2026-09-20T00:00:00Z",
 "issued_utc": "2026-08-03T16:03:42Z",
 "issuer": {
  "algo": "Ed25519",
  "key_id": "ns-passport-2026-07",
  "name": "NakedSignal",
  "public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
 },
 "observed": null,
 "observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
 "passport_version": "0.1",
 "signature": "zEkY2QAqT3GIWy9V56ID/nMWM2dCMjZ2x6Jt76mTKQ+07enpAzrNWvyW1ghU0utdkakq1Z4GyzFcizjfv1rCCQ==",
 "spec_version": "MedEval-1 v0.1",
 "subject": {
  "behaviour_version": "1.0.0",
  "code_fingerprint": null,
  "code_fingerprints": [
   "b444c02112cc",
   "de32bca72d56"
  ],
  "family": "shallow",
  "model_id": "tab_mlp",
  "model_name": "MLP (128, 64)",
  "params": "2 hidden layers, adam, early stopping when n>=200"
 }
}

Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.