NakedSignal OS · Documentation
Back to the consoleDocumentation

Logistic regression

LIVE

The signed evidence record for tab_logreg. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.

SIGNATURE VALIDexit 0signature valid, issued by NakedSignal under key ns-passport-2026-07, not expired
SubjectLogistic regression tab_logreg
Configurationstandardised, L2, C=1
Code fingerprintb444c02112cc, de32bca72d56
Spec versionMedEval-1 v0.1
Issued2026-08-03T16:03:42Z by NakedSignal
Expires2026-09-20T00:00:00Zthe sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first
key ns-passport-2026-07 · Ed25519

The check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.

Measured by NakedSignal

Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 21 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.

TaskDomainn testaccuracycalibrationuncertaintysubgroupcorruptionacquisitionlimited datacostspec sensitivityIndex
adult_incomelabour1628153.28661997.09226991.08649429.4360990.644038n/a79.10956n/an/a70.641549
anuran_mfccecology99572.156942058.57515976.08922264.782891n/a87.030642n/an/a58.950366
breast_wdbcmedicine14398.11320886.14064199.45069296.22641579.717474n/a92.827635n/an/a91.618086
coil2000_caravanfinancial40001.57434694.04880294.400346081.516422n/a66.666667n/an/a47.275486
diabetes_pimamedicine19335.57213967.02773868.83329232.72727390.583491n/a96.614497n/an/a60.64544
dry_beanagriculture340493.09883290.47455598.80700787.96155875.224967n/a96.569948n/an/a89.218101
german_creditfinancial25032.38095262.53270375.10029127.62923496.732026n/a62.941176n/an/a56.039968
heart_clevelandmedicine7570.71428661.29050183.9364816084.624018n/a83.838384n/an/a72.605785
landsat_statlogearth200074.63175486.99191395.53959614.69194367.25434n/a97.095518n/an/a71.051886
occupancy_detection_2energy975279.95372683.57598498.41234974.91459288.425888n/a96.840405n/an/a84.713377
occupancy_detectionenergy266595.47294789.9259899.43803293.62079188.191024n/a92.965901n/an/a92.79503
secomindustrial392061.81427885.2305430not scoredn/a0n/an/a21.202385
steel_plate_faultsindustrial48670.43492566.00421882.42916334.17596866.145126n/a92.579436n/an/a67.058277
wine_quality_redagriculture40050.38689669.31806168.82126244.85981370.833887n/a63.096663n/an/a59.705537
wine_quality_whiteagriculture122538.22684485.63197976.669329063.772944n/a66.666667n/an/a52.273715

Cross-site retention · occupancy

retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column

Sitesoccupancy_detection, occupancy_detection_2
Mean retention0.918725
Worst retention0.837449

Private held-out track

this subject was not submitted to a sealed generation.

Known limitations

Subgroup metadataabsent on 9 of 15 corpora, so the subgroup score falls back to the worst diagnostic class. That absence is a disclosure, not a pass.
Training set cappedyes on at least one task — see the raw JSON for which
Observed in deploymentnothing — see the third panel

Audit chain

Bundles re-derived48 of 48
Ledger100 entries, head 6b7b371ff650a8ed, chain intact
Re-run it yourselfpython src/heldout/audit.py 1f66c7719ab3943c6fcc
21 source files, with hashes
results/cross_site/occupancy__tab_logreg.json39f6dd624d9723016056f3deb5caf4f5be5c852ab6ec756c45afd28732dd3d52
results/heldout/audit_report.jsondf460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2
results/heldout/cycle.jsone012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502
results/heldout/manifests/gen1.seal.json91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81
results/heldout/manifests/gen2.seal.jsonf61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547
results/heldout/manifests/public.seal.json1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c
results/records/adult_income__tab_logreg.json9d49f868e37c94bdcbde73d8ea4f043ccd8d8fc63c89550f8e2d95605e67deb0
results/records/anuran_mfcc__tab_logreg.jsonc1271e3ccaa18a559772a12ca02fb9265413e70bb06f299c2c87b78c9f5ac6f0
results/records/breast_wdbc__tab_logreg.json9756b728ad8a36e7bd2806f46f3fb6b9c2bd8590a97bcf5c22fc66f28fb165be
results/records/coil2000_caravan__tab_logreg.json252f17a68d072d2cb05ee5514df2f21bdf1d6caa45904c07260ed5c59778507b
results/records/diabetes_pima__tab_logreg.jsond8c33c2b3ecb291523db6e86a5fb15de72eab3df9c9f8050abf5a97a862441d3
results/records/dry_bean__tab_logreg.json746ef76d9275ef54925dba9afd291e7dcc4986601dfbf86802325ca2bffa87ee
results/records/german_credit__tab_logreg.jsonef0656154fcaed4e3ec40dfe3f52387aed341910b42f970cb4feb3ea9a0dcdd9
results/records/heart_cleveland__tab_logreg.jsonee6774fabcb5b6f03727dfba90c2b3081f6f98b06172ecfc2f9e4241b3432ee2
results/records/landsat_statlog__tab_logreg.json2fe08d5aceaf2cfc8e5eb9cbdf58169f99876093ad34aea9e16de93890ee949f
results/records/occupancy_detection_2__tab_logreg.json33859c9106ad2179d4b479a3b53c112f6e1e36dfd7324b915658e8ab4f751bce
results/records/occupancy_detection__tab_logreg.json7a7adae2920572b0569780fdcc65883daabc66360174c0989ebfb3727800090d
results/records/secom__tab_logreg.json5e2c3e5e4a5dea390dd9968e2121956ee3e704ff150bd3a93301bd52d299e797
results/records/steel_plate_faults__tab_logreg.jsonbcb35954c6c21b3a1a223d12a924f3ff3ca1cebc8016823872b1c305aa3b4e6b
results/records/wine_quality_red__tab_logreg.json528c7eedc18d0baf7620c8f7ba6e36991a41f7675d52c150ca92ce4ae23452fe
results/records/wine_quality_white__tab_logreg.jsona1caccca58fe53897a0124cf62fc019fe3b4bc6b3a15aee3d5c22da9d3460e68
Declared by vendor

Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.

No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.

Observed in deployment

Production history and drift — how the model has actually behaved since it was deployed, on real traffic.

Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.

Raw JSON — the whole signed document, 76,592 characters
{
 "canonicalisation": {
  "float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
  "form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
  "signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
 },
 "computed": {
  "audit": {
   "audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
   "bundles": 48,
   "ledger_entries": 100,
   "ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
   "ledger_intact": true,
   "rederived": 48
  },
  "cross_site": {
   "code_fingerprint": "b444c02112cc",
   "device": "mps",
   "family": "occupancy",
   "matrix": {
    "occupancy_detection": {
     "occupancy_detection": {
      "balanced_accuracy": 0.977365,
      "chance_corrected": 0.954729,
      "ece": 0.020148
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.899769,
      "chance_corrected": 0.799537,
      "ece": 0.032848,
      "retention": 0.837449
     }
    },
    "occupancy_detection_2": {
     "occupancy_detection": {
      "balanced_accuracy": 0.977365,
      "chance_corrected": 0.954729,
      "ece": 0.020148,
      "retention": 1
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.899769,
      "chance_corrected": 0.799537,
      "ece": 0.032848
     }
    }
   },
   "mean_retention": 0.918725,
   "note": "retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column",
   "score": 91.872451,
   "site_labels": {
    "occupancy_detection": "test period 1",
    "occupancy_detection_2": "test period 2"
   },
   "sites": [
    "occupancy_detection",
    "occupancy_detection_2"
   ],
   "spec_version": "MedEval-1 v0.1",
   "worst_retention": 0.837449
  },
  "evaluations": [
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 70.641549
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.851299,
      "auc": 0.900916,
      "balanced_accuracy": 0.766433,
      "ci95": [
       0.758858,
       0.775213
      ],
      "n_test": 16281,
      "score": 53.286619
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.851299,
      "brier": 0.206377,
      "ece": 0.005815,
      "mean_confidence": 0.856261,
      "overconfidence": 0.004961,
      "score": 97.092269
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.547023,
      "mean_agreement": 0.950926,
      "mean_kappa": 0.834848,
      "mean_retention": 0.90644,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9975,
         0.9833,
         0.97
        ],
        "mean_kappa": 0.95156,
        "mean_retention": 0.999333,
        "retention_by_severity": [
         0.998,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9692,
         0.9492,
         0.8283
        ],
        "mean_kappa": 0.66839,
        "mean_retention": 0.725449,
        "retention_by_severity": [
         0.9936,
         0.9041,
         0.2786
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9992,
         0.9942,
         0.8675
        ],
        "mean_kappa": 0.884593,
        "mean_retention": 0.99454,
        "retention_by_severity": [
         0.998,
         0.9856,
         1
        ],
        "scored": true
       }
      },
      "score": 90.644038,
      "worst_retention": 0.278623
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.532866,
      "mean_retention": 0.791096,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.516932,
        "n_labels": 200,
        "retention": 0.970098
       },
       "n20": {
        "chance_corrected": 0.437943,
        "n_labels": 40,
        "retention": 0.821862
       },
       "n5": {
        "chance_corrected": 0.309769,
        "n_labels": 10,
        "retention": 0.581327
       }
      },
      "score": 79.10956
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.00671,
        "per_group": {
         "45-59": 0.758929,
         "60+": 0.765639,
         "<45": 0.759156
        },
        "worst": 0.758929
       },
       "race": {
        "gap": 0.120628,
        "per_group": {
         "Amer-Indian-Eskimo": 0.64718,
         "Asian-Pac-Islander": 0.766549,
         "Black": 0.721149,
         "Other": 0.706364,
         "White": 0.767809
        },
        "worst": 0.64718
       },
       "sex": {
        "gap": 0.000022,
        "per_group": {
         "Female": 0.757175,
         "Male": 0.757197
        },
        "worst": 0.757175
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 29.43609,
      "worst_class_recall": 0.605564
     },
     "uncertainty": {
      "full_accuracy": 0.851299,
      "risk_coverage_auc": 0.955432,
      "score": 91.086494,
      "selective_acc_at_50": 0.978378,
      "selective_acc_at_80": 0.913321
     }
    },
    "domain": "labour",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 27676,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": true
    },
    "modality": "Tabular / socioeconomic",
    "n_classes": 2,
    "n_test": 16281,
    "runtime_s": 0.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (UCI adult.test scored whole); validation carved from adult.data only (stratified 15%, seed 20260727)",
    "task": "adult_income",
    "task_name": "Census income (>$50K)",
    "timestamp": "2026-08-01T19:16:09+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 58.950366
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.620101,
      "auc": 0.92757,
      "balanced_accuracy": 0.749412,
      "ci95": [
       0.707507,
       0.788672
      ],
      "n_test": 995,
      "score": 72.156942
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.620101,
      "brier": 0.65531,
      "ece": 0.260751,
      "mean_confidence": 0.878081,
      "overconfidence": 0.257981,
      "score": 0
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.721569,
      "mean_agreement": 0.736683,
      "mean_kappa": 0.697545,
      "mean_retention": 0.647829,
      "n_cases": 995,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9497,
         0.8693,
         0.7538
        ],
        "mean_kappa": 0.834531,
        "mean_retention": 0.875213,
        "retention_by_severity": [
         0.9282,
         0.8991,
         0.7983
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.7598,
         0.4935,
         0.3487
        ],
        "mean_kappa": 0.468494,
        "mean_retention": 0.311539,
        "retention_by_severity": [
         0.6965,
         0.1654,
         0.0727
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9779,
         0.8774,
         0.6
        ],
        "mean_kappa": 0.789612,
        "mean_retention": 0.756735,
        "retention_by_severity": [
         0.9967,
         0.8264,
         0.4471
        ],
        "scored": true
       }
      },
      "score": 64.782891,
      "worst_retention": 0.07272
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.721569,
      "mean_retention": 0.870306,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.776993,
        "n_labels": 794,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.727142,
        "n_labels": 200,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.440821,
        "n_labels": 50,
        "retention": 0.610919
       }
      },
      "score": 87.030642
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "family": {
        "gap": 0.065627,
        "per_group": {
         "Hylidae": 0.85043,
         "Leptodactylidae": 0.784803
        },
        "worst": 0.784803
       },
       "genus": {
        "gap": null,
        "per_group": {
         "Hypsiboas": 0.925172
        },
        "worst": null
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 76.089222,
      "worst_class_recall": 0.223975
     },
     "uncertainty": {
      "full_accuracy": 0.620101,
      "risk_coverage_auc": 0.627176,
      "score": 58.575159,
      "selective_acc_at_50": 0.648594,
      "selective_acc_at_80": 0.644472
     }
    },
    "domain": "ecology",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 5073,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Bioacoustic / MFCC",
    "n_classes": 10,
    "n_test": 995,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "no official split; minted at the standard seed (20260727) by partitioning the 60 RecordID recording groups 60/15/15/25 and taking whole recordings, so no recording spans two splits. A row-level random split would place near-duplicate syllables from one recording on both sides and score memorisation (~99%); class balance is therefore uncontrolled across splits and is not stratified",
    "task": "anuran_mfcc",
    "task_name": "Anuran call species (10-class)",
    "timestamp": "2026-08-01T19:14:27+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 91.618086
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.993007,
      "auc": 0.992662,
      "balanced_accuracy": 0.990566,
      "ci95": [
       0.968077,
       1
      ],
      "n_test": 143,
      "score": 98.113208
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.993007,
      "brier": 0.026788,
      "ece": 0.027719,
      "mean_confidence": 0.971885,
      "overconfidence": -0.021122,
      "score": 86.140641
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.981132,
      "mean_agreement": 0.874903,
      "mean_kappa": 0.78041,
      "mean_retention": 0.797175,
      "n_cases": 143,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.986,
         0.965,
         0.8951
        ],
        "mean_kappa": 0.894183,
        "mean_retention": 0.927137,
        "retention_by_severity": [
         0.9774,
         0.9434,
         0.8607
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.993,
         0.993,
         0.9371
        ],
        "mean_kappa": 0.943292,
        "mean_retention": 0.937393,
        "retention_by_severity": [
         0.9887,
         0.9887,
         0.8348
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.993,
         0.7483,
         0.3636
        ],
        "mean_kappa": 0.503756,
        "mean_retention": 0.526994,
        "retention_by_severity": [
         0.9887,
         0.5923,
         0
        ],
        "scored": true
       }
      },
      "score": 79.717474,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.981132,
      "mean_retention": 0.928276,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.970021,
        "n_labels": 200,
        "retention": 0.988675
       },
       "n20": {
        "chance_corrected": 0.814465,
        "n_labels": 40,
        "retention": 0.830128
       },
       "n5": {
        "chance_corrected": 0.947799,
        "n_labels": 10,
        "retention": 0.966026
       }
      },
      "score": 92.827635
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 96.226415,
      "worst_class_recall": 0.981132
     },
     "uncertainty": {
      "full_accuracy": 0.993007,
      "risk_coverage_auc": 0.997253,
      "score": 99.450692,
      "selective_acc_at_50": 1,
      "selective_acc_at_80": 0.991228
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 341,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 2,
    "n_test": 143,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "breast_wdbc",
    "task_name": "Breast cytology (WDBC)",
    "timestamp": "2026-08-01T19:13:37+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 47.275486
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.9405,
      "auc": 0.726572,
      "balanced_accuracy": 0.507872,
      "ci95": [
       0.501587,
       0.516474
      ],
      "n_test": 4000,
      "score": 1.574346
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.9405,
      "brier": 0.107726,
      "ece": 0.011902,
      "mean_confidence": 0.943005,
      "overconfidence": 0.002505,
      "score": 94.048802
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.040202,
      "mean_agreement": 0.951296,
      "mean_kappa": 0.483838,
      "mean_retention": 0.815164,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9983,
         0.9958,
         0.9883
        ],
        "mean_kappa": 0.644412,
        "mean_retention": 0.985325,
        "retention_by_severity": [
         0.956,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9967,
         0.9917,
         0.9858
        ],
        "mean_kappa": 0.234038,
        "mean_retention": 0.460168,
        "retention_by_severity": [
         0.956,
         0.2893,
         0.1352
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         1,
         0.9958,
         0.6092
        ],
        "mean_kappa": 0.573063,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 81.516422,
      "worst_retention": 0.13522
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.015743,
      "mean_retention": 0.666667,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.274787,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0,
        "n_labels": 40,
        "retention": 0
       },
       "n5": {
        "chance_corrected": 0.10187,
        "n_labels": 10,
        "retention": 1
       }
      },
      "score": 66.666667
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.016807
     },
     "uncertainty": {
      "full_accuracy": 0.9405,
      "risk_coverage_auc": 0.972002,
      "score": 94.400346,
      "selective_acc_at_50": 0.9735,
      "selective_acc_at_80": 0.961875
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 4948,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / insurance",
    "n_classes": 2,
    "n_test": 4000,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (CoIL Challenge 2000: ticdata2000.txt 5822 train / ticeval2000.txt 4000 test). The test labels were released separately, in tictgts2000.txt, only after the challenge closed -- they were not available to anyone tuning on this benchmark at the time. Test scored whole; validation carved from ticdata2000.txt only (stratified 15%, seed 20260727).",
    "task": "coil2000_caravan",
    "task_name": "Caravan insurance purchase (CoIL 2000)",
    "timestamp": "2026-08-01T19:15:22+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 60.64544
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.725389,
      "auc": 0.783345,
      "balanced_accuracy": 0.677861,
      "ci95": [
       0.613723,
       0.751398
      ],
      "n_test": 193,
      "score": 35.572139
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.725389,
      "brier": 0.361033,
      "ece": 0.065945,
      "mean_confidence": 0.790776,
      "overconfidence": 0.065387,
      "score": 67.027738
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.355721,
      "mean_agreement": 0.893495,
      "mean_kappa": 0.805428,
      "mean_retention": 0.905835,
      "n_cases": 193,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9896,
         0.9741,
         0.9171
        ],
        "mean_kappa": 0.906041,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9741,
         0.9378,
         0.9482
        ],
        "mean_kappa": 0.882132,
        "mean_retention": 0.931846,
        "retention_by_severity": [
         0.9134,
         0.8821,
         1
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9845,
         0.9275,
         0.3886
        ],
        "mean_kappa": 0.628111,
        "mean_retention": 0.785659,
        "retention_by_severity": [
         0.9973,
         1,
         0.3596
        ],
        "scored": true
       }
      },
      "score": 90.583491,
      "worst_retention": 0.35964
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.355721,
      "mean_retention": 0.966145,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.401801,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.429756,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.319593,
        "n_labels": 10,
        "retention": 0.898435
       }
      },
      "score": 96.614497
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.009091,
        "per_group": {
         "45-59": 0.663636,
         "<45": 0.672727
        },
        "worst": 0.663636
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 32.727273,
      "worst_class_recall": 0.522388
     },
     "uncertainty": {
      "full_accuracy": 0.725389,
      "risk_coverage_auc": 0.844166,
      "score": 68.833292,
      "selective_acc_at_50": 0.854167,
      "selective_acc_at_80": 0.785714
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 460,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 193,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "diabetes_pima",
    "task_name": "Diabetes onset (Pima)",
    "timestamp": "2026-08-01T19:13:35+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 89.218101
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.93067,
      "auc": 0.995574,
      "balanced_accuracy": 0.940847,
      "ci95": [
       0.933445,
       0.949987
      ],
      "n_test": 3404,
      "score": 93.098832
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.93067,
      "brier": 0.105137,
      "ece": 0.019051,
      "mean_confidence": 0.915137,
      "overconfidence": -0.015533,
      "score": 90.474555
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.93135,
      "mean_agreement": 0.778889,
      "mean_kappa": 0.7413,
      "mean_retention": 0.75225,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9775,
         0.9025,
         0.7675
        ],
        "mean_kappa": 0.858214,
        "mean_retention": 0.898208,
        "retention_by_severity": [
         0.9843,
         0.9204,
         0.79
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9575,
         0.8825,
         0.5983
        ],
        "mean_kappa": 0.768035,
        "mean_retention": 0.720179,
        "retention_by_severity": [
         0.9221,
         0.8503,
         0.3881
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9817,
         0.8467,
         0.0958
        ],
        "mean_kappa": 0.597651,
        "mean_retention": 0.638363,
        "retention_by_severity": [
         1,
         0.9151,
         0
        ],
        "scored": true
       }
      },
      "score": 75.224967,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.930988,
      "mean_retention": 0.965699,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.925595,
        "n_labels": 700,
        "retention": 0.994207
       },
       "n20": {
        "chance_corrected": 0.905913,
        "n_labels": 140,
        "retention": 0.973066
       },
       "n5": {
        "chance_corrected": 0.865657,
        "n_labels": 35,
        "retention": 0.929825
       }
      },
      "score": 96.569948
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 87.961558,
      "worst_class_recall": 0.896813
     },
     "uncertainty": {
      "full_accuracy": 0.93067,
      "risk_coverage_auc": 0.989774,
      "score": 98.807007,
      "selective_acc_at_50": 0.99765,
      "selective_acc_at_80": 0.983841
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 8166,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 3404,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. Parsed from the ARFF in the donor's zip, whose nominal class order is the class order used here.",
    "task": "dry_bean",
    "task_name": "Dry bean variety (7-class)",
    "timestamp": "2026-08-01T19:15:26+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 56.039968
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.756,
      "auc": 0.791771,
      "balanced_accuracy": 0.661905,
      "ci95": [
       0.611306,
       0.722672
      ],
      "n_test": 250,
      "score": 32.380952
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.756,
      "brier": 0.332951,
      "ece": 0.074935,
      "mean_confidence": 0.79844,
      "overconfidence": 0.04244,
      "score": 62.532703
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.32381,
      "mean_agreement": 0.953333,
      "mean_kappa": 0.858539,
      "mean_retention": 0.96732,
      "n_cases": 250,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.996,
         0.984,
         0.964
        ],
        "mean_kappa": 0.940772,
        "mean_retention": 0.958824,
        "retention_by_severity": [
         1,
         1,
         0.8765
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.936,
         0.92,
         0.88
        ],
        "mean_kappa": 0.736584,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.988,
         0.98,
         0.932
        ],
        "mean_kappa": 0.898261,
        "mean_retention": 0.943137,
        "retention_by_severity": [
         0.9353,
         0.9824,
         0.9118
        ],
        "scored": true
       }
      },
      "score": 96.732026,
      "worst_retention": 0.876471
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.32381,
      "mean_retention": 0.629412,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.466667,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.213333,
        "n_labels": 40,
        "retention": 0.658824
       },
       "n5": {
        "chance_corrected": 0.074286,
        "n_labels": 10,
        "retention": 0.229412
       }
      },
      "score": 62.941176
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.155835,
        "per_group": {
         "45-59": 0.793981,
         "<45": 0.638146
        },
        "worst": 0.638146
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 27.629234,
      "worst_class_recall": 0.426667
     },
     "uncertainty": {
      "full_accuracy": 0.756,
      "risk_coverage_auc": 0.875501,
      "score": 75.100291,
      "selective_acc_at_50": 0.896,
      "selective_acc_at_80": 0.805
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 600,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / financial",
    "n_classes": 2,
    "n_test": 250,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "german_credit",
    "task_name": "Consumer credit risk (Statlog)",
    "timestamp": "2026-08-01T19:13:39+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 72.605785
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.853333,
      "auc": 0.903571,
      "balanced_accuracy": 0.853571,
      "ci95": [
       0.773817,
       0.928504
      ],
      "n_test": 75,
      "score": 70.714286
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.853333,
      "brier": 0.246061,
      "ece": 0.077419,
      "mean_confidence": 0.829689,
      "overconfidence": -0.023644,
      "score": 61.290501
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.707143,
      "mean_agreement": 0.89037,
      "mean_kappa": 0.777977,
      "mean_retention": 0.84624,
      "n_cases": 75,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.96,
         0.96,
         0.8667
        ],
        "mean_kappa": 0.857986,
        "mean_retention": 0.90404,
        "retention_by_severity": [
         0.9646,
         0.9646,
         0.7828
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.96,
         0.9333,
         0.6933
        ],
        "mean_kappa": 0.71917,
        "mean_retention": 0.789562,
        "retention_by_severity": [
         0.9545,
         0.8838,
         0.5303
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9467,
         0.7733,
         0.92
        ],
        "mean_kappa": 0.756777,
        "mean_retention": 0.845118,
        "retention_by_severity": [
         0.9141,
         0.6919,
         0.9293
        ],
        "scored": true
       }
      },
      "score": 84.624018,
      "worst_retention": 0.530303
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.707143,
      "mean_retention": 0.838384,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.707143,
        "n_labels": 178,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.578571,
        "n_labels": 40,
        "retention": 0.818182
       },
       "n5": {
        "chance_corrected": 0.492857,
        "n_labels": 10,
        "retention": 0.69697
       }
      },
      "score": 83.838384
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": null,
        "per_group": {
         "45-59": 0.85
        },
        "worst": null
       },
       "sex": {
        "gap": 0.076552,
        "per_group": {
         "female": 0.8,
         "male": 0.876552
        },
        "worst": 0.8
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 60,
      "worst_class_recall": 0.85
     },
     "uncertainty": {
      "full_accuracy": 0.853333,
      "risk_coverage_auc": 0.919682,
      "score": 83.936481,
      "selective_acc_at_50": 0.921053,
      "selective_acc_at_80": 0.883333
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 178,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 75,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "heart_cleveland",
    "task_name": "Coronary artery disease (Cleveland)",
    "timestamp": "2026-08-01T19:13:33+00:00"
   },
   {
    "code_fingerprint": "de32bca72d56",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 71.051886
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.8355,
      "auc": 0.975552,
      "balanced_accuracy": 0.788598,
      "ci95": [
       0.774709,
       0.805025
      ],
      "n_test": 2000,
      "score": 74.631754
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.8355,
      "brier": 0.212281,
      "ece": 0.026016,
      "mean_confidence": 0.858246,
      "overconfidence": 0.022746,
      "score": 86.991913
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.742217,
      "mean_agreement": 0.695,
      "mean_kappa": 0.635778,
      "mean_retention": 0.672543,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.975,
         0.9033,
         0.805
        ],
        "mean_kappa": 0.868651,
        "mean_retention": 0.93369,
        "retention_by_severity": [
         0.9803,
         0.9627,
         0.8581
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9108,
         0.675,
         0.2592
        ],
        "mean_kappa": 0.536669,
        "mean_retention": 0.547975,
        "retention_by_severity": [
         0.9316,
         0.6109,
         0.1015
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9683,
         0.63,
         0.1283
        ],
        "mean_kappa": 0.502014,
        "mean_retention": 0.535965,
        "retention_by_severity": [
         0.9889,
         0.5961,
         0.0229
        ],
        "scored": true
       }
      },
      "score": 67.25434,
      "worst_retention": 0.022933
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.746318,
      "mean_retention": 0.970955,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.775168,
        "n_labels": 600,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.732551,
        "n_labels": 120,
        "retention": 0.981554
       },
       "n5": {
        "chance_corrected": 0.695054,
        "n_labels": 30,
        "retention": 0.931311
       }
      },
      "score": 97.095518
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 14.691943,
      "worst_class_recall": 0.2891
     },
     "uncertainty": {
      "full_accuracy": 0.8355,
      "risk_coverage_auc": 0.96283,
      "score": 95.539596,
      "selective_acc_at_50": 0.987,
      "selective_acc_at_80": 0.92375
     }
    },
    "domain": "earth",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 3769,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / multispectral",
    "n_classes": 6,
    "n_test": 2000,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (Statlog sat.trn 4435 / sat.tst 2000), scored whole, with validation carved from sat.trn only (stratified 15%, seed 20260727). LEAKS, AND THE ACCURACIES HERE ARE OPTIMISTIC BECAUSE OF IT: each row is a 3x3 ground window, and 1,939 of the 2,000 test windows (97%) have an immediate neighbour in the training file sharing six of their nine ground pixels -- the two files are interleaved windows of one 82x100 scene, not separated regions. No 36-vector appears in both files, so byte-level deduplication passes while the test set overlaps training data on the ground; byte disjointness is not independence. Measured in variation/modalities/multispectral.md section 8. Raw class codes are 1,2,3,4,5,7 -- code 6 was withdrawn by the donors -- and are mapped positionally onto 0..5, not by subtracting one.",
    "task": "landsat_statlog",
    "task_name": "Landsat satellite land cover (6-class)",
    "timestamp": "2026-08-03T16:02:36+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 84.713377
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.914377,
      "auc": 0.980323,
      "balanced_accuracy": 0.899769,
      "ci95": [
       0.892369,
       0.908635
      ],
      "n_test": 9752,
      "score": 79.953726
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.914377,
      "brier": 0.095536,
      "ece": 0.032848,
      "mean_confidence": 0.936639,
      "overconfidence": 0.022263,
      "score": 83.575984
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.828612,
      "mean_agreement": 0.945648,
      "mean_kappa": 0.805737,
      "mean_retention": 0.884259,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.99,
         0.9808,
         0.9608
        ],
        "mean_kappa": 0.934114,
        "mean_retention": 0.98611,
        "retention_by_severity": [
         1,
         1,
         0.9583
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         1,
         0.9492,
         0.9375
        ],
        "mean_kappa": 0.892998,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9767,
         0.9375,
         0.7783
        ],
        "mean_kappa": 0.590098,
        "mean_retention": 0.666667,
        "retention_by_severity": [
         1,
         1,
         0
        ],
        "scored": true
       }
      },
      "score": 88.425888,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.799537,
      "mean_retention": 0.968404,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.971108,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.820101,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.723751,
        "n_labels": 10,
        "retention": 0.905212
       }
      },
      "score": 96.840405
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 74.914592,
      "worst_class_recall": 0.874573
     },
     "uncertainty": {
      "full_accuracy": 0.914377,
      "risk_coverage_auc": 0.992062,
      "score": 98.412349,
      "selective_acc_at_50": 0.99959,
      "selective_acc_at_80": 0.995642
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 9752,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 2 (datatest2.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection_2",
    "task_name": "Room occupancy from sensors (period 2)",
    "timestamp": "2026-08-01T19:15:19+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 92.79503
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.974859,
      "auc": 0.992121,
      "balanced_accuracy": 0.977365,
      "ci95": [
       0.971977,
       0.982285
      ],
      "n_test": 2665,
      "score": 95.472947
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.974859,
      "brier": 0.041825,
      "ece": 0.020148,
      "mean_confidence": 0.960381,
      "overconfidence": -0.014478,
      "score": 89.92598
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.948521,
      "mean_agreement": 0.951574,
      "mean_kappa": 0.873249,
      "mean_retention": 0.88191,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.995,
         0.9883,
         0.9675
        ],
        "mean_kappa": 0.964616,
        "mean_retention": 0.978229,
        "retention_by_severity": [
         1,
         0.9911,
         0.9436
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         1,
         0.9942,
         0.9967
        ],
        "mean_kappa": 0.993464,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9967,
         0.995,
         0.6308
        ],
        "mean_kappa": 0.661668,
        "mean_retention": 0.667501,
        "retention_by_severity": [
         1,
         1,
         0.0025
        ],
        "scored": true
       }
      },
      "score": 88.191024,
      "worst_retention": 0.002504
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.954729,
      "mean_retention": 0.929659,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.96294,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.895604,
        "n_labels": 40,
        "retention": 0.938071
       },
       "n5": {
        "chance_corrected": 0.812385,
        "n_labels": 10,
        "retention": 0.850906
       }
      },
      "score": 92.965901
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 93.620791,
      "worst_class_recall": 0.968104
     },
     "uncertainty": {
      "full_accuracy": 0.974859,
      "risk_coverage_auc": 0.99719,
      "score": 99.438032,
      "selective_acc_at_50": 1,
      "selective_acc_at_80": 0.99531
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 2665,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 1 (datatest.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection",
    "task_name": "Room occupancy from sensors (period 1)",
    "timestamp": "2026-08-01T19:15:16+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.58,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 21.202385
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.92602,
      "auc": 0.608696,
      "balanced_accuracy": 0.493207,
      "ci95": [
       0.486554,
       0.49863
      ],
      "n_test": 392,
      "score": 0
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.92602,
      "brier": 0.149519,
      "ece": 0.076371,
      "mean_confidence": 0.996369,
      "overconfidence": 0.070348,
      "score": 61.814278
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 392,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 0,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9872,
         0.9668,
         0.9388
        ],
        "mean_kappa": 0.081303,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9949,
         0.9949,
         0.9643
        ],
        "mean_kappa": 0.579204,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9974,
         0.9974,
         0.0051
        ],
        "mean_kappa": 0.584876,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0,
      "mean_retention": 0,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.124094,
        "n_labels": 176,
        "retention": 0
       },
       "n20": {
        "chance_corrected": 0,
        "n_labels": 40,
        "retention": 0
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 0
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.92602,
      "risk_coverage_auc": 0.926153,
      "score": 85.230543,
      "selective_acc_at_50": 0.943878,
      "selective_acc_at_80": 0.933121
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 940,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / process sensors",
    "n_classes": 2,
    "n_test": 392,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "chronological split minted by MedEval-1 on the production timestamp shipped in secom_labels.data: earliest 60% train, next 15% validation, latest 25% test. A random split is not defensible on a process-monitoring corpus -- tools drift and are recalibrated, so shuffling lets the model interpolate a tool's state from wafers measured minutes on either side, while production always extrapolates forward. Missing values imputed with the TRAIN median only; channels all-missing or constant on the training period dropped (468 of 590 channels kept).",
    "task": "secom",
    "task_name": "Semiconductor fabrication pass/fail",
    "timestamp": "2026-08-01T19:13:53+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 67.058277
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.713992,
      "auc": 0.932632,
      "balanced_accuracy": 0.746585,
      "ci95": [
       0.700104,
       0.79283
      ],
      "n_test": 486,
      "score": 70.434925
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.713992,
      "brier": 0.405677,
      "ece": 0.067992,
      "mean_confidence": 0.70528,
      "overconfidence": -0.008712,
      "score": 66.004218
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.704349,
      "mean_agreement": 0.758116,
      "mean_kappa": 0.671006,
      "mean_retention": 0.661451,
      "n_cases": 486,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9753,
         0.9033,
         0.7551
        ],
        "mean_kappa": 0.843901,
        "mean_retention": 0.868935,
        "retention_by_severity": [
         0.9531,
         0.9134,
         0.7403
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.7366,
         0.6667,
         0.535
        ],
        "mean_kappa": 0.514781,
        "mean_retention": 0.486918,
        "retention_by_severity": [
         0.7054,
         0.5143,
         0.241
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9774,
         0.7305,
         0.5432
        ],
        "mean_kappa": 0.654336,
        "mean_retention": 0.628501,
        "retention_by_severity": [
         0.9859,
         0.6761,
         0.2235
        ],
        "scored": true
       }
      },
      "score": 66.145126,
      "worst_retention": 0.223531
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.704349,
      "mean_retention": 0.925794,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.723243,
        "n_labels": 571,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.635584,
        "n_labels": 140,
        "retention": 0.902371
       },
       "n5": {
        "chance_corrected": 0.616314,
        "n_labels": 35,
        "retention": 0.875012
       }
      },
      "score": 92.579436
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "steel_type": {
        "gap": 0.182203,
        "per_group": {
         "A300": 0.435794,
         "A400": 0.617997
        },
        "worst": 0.435794
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 34.175968,
      "worst_class_recall": 0.564356
     },
     "uncertainty": {
      "full_accuracy": 0.713992,
      "risk_coverage_auc": 0.849393,
      "score": 82.429163,
      "selective_acc_at_50": 0.82716,
      "selective_acc_at_80": 0.771208
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 1164,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 486,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. The label is the argmax of the seven trailing one-hot fault columns, which are dropped from the features.",
    "task": "steel_plate_faults",
    "task_name": "Steel plate surface faults (7-class)",
    "timestamp": "2026-08-01T19:13:43+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 59.705537
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.75,
      "auc": 0.815923,
      "balanced_accuracy": 0.751934,
      "ci95": [
       0.709575,
       0.799232
      ],
      "n_test": 400,
      "score": 50.386896
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.75,
      "brier": 0.346947,
      "ece": 0.061364,
      "mean_confidence": 0.75821,
      "overconfidence": 0.00821,
      "score": 69.318061
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.503869,
      "mean_agreement": 0.819722,
      "mean_kappa": 0.637728,
      "mean_retention": 0.708339,
      "n_cases": 400,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9775,
         0.9525,
         0.875
        ],
        "mean_kappa": 0.869658,
        "mean_retention": 0.970815,
        "retention_by_severity": [
         0.9949,
         0.9963,
         0.9212
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9025,
         0.73,
         0.6975
        ],
        "mean_kappa": 0.553853,
        "mean_retention": 0.598458,
        "retention_by_severity": [
         0.8669,
         0.5088,
         0.4197
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.96,
         0.7725,
         0.51
        ],
        "mean_kappa": 0.489673,
        "mean_retention": 0.555744,
        "retention_by_severity": [
         1,
         0.6672,
         0
        ],
        "scored": true
       }
      },
      "score": 70.833887,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.503869,
      "mean_retention": 0.630967,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.518842,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.449905,
        "n_labels": 40,
        "retention": 0.8929
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 63.096663
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 44.859813,
      "worst_class_recall": 0.724299
     },
     "uncertainty": {
      "full_accuracy": 0.75,
      "risk_coverage_auc": 0.844106,
      "score": 68.821262,
      "selective_acc_at_50": 0.825,
      "selective_acc_at_80": 0.796875
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 959,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 400,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (red wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_red",
    "task_name": "Wine quality, red (binarised)",
    "timestamp": "2026-08-01T19:13:41+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 52.273715
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.757551,
      "auc": 0.827081,
      "balanced_accuracy": 0.691134,
      "ci95": [
       0.664891,
       0.716687
      ],
      "n_test": 1225,
      "score": 38.226844
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.757551,
      "brier": 0.315442,
      "ece": 0.028736,
      "mean_confidence": 0.752467,
      "overconfidence": -0.005084,
      "score": 85.631979
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.38084,
      "mean_agreement": 0.803611,
      "mean_kappa": 0.538934,
      "mean_retention": 0.637729,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9875,
         0.935,
         0.845
        ],
        "mean_kappa": 0.805296,
        "mean_retention": 0.938348,
        "retention_by_severity": [
         0.9867,
         0.9268,
         0.9015
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.89,
         0.7867,
         0.7808
        ],
        "mean_kappa": 0.311528,
        "mean_retention": 0.318974,
        "retention_by_severity": [
         0.7204,
         0.0906,
         0.1459
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9783,
         0.7942,
         0.235
        ],
        "mean_kappa": 0.499976,
        "mean_retention": 0.655867,
        "retention_by_severity": [
         0.9676,
         1,
         0
        ],
        "scored": true
       }
      },
      "score": 63.772944,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.382268,
      "mean_retention": 0.666667,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.410205,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.454601,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 66.666667
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.490244
     },
     "uncertainty": {
      "full_accuracy": 0.757551,
      "risk_coverage_auc": 0.883347,
      "score": 76.669329,
      "selective_acc_at_50": 0.897059,
      "selective_acc_at_80": 0.811224
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 2938,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 1225,
    "runtime_s": 0.1,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (white wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_white",
    "task_name": "Wine quality, white (binarised)",
    "timestamp": "2026-08-01T19:14:08+00:00"
   }
  ],
  "heldout": null,
  "sources": [
   {
    "path": "results/cross_site/occupancy__tab_logreg.json",
    "sha256": "39f6dd624d9723016056f3deb5caf4f5be5c852ab6ec756c45afd28732dd3d52"
   },
   {
    "path": "results/heldout/audit_report.json",
    "sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
   },
   {
    "path": "results/heldout/cycle.json",
    "sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
   },
   {
    "path": "results/heldout/manifests/gen1.seal.json",
    "sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
   },
   {
    "path": "results/heldout/manifests/gen2.seal.json",
    "sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
   },
   {
    "path": "results/heldout/manifests/public.seal.json",
    "sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
   },
   {
    "path": "results/records/adult_income__tab_logreg.json",
    "sha256": "9d49f868e37c94bdcbde73d8ea4f043ccd8d8fc63c89550f8e2d95605e67deb0"
   },
   {
    "path": "results/records/anuran_mfcc__tab_logreg.json",
    "sha256": "c1271e3ccaa18a559772a12ca02fb9265413e70bb06f299c2c87b78c9f5ac6f0"
   },
   {
    "path": "results/records/breast_wdbc__tab_logreg.json",
    "sha256": "9756b728ad8a36e7bd2806f46f3fb6b9c2bd8590a97bcf5c22fc66f28fb165be"
   },
   {
    "path": "results/records/coil2000_caravan__tab_logreg.json",
    "sha256": "252f17a68d072d2cb05ee5514df2f21bdf1d6caa45904c07260ed5c59778507b"
   },
   {
    "path": "results/records/diabetes_pima__tab_logreg.json",
    "sha256": "d8c33c2b3ecb291523db6e86a5fb15de72eab3df9c9f8050abf5a97a862441d3"
   },
   {
    "path": "results/records/dry_bean__tab_logreg.json",
    "sha256": "746ef76d9275ef54925dba9afd291e7dcc4986601dfbf86802325ca2bffa87ee"
   },
   {
    "path": "results/records/german_credit__tab_logreg.json",
    "sha256": "ef0656154fcaed4e3ec40dfe3f52387aed341910b42f970cb4feb3ea9a0dcdd9"
   },
   {
    "path": "results/records/heart_cleveland__tab_logreg.json",
    "sha256": "ee6774fabcb5b6f03727dfba90c2b3081f6f98b06172ecfc2f9e4241b3432ee2"
   },
   {
    "path": "results/records/landsat_statlog__tab_logreg.json",
    "sha256": "2fe08d5aceaf2cfc8e5eb9cbdf58169f99876093ad34aea9e16de93890ee949f"
   },
   {
    "path": "results/records/occupancy_detection_2__tab_logreg.json",
    "sha256": "33859c9106ad2179d4b479a3b53c112f6e1e36dfd7324b915658e8ab4f751bce"
   },
   {
    "path": "results/records/occupancy_detection__tab_logreg.json",
    "sha256": "7a7adae2920572b0569780fdcc65883daabc66360174c0989ebfb3727800090d"
   },
   {
    "path": "results/records/secom__tab_logreg.json",
    "sha256": "5e2c3e5e4a5dea390dd9968e2121956ee3e704ff150bd3a93301bd52d299e797"
   },
   {
    "path": "results/records/steel_plate_faults__tab_logreg.json",
    "sha256": "bcb35954c6c21b3a1a223d12a924f3ff3ca1cebc8016823872b1c305aa3b4e6b"
   },
   {
    "path": "results/records/wine_quality_red__tab_logreg.json",
    "sha256": "528c7eedc18d0baf7620c8f7ba6e36991a41f7675d52c150ca92ce4ae23452fe"
   },
   {
    "path": "results/records/wine_quality_white__tab_logreg.json",
    "sha256": "a1caccca58fe53897a0124cf62fc019fe3b4bc6b3a15aee3d5c22da9d3460e68"
   }
  ]
 },
 "declared": null,
 "declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
 "expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
 "expiry_utc": "2026-09-20T00:00:00Z",
 "issued_utc": "2026-08-03T16:03:42Z",
 "issuer": {
  "algo": "Ed25519",
  "key_id": "ns-passport-2026-07",
  "name": "NakedSignal",
  "public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
 },
 "observed": null,
 "observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
 "passport_version": "0.1",
 "signature": "tn7ly4md5fR34HyiyAR06s1iIi2FMO5XyDGxSdbQo7oNEkv7iVBFGdtTBpbHgSUJukFh0VOYCZoZLBxLlsDxCA==",
 "spec_version": "MedEval-1 v0.1",
 "subject": {
  "behaviour_version": "1.0.0",
  "code_fingerprint": null,
  "code_fingerprints": [
   "b444c02112cc",
   "de32bca72d56"
  ],
  "family": "classical",
  "model_id": "tab_logreg",
  "model_name": "Logistic regression",
  "params": "standardised, L2, C=1"
 }
}

Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.