NakedSignal OS · Documentation
Back to the consoleDocumentation

Random forest

LIVE

The signed evidence record for tab_rf. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.

SIGNATURE VALIDexit 0signature valid, issued by NakedSignal under key ns-passport-2026-07, not expired
SubjectRandom forest tab_rf
Configuration400 trees
Code fingerprintb444c02112cc, de32bca72d56
Spec versionMedEval-1 v0.1
Issued2026-08-03T16:03:42Z by NakedSignal
Expires2026-09-20T00:00:00Zthe sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first
key ns-passport-2026-07 · Ed25519

The check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.

Measured by NakedSignal

Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 21 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.

TaskDomainn testaccuracycalibrationuncertaintysubgroupcorruptionacquisitionlimited datacostspec sensitivityIndex
adult_incomelabour1628152.76495585.02509989.60847639.2481283.81681n/a89.953523n/an/a69.440867
anuran_mfccecology99574.72148127.16331784.01542375.60975683.574454n/a87.122211n/an/a70.002645
breast_wdbcmedicine14389.11949783.02447698.93157684.9056686.907656n/a98.376853n/an/a88.713661
coil2000_caravanfinancial40005.07976785.47225393.025273031.47343n/a83.057034n/an/a38.793123
diabetes_pimamedicine19337.66879968.11528573.39298633.63636497.4703n/a92.955975n/an/a62.857956
dry_beanagriculture340493.06314592.63880798.90486887.4304588.317564n/a93.273468n/an/a91.710502
german_creditfinancial25022.85714372.20575.04518321.29629694.62963n/a68.333333n/an/a54.183229
heart_clevelandmedicine7565.35714348.5590.23198259.86206977.717061n/a88.342441n/an/a68.243974
landsat_statlogearth200086.74318669.49187598.13305155.6398184.792489n/a88.843692n/an/a79.953503
occupancy_detection_2energy975294.90912210.7868998.19039791.37998281.645552n/a89.682104n/an/a76.298286
occupancy_detectionenergy266589.93264484.85506698.18871486.0082385.305365n/a94.718527n/an/a88.676156
secomindustrial392066.24043487.0637040not scoredn/a0n/an/a22.352486
steel_plate_faultsindustrial48676.0138854.83796391.87663548.19550172.442213n/a89.679257n/an/a70.252251
wine_quality_redagriculture40060.97377173.22812578.24338654.20560780.266447n/a64.639473n/an/a67.604273
wine_quality_whiteagriculture122562.7592474.48469488.85766340.48780563.894206n/a61.362516n/an/a63.660968

Cross-site retention · occupancy

retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column

Sitesoccupancy_detection, occupancy_detection_2
Mean retention0.973783
Worst retention0.947566

Private held-out track

this subject was not submitted to a sealed generation.

Known limitations

Subgroup metadataabsent on 9 of 15 corpora, so the subgroup score falls back to the worst diagnostic class. That absence is a disclosure, not a pass.
Training set cappedyes on at least one task — see the raw JSON for which
Observed in deploymentnothing — see the third panel

Audit chain

Bundles re-derived48 of 48
Ledger100 entries, head 6b7b371ff650a8ed, chain intact
Re-run it yourselfpython src/heldout/audit.py 1f66c7719ab3943c6fcc
21 source files, with hashes
results/cross_site/occupancy__tab_rf.jsona405d2c58de366fe1d1affa19d29be533b62395b8de360dd8e1b9d63f0ee1713
results/heldout/audit_report.jsondf460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2
results/heldout/cycle.jsone012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502
results/heldout/manifests/gen1.seal.json91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81
results/heldout/manifests/gen2.seal.jsonf61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547
results/heldout/manifests/public.seal.json1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c
results/records/adult_income__tab_rf.json0c72895bc452f76c94279742836ab08f6a3b90a995b96b35d1fc1ba390a5e3da
results/records/anuran_mfcc__tab_rf.json2aa16d343d4791780b6688e3cbdec9935e8bc1207243f0c4a3d814767a88dc3e
results/records/breast_wdbc__tab_rf.jsonb7fc1d6a57687ecba17348b2d8fbbea4f67ee453bb423b3148d82f67be2ffb52
results/records/coil2000_caravan__tab_rf.json64ae58d98289e58681d35676afc9effd95649504396200c830fae76b85b36f91
results/records/diabetes_pima__tab_rf.jsonaa7459b59b46fc8a1d1624f4d7c8ea2b1e7c4acd75206c45c5e3ab218477ef11
results/records/dry_bean__tab_rf.json6b7012e82fb4d774eb7443cf3e5b6b335484b897b5be188ec49212af3ed727a4
results/records/german_credit__tab_rf.json24ff6f5c6838f28b152a67cef79296d839893be12124bb44b60f2624dcb22c65
results/records/heart_cleveland__tab_rf.json4e3b6e895abda7ad0c9377c4fdf3c85e91e061d7fb2705de140ae2b116a2fb4f
results/records/landsat_statlog__tab_rf.json62fa9186b3120e32280b31f0b83e1a03855fe9020b74ac04bd2bdf8d6d2d67f9
results/records/occupancy_detection_2__tab_rf.json7e81c304359ec6370e7e9464502c0f66c2368e2b5b036e558c8c5ffd35d7316e
results/records/occupancy_detection__tab_rf.jsonfe8a6aea38b07daf1ae9f259f3a6e3eb1783d625afd1bfe80b95c96ea988872f
results/records/secom__tab_rf.json6450800d5380b41fabcb5512ad4df69fee970347b54814ad5793ca20f71b3947
results/records/steel_plate_faults__tab_rf.json34a1824eb2b617a095e35ef2f305eea48af27d9386eb7cb39713abaf85847083
results/records/wine_quality_red__tab_rf.jsonf02c6fee1004ba93bcc07253fb83e41177515660541d7f7b5a2e3687131473c2
results/records/wine_quality_white__tab_rf.json16bcf923dd4d658453df5e051aa97fb67046704ce9379fb9e22c380c88045ba6
Declared by vendor

Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.

No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.

Observed in deployment

Production history and drift — how the model has actually behaved since it was deployed, on real traffic.

Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.

Raw JSON — the whole signed document, 76,621 characters
{
 "canonicalisation": {
  "float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
  "form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
  "signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
 },
 "computed": {
  "audit": {
   "audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
   "bundles": 48,
   "ledger_entries": 100,
   "ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
   "ledger_intact": true,
   "rederived": 48
  },
  "cross_site": {
   "code_fingerprint": "b444c02112cc",
   "device": "mps",
   "family": "occupancy",
   "matrix": {
    "occupancy_detection": {
     "occupancy_detection": {
      "balanced_accuracy": 0.949663,
      "chance_corrected": 0.899326,
      "ece": 0.03029
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.974546,
      "chance_corrected": 0.949091,
      "ece": 0.178426,
      "retention": 1
     }
    },
    "occupancy_detection_2": {
     "occupancy_detection": {
      "balanced_accuracy": 0.949663,
      "chance_corrected": 0.899326,
      "ece": 0.03029,
      "retention": 0.947566
     },
     "occupancy_detection_2": {
      "balanced_accuracy": 0.974546,
      "chance_corrected": 0.949091,
      "ece": 0.178426
     }
    }
   },
   "mean_retention": 0.973783,
   "note": "retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column",
   "score": 97.378293,
   "site_labels": {
    "occupancy_detection": "test period 1",
    "occupancy_detection_2": "test period 2"
   },
   "sites": [
    "occupancy_detection",
    "occupancy_detection_2"
   ],
   "spec_version": "MedEval-1 v0.1",
   "worst_retention": 0.947566
  },
  "evaluations": [
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 69.440867
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.844297,
      "auc": 0.889154,
      "balanced_accuracy": 0.763825,
      "ci95": [
       0.757015,
       0.772402
      ],
      "n_test": 16281,
      "score": 52.764955
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.844297,
      "brier": 0.219581,
      "ece": 0.02995,
      "mean_confidence": 0.87282,
      "overconfidence": 0.028523,
      "score": 85.025099
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.555881,
      "mean_agreement": 0.937037,
      "mean_kappa": 0.784667,
      "mean_retention": 0.838168,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9908,
         0.97,
         0.9525
        ],
        "mean_kappa": 0.915722,
        "mean_retention": 0.974271,
        "retention_by_severity": [
         1,
         0.9667,
         0.9561
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.94,
         0.8875,
         0.8167
        ],
        "mean_kappa": 0.557807,
        "mean_retention": 0.608112,
        "retention_by_severity": [
         0.9079,
         0.6445,
         0.2719
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.995,
         0.9883,
         0.8925
        ],
        "mean_kappa": 0.880471,
        "mean_retention": 0.932122,
        "retention_by_severity": [
         0.9835,
         0.967,
         0.8458
        ],
        "scored": true
       }
      },
      "score": 83.81681,
      "worst_retention": 0.271856
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.52765,
      "mean_retention": 0.899535,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.584809,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.532786,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.368619,
        "n_labels": 10,
        "retention": 0.698606
       }
      },
      "score": 89.953523
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.011916,
        "per_group": {
         "45-59": 0.754575,
         "60+": 0.745397,
         "<45": 0.757312
        },
        "worst": 0.745397
       },
       "race": {
        "gap": 0.082085,
        "per_group": {
         "Amer-Indian-Eskimo": 0.696241,
         "Asian-Pac-Islander": 0.778325,
         "Black": 0.75053,
         "Other": 0.717273,
         "White": 0.762212
        },
        "worst": 0.696241
       },
       "sex": {
        "gap": 0.005315,
        "per_group": {
         "Female": 0.757978,
         "Male": 0.752662
        },
        "worst": 0.752662
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 39.24812,
      "worst_class_recall": 0.611284
     },
     "uncertainty": {
      "full_accuracy": 0.844297,
      "risk_coverage_auc": 0.948042,
      "score": 89.608476,
      "selective_acc_at_50": 0.96683,
      "selective_acc_at_80": 0.907946
     }
    },
    "domain": "labour",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 27676,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": true
    },
    "modality": "Tabular / socioeconomic",
    "n_classes": 2,
    "n_test": 16281,
    "runtime_s": 1.8,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (UCI adult.test scored whole); validation carved from adult.data only (stratified 15%, seed 20260727)",
    "task": "adult_income",
    "task_name": "Census income (>$50K)",
    "timestamp": "2026-08-01T19:16:09+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 70.002645
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.590955,
      "auc": 0.989563,
      "balanced_accuracy": 0.772493,
      "ci95": [
       0.736747,
       0.794851
      ],
      "n_test": 995,
      "score": 74.721481
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.590955,
      "brier": 0.50018,
      "ece": 0.145673,
      "mean_confidence": 0.533367,
      "overconfidence": -0.057588,
      "score": 27.163317
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.747215,
      "mean_agreement": 0.784366,
      "mean_kappa": 0.752907,
      "mean_retention": 0.835745,
      "n_cases": 995,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9678,
         0.9397,
         0.8754
        ],
        "mean_kappa": 0.91344,
        "mean_retention": 0.988336,
        "retention_by_severity": [
         0.9953,
         0.9944,
         0.9753
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.804,
         0.6583,
         0.1467
        ],
        "mean_kappa": 0.477393,
        "mean_retention": 0.561256,
        "retention_by_severity": [
         0.9149,
         0.705,
         0.0639
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9769,
         0.9357,
         0.7548
        ],
        "mean_kappa": 0.867887,
        "mean_retention": 0.957642,
        "retention_by_severity": [
         1,
         1,
         0.8729
        ],
        "scored": true
       }
      },
      "score": 83.574454,
      "worst_retention": 0.063933
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.747215,
      "mean_retention": 0.871222,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.790268,
        "n_labels": 794,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.74041,
        "n_labels": 200,
        "retention": 0.990893
       },
       "n5": {
        "chance_corrected": 0.465345,
        "n_labels": 50,
        "retention": 0.622773
       }
      },
      "score": 87.122211
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "family": {
        "gap": 0.109034,
        "per_group": {
         "Hylidae": 0.889522,
         "Leptodactylidae": 0.780488
        },
        "worst": 0.780488
       },
       "genus": {
        "gap": null,
        "per_group": {
         "Hypsiboas": 0.960862
        },
        "worst": null
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 75.609756,
      "worst_class_recall": 0.160883
     },
     "uncertainty": {
      "full_accuracy": 0.590955,
      "risk_coverage_auc": 0.856139,
      "score": 84.015423,
      "selective_acc_at_50": 0.925703,
      "selective_acc_at_80": 0.703518
     }
    },
    "domain": "ecology",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 5073,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Bioacoustic / MFCC",
    "n_classes": 10,
    "n_test": 995,
    "runtime_s": 1.8,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "no official split; minted at the standard seed (20260727) by partitioning the 60 RecordID recording groups 60/15/15/25 and taking whole recordings, so no recording spans two splits. A row-level random split would place near-duplicate syllables from one recording on both sides and score memorisation (~99%); class balance is therefore uncontrolled across splits and is not stratified",
    "task": "anuran_mfcc",
    "task_name": "Anuran call species (10-class)",
    "timestamp": "2026-08-01T19:14:27+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 88.713661
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.951049,
      "auc": 0.989623,
      "balanced_accuracy": 0.945597,
      "ci95": [
       0.898492,
       0.979619
      ],
      "n_test": 143,
      "score": 89.119497
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.951049,
      "brier": 0.064681,
      "ece": 0.033951,
      "mean_confidence": 0.940944,
      "overconfidence": -0.010105,
      "score": 83.024476
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.891195,
      "mean_agreement": 0.903652,
      "mean_kappa": 0.834613,
      "mean_retention": 0.869077,
      "n_cases": 143,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.986,
         0.972
        ],
        "mean_kappa": 0.969945,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.993,
         0.979,
         0.958
        ],
        "mean_kappa": 0.949588,
        "mean_retention": 0.973418,
        "retention_by_severity": [
         1,
         0.9962,
         0.924
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.993,
         0.8811,
         0.3706
        ],
        "mean_kappa": 0.584306,
        "mean_retention": 0.633812,
        "retention_by_severity": [
         1,
         0.889,
         0.0125
        ],
        "scored": true
       }
      },
      "score": 86.907656,
      "worst_retention": 0.012468
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.891195,
      "mean_retention": 0.983769,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.895597,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.847799,
        "n_labels": 40,
        "retention": 0.951306
       },
       "n5": {
        "chance_corrected": 0.925577,
        "n_labels": 10,
        "retention": 1
       }
      },
      "score": 98.376853
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 84.90566,
      "worst_class_recall": 0.924528
     },
     "uncertainty": {
      "full_accuracy": 0.951049,
      "risk_coverage_auc": 0.994658,
      "score": 98.931576,
      "selective_acc_at_50": 1,
      "selective_acc_at_80": 0.991228
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 341,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 2,
    "n_test": 143,
    "runtime_s": 1.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "breast_wdbc",
    "task_name": "Breast cytology (WDBC)",
    "timestamp": "2026-08-01T19:13:37+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 38.793123
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.93275,
      "auc": 0.679491,
      "balanced_accuracy": 0.525399,
      "ci95": [
       0.512788,
       0.540416
      ],
      "n_test": 4000,
      "score": 5.079767
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.93275,
      "brier": 0.119531,
      "ece": 0.029055,
      "mean_confidence": 0.936037,
      "overconfidence": 0.003287,
      "score": 85.472253
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.058154,
      "mean_agreement": 0.989907,
      "mean_kappa": 0.453368,
      "mean_retention": 0.314734,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9992,
         0.9983,
         0.9917
        ],
        "mean_kappa": 0.861544,
        "mean_retention": 0.621014,
        "retention_by_severity": [
         0.9848,
         0.5087,
         0.3696
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9858,
         0.9833,
         0.9833
        ],
        "mean_kappa": 0.085881,
        "mean_retention": 0,
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9992,
         0.9842,
         0.9842
        ],
        "mean_kappa": 0.41268,
        "mean_retention": 0.323188,
        "retention_by_severity": [
         0.7543,
         0.2152,
         0
        ],
        "scored": true
       }
      },
      "score": 31.47343,
      "worst_retention": 0
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.050798,
      "mean_retention": 0.83057,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.281783,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.081358,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.024978,
        "n_labels": 10,
        "retention": 0.491711
       }
      },
      "score": 83.057034
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0.063025
     },
     "uncertainty": {
      "full_accuracy": 0.93275,
      "risk_coverage_auc": 0.965126,
      "score": 93.025273,
      "selective_acc_at_50": 0.9685,
      "selective_acc_at_80": 0.954688
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 4948,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / insurance",
    "n_classes": 2,
    "n_test": 4000,
    "runtime_s": 1.6,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (CoIL Challenge 2000: ticdata2000.txt 5822 train / ticeval2000.txt 4000 test). The test labels were released separately, in tictgts2000.txt, only after the challenge closed -- they were not available to anyone tuning on this benchmark at the time. Test scored whole; validation carved from ticdata2000.txt only (stratified 15%, seed 20260727).",
    "task": "coil2000_caravan",
    "task_name": "Caravan insurance purchase (CoIL 2000)",
    "timestamp": "2026-08-01T19:15:22+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 62.857956
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.725389,
      "auc": 0.795783,
      "balanced_accuracy": 0.688344,
      "ci95": [
       0.623281,
       0.75972
      ],
      "n_test": 193,
      "score": 37.668799
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.725389,
      "brier": 0.350497,
      "ece": 0.063769,
      "mean_confidence": 0.778148,
      "overconfidence": 0.052759,
      "score": 68.115285
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.376688,
      "mean_agreement": 0.926885,
      "mean_kappa": 0.84281,
      "mean_retention": 0.974703,
      "n_cases": 193,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.9534,
         0.9326
        ],
        "mean_kappa": 0.912054,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9637,
         0.9378,
         0.9016
        ],
        "mean_kappa": 0.844912,
        "mean_retention": 0.924109,
        "retention_by_severity": [
         0.9418,
         0.9443,
         0.8862
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9741,
         0.943,
         0.7358
        ],
        "mean_kappa": 0.771464,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 97.4703,
      "worst_retention": 0.886164
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.376688,
      "mean_retention": 0.92956,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.362118,
        "n_labels": 200,
        "retention": 0.961321
       },
       "n20": {
        "chance_corrected": 0.424662,
        "n_labels": 40,
        "retention": 1
       },
       "n5": {
        "chance_corrected": 0.311656,
        "n_labels": 10,
        "retention": 0.827358
       }
      },
      "score": 92.955975
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.018182,
        "per_group": {
         "45-59": 0.668182,
         "<45": 0.686364
        },
        "worst": 0.668182
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 33.636364,
      "worst_class_recall": 0.567164
     },
     "uncertainty": {
      "full_accuracy": 0.725389,
      "risk_coverage_auc": 0.866965,
      "score": 73.392986,
      "selective_acc_at_50": 0.875,
      "selective_acc_at_80": 0.75974
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 460,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 193,
    "runtime_s": 1.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "diabetes_pima",
    "task_name": "Diabetes onset (Pima)",
    "timestamp": "2026-08-01T19:13:35+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 91.710502
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.931257,
      "auc": 0.994913,
      "balanced_accuracy": 0.940541,
      "ci95": [
       0.932217,
       0.948192
      ],
      "n_test": 3404,
      "score": 93.063145
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.931257,
      "brier": 0.103726,
      "ece": 0.014722,
      "mean_confidence": 0.917785,
      "overconfidence": -0.013472,
      "score": 92.638807
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.932822,
      "mean_agreement": 0.875278,
      "mean_kappa": 0.846422,
      "mean_retention": 0.883176,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9842,
         0.9667,
         0.9375
        ],
        "mean_kappa": 0.954775,
        "mean_retention": 0.977572,
        "retention_by_severity": [
         0.9928,
         0.9842,
         0.9557
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9483,
         0.9492,
         0.7308
        ],
        "mean_kappa": 0.849801,
        "mean_retention": 0.898907,
        "retention_by_severity": [
         0.982,
         0.9902,
         0.7246
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.985,
         0.9258,
         0.45
        ],
        "mean_kappa": 0.734689,
        "mean_retention": 0.773048,
        "retention_by_severity": [
         0.999,
         0.9599,
         0.3602
        ],
        "scored": true
       }
      },
      "score": 88.317564,
      "worst_retention": 0.360234
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.930631,
      "mean_retention": 0.932735,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.905675,
        "n_labels": 700,
        "retention": 0.973184
       },
       "n20": {
        "chance_corrected": 0.885251,
        "n_labels": 140,
        "retention": 0.951236
       },
       "n5": {
        "chance_corrected": 0.813171,
        "n_labels": 35,
        "retention": 0.873784
       }
      },
      "score": 93.273468
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 87.43045,
      "worst_class_recall": 0.892261
     },
     "uncertainty": {
      "full_accuracy": 0.931257,
      "risk_coverage_auc": 0.990613,
      "score": 98.904868,
      "selective_acc_at_50": 0.99765,
      "selective_acc_at_80": 0.986779
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 8166,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 3404,
    "runtime_s": 2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. Parsed from the ARFF in the donor's zip, whose nominal class order is the class order used here.",
    "task": "dry_bean",
    "task_name": "Dry bean variety (7-class)",
    "timestamp": "2026-08-01T19:15:26+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 54.183229
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.748,
      "auc": 0.79901,
      "balanced_accuracy": 0.614286,
      "ci95": [
       0.560355,
       0.674563
      ],
      "n_test": 250,
      "score": 22.857143
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.748,
      "brier": 0.322187,
      "ece": 0.05559,
      "mean_confidence": 0.73835,
      "overconfidence": -0.00965,
      "score": 72.205
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.228571,
      "mean_agreement": 0.963111,
      "mean_kappa": 0.836647,
      "mean_retention": 0.946296,
      "n_cases": 250,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.996,
         0.992,
         0.98
        ],
        "mean_kappa": 0.949124,
        "mean_retention": 0.980556,
        "retention_by_severity": [
         1,
         1,
         0.9417
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.944,
         0.928,
         0.908
        ],
        "mean_kappa": 0.674033,
        "mean_retention": 0.858333,
        "retention_by_severity": [
         0.95,
         1,
         0.625
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         1,
         0.972,
         0.948
        ],
        "mean_kappa": 0.886785,
        "mean_retention": 1,
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": true
       }
      },
      "score": 94.62963,
      "worst_retention": 0.625
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.228571,
      "mean_retention": 0.683333,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.44381,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.226667,
        "n_labels": 40,
        "retention": 0.991667
       },
       "n5": {
        "chance_corrected": 0.013333,
        "n_labels": 10,
        "retention": 0.058333
       }
      },
      "score": 68.333333
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": 0.00047,
        "per_group": {
         "45-59": 0.606481,
         "<45": 0.606952
        },
        "worst": 0.606481
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 21.296296,
      "worst_class_recall": 0.28
     },
     "uncertainty": {
      "full_accuracy": 0.748,
      "risk_coverage_auc": 0.875226,
      "score": 75.045183,
      "selective_acc_at_50": 0.904,
      "selective_acc_at_80": 0.795
     }
    },
    "domain": "financial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 600,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / financial",
    "n_classes": 2,
    "n_test": 250,
    "runtime_s": 1.3,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "german_credit",
    "task_name": "Consumer credit risk (Statlog)",
    "timestamp": "2026-08-01T19:13:39+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 68.243974
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.826667,
      "auc": 0.922143,
      "balanced_accuracy": 0.826786,
      "ci95": [
       0.757069,
       0.901891
      ],
      "n_test": 75,
      "score": 65.357143
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.826667,
      "brier": 0.235851,
      "ece": 0.1029,
      "mean_confidence": 0.762767,
      "overconfidence": -0.0639,
      "score": 48.55
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.653571,
      "mean_agreement": 0.851852,
      "mean_kappa": 0.697847,
      "mean_retention": 0.777171,
      "n_cases": 75,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9867,
         0.9467,
         0.9067
        ],
        "mean_kappa": 0.892895,
        "mean_retention": 0.941712,
        "retention_by_severity": [
         0.9617,
         0.9945,
         0.8689
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.88,
         0.7333,
         0.5733
        ],
        "mean_kappa": 0.442949,
        "mean_retention": 0.497268,
        "retention_by_severity": [
         0.8525,
         0.4645,
         0.1749
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9733,
         0.8,
         0.8667
        ],
        "mean_kappa": 0.757696,
        "mean_retention": 0.892532,
        "retention_by_severity": [
         1,
         0.7596,
         0.918
        ],
        "scored": true
       }
      },
      "score": 77.717061,
      "worst_retention": 0.174863
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.653571,
      "mean_retention": 0.883424,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.621429,
        "n_labels": 178,
        "retention": 0.95082
       },
       "n20": {
        "chance_corrected": 0.635714,
        "n_labels": 40,
        "retention": 0.972678
       },
       "n5": {
        "chance_corrected": 0.475,
        "n_labels": 10,
        "retention": 0.726776
       }
      },
      "score": 88.342441
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "age_band": {
        "gap": null,
        "per_group": {
         "45-59": 0.825
        },
        "worst": null
       },
       "sex": {
        "gap": 0.10069,
        "per_group": {
         "female": 0.9,
         "male": 0.79931
        },
        "worst": 0.79931
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 59.862069,
      "worst_class_recall": 0.825
     },
     "uncertainty": {
      "full_accuracy": 0.826667,
      "risk_coverage_auc": 0.95116,
      "score": 90.231982,
      "selective_acc_at_50": 1,
      "selective_acc_at_80": 0.9
     }
    },
    "domain": "medicine",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 178,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / clinical",
    "n_classes": 2,
    "n_test": 75,
    "runtime_s": 1.2,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
    "task": "heart_cleveland",
    "task_name": "Coronary artery disease (Cleveland)",
    "timestamp": "2026-08-01T19:13:33+00:00"
   },
   {
    "code_fingerprint": "de32bca72d56",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 79.953503
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.91,
      "auc": 0.989742,
      "balanced_accuracy": 0.889527,
      "ci95": [
       0.873785,
       0.903103
      ],
      "n_test": 2000,
      "score": 86.743186
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.91,
      "brier": 0.14278,
      "ece": 0.061016,
      "mean_confidence": 0.849226,
      "overconfidence": -0.060774,
      "score": 69.491875
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.872609,
      "mean_agreement": 0.857963,
      "mean_kappa": 0.82291,
      "mean_retention": 0.847925,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9942,
         0.9825,
         0.9583
        ],
        "mean_kappa": 0.973131,
        "mean_retention": 0.982722,
        "retention_by_severity": [
         0.9978,
         0.9817,
         0.9687
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9542,
         0.7817,
         0.56
        ],
        "mean_kappa": 0.708097,
        "mean_retention": 0.749326,
        "retention_by_severity": [
         0.9727,
         0.7865,
         0.4887
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9667,
         0.8742,
         0.65
        ],
        "mean_kappa": 0.787501,
        "mean_retention": 0.811726,
        "retention_by_severity": [
         0.9714,
         0.8609,
         0.6029
        ],
        "scored": true
       }
      },
      "score": 84.792489,
      "worst_retention": 0.488722
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.867432,
      "mean_retention": 0.888437,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.833674,
        "n_labels": 600,
        "retention": 0.961083
       },
       "n20": {
        "chance_corrected": 0.783786,
        "n_labels": 120,
        "retention": 0.90357
       },
       "n5": {
        "chance_corrected": 0.694516,
        "n_labels": 30,
        "retention": 0.800658
       }
      },
      "score": 88.843692
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 55.63981,
      "worst_class_recall": 0.630332
     },
     "uncertainty": {
      "full_accuracy": 0.91,
      "risk_coverage_auc": 0.984442,
      "score": 98.133051,
      "selective_acc_at_50": 0.998,
      "selective_acc_at_80": 0.971875
     }
    },
    "domain": "earth",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 3769,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / multispectral",
    "n_classes": 6,
    "n_test": 2000,
    "runtime_s": 1.8,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split (Statlog sat.trn 4435 / sat.tst 2000), scored whole, with validation carved from sat.trn only (stratified 15%, seed 20260727). LEAKS, AND THE ACCURACIES HERE ARE OPTIMISTIC BECAUSE OF IT: each row is a 3x3 ground window, and 1,939 of the 2,000 test windows (97%) have an immediate neighbour in the training file sharing six of their nine ground pixels -- the two files are interleaved windows of one 82x100 scene, not separated regions. No 36-vector appears in both files, so byte-level deduplication passes while the test set overlaps training data on the ground; byte disjointness is not independence. Measured in variation/modalities/multispectral.md section 8. Raw class codes are 1,2,3,4,5,7 -- code 6 was withdrawn by the donors -- and are mapped positionally onto 0..5, not by subtracting one.",
    "task": "landsat_statlog",
    "task_name": "Landsat satellite land cover (6-class)",
    "timestamp": "2026-08-03T16:02:36+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 76.298286
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.964315,
      "auc": 0.993986,
      "balanced_accuracy": 0.974546,
      "ci95": [
       0.97196,
       0.977215
      ],
      "n_test": 9752,
      "score": 94.909122
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.964315,
      "brier": 0.132329,
      "ece": 0.178426,
      "mean_confidence": 0.788762,
      "overconfidence": -0.175553,
      "score": 10.78689
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.95307,
      "mean_agreement": 0.946759,
      "mean_kappa": 0.802162,
      "mean_retention": 0.816456,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9958,
         0.9875,
         0.97
        ],
        "mean_kappa": 0.956877,
        "mean_retention": 0.981797,
        "retention_by_severity": [
         0.9989,
         0.9923,
         0.9542
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9933,
         0.9858,
         0.77
        ],
        "mean_kappa": 0.660164,
        "mean_retention": 0.669132,
        "retention_by_severity": [
         0.9912,
         0.9866,
         0.0296
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9958,
         0.9783,
         0.8442
        ],
        "mean_kappa": 0.789446,
        "mean_retention": 0.798438,
        "retention_by_severity": [
         0.9945,
         1,
         0.4008
        ],
        "scored": true
       }
      },
      "score": 81.645552,
      "worst_retention": 0.029616
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.949091,
      "mean_retention": 0.896821,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.967213,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.875695,
        "n_labels": 40,
        "retention": 0.922667
       },
       "n5": {
        "chance_corrected": 0.728708,
        "n_labels": 10,
        "retention": 0.767796
       }
      },
      "score": 89.682104
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 91.379982,
      "worst_class_recall": 0.9569
     },
     "uncertainty": {
      "full_accuracy": 0.964315,
      "risk_coverage_auc": 0.990952,
      "score": 98.190397,
      "selective_acc_at_50": 0.994053,
      "selective_acc_at_80": 0.991156
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 9752,
    "runtime_s": 1.6,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 2 (datatest2.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection_2",
    "task_name": "Room occupancy from sensors (period 2)",
    "timestamp": "2026-08-01T19:15:19+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 88.676156
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.954972,
      "auc": 0.984933,
      "balanced_accuracy": 0.949663,
      "ci95": [
       0.939995,
       0.9574
      ],
      "n_test": 2665,
      "score": 89.932644
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.954972,
      "brier": 0.078528,
      "ece": 0.03029,
      "mean_confidence": 0.924682,
      "overconfidence": -0.03029,
      "score": 84.855066
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.880022,
      "mean_agreement": 0.931944,
      "mean_kappa": 0.818571,
      "mean_retention": 0.853054,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9942,
         0.9908,
         0.9783
        ],
        "mean_kappa": 0.972735,
        "mean_retention": 0.99876,
        "retention_by_severity": [
         1,
         1,
         0.9963
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         1,
         0.97,
         0.6625
        ],
        "mean_kappa": 0.649159,
        "mean_retention": 0.670266,
        "retention_by_severity": [
         1,
         1,
         0.0108
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9933,
         0.9617,
         0.8367
        ],
        "mean_kappa": 0.83382,
        "mean_retention": 0.890136,
        "retention_by_severity": [
         1,
         1,
         0.6704
        ],
        "scored": true
       }
      },
      "score": 85.305365,
      "worst_retention": 0.010797
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.899326,
      "mean_retention": 0.947185,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.963836,
        "n_labels": 200,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.817198,
        "n_labels": 40,
        "retention": 0.908678
       },
       "n5": {
        "chance_corrected": 0.838962,
        "n_labels": 10,
        "retention": 0.932878
       }
      },
      "score": 94.718527
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 86.00823,
      "worst_class_recall": 0.930041
     },
     "uncertainty": {
      "full_accuracy": 0.954972,
      "risk_coverage_auc": 0.990944,
      "score": 98.188714,
      "selective_acc_at_50": 0.991742,
      "selective_acc_at_80": 0.983583
     }
    },
    "domain": "energy",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 6922,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / environmental sensors",
    "n_classes": 2,
    "n_test": 2665,
    "runtime_s": 1.5,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 1 (datatest.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
    "task": "occupancy_detection",
    "task_name": "Room occupancy from sensors (period 1)",
    "timestamp": "2026-08-01T19:15:16+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.58,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 22.352486
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.936224,
      "auc": 0.447747,
      "balanced_accuracy": 0.498641,
      "ci95": [
       0.495902,
       0.5
      ],
      "n_test": 392,
      "score": 0
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.936224,
      "brier": 0.13808,
      "ece": 0.067519,
      "mean_confidence": 0.87479,
      "overconfidence": -0.061435,
      "score": 66.240434
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0,
      "mean_agreement": null,
      "mean_kappa": null,
      "mean_retention": null,
      "n_cases": 392,
      "not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 0,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         1,
         0.9974,
         0.9872
        ],
        "mean_kappa": 0.649369,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9974,
         0.9974,
         0.9974
        ],
        "mean_kappa": 0,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       },
       "unit_scale": {
        "agreement_by_severity": [
         1,
         1,
         0.5128
        ],
        "mean_kappa": 0.668438,
        "mean_retention": 0,
        "not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
        "retention_by_severity": [
         0,
         0,
         0
        ],
        "scored": false
       }
      },
      "score": null,
      "worst_retention": null
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0,
      "mean_retention": 0,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.014493,
        "n_labels": 176,
        "retention": 0
       },
       "n20": {
        "chance_corrected": 0.115036,
        "n_labels": 40,
        "retention": 0
       },
       "n5": {
        "chance_corrected": 0,
        "n_labels": 10,
        "retention": 0
       }
      },
      "score": 0
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 0,
      "worst_class_recall": 0
     },
     "uncertainty": {
      "full_accuracy": 0.936224,
      "risk_coverage_auc": 0.935319,
      "score": 87.063704,
      "selective_acc_at_50": 0.928571,
      "selective_acc_at_80": 0.929936
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 940,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / process sensors",
    "n_classes": 2,
    "n_test": 392,
    "runtime_s": 1.8,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "chronological split minted by MedEval-1 on the production timestamp shipped in secom_labels.data: earliest 60% train, next 15% validation, latest 25% test. A random split is not defensible on a process-monitoring corpus -- tools drift and are recalibrated, so shuffling lets the model interpolate a tool's state from wafers measured minutes on either side, while production always extrapolates forward. Missing values imputed with the TRAIN median only; channels all-missing or constant on the training period dropped (468 of 590 channels kept).",
    "task": "secom",
    "task_name": "Semiconductor fabrication pass/fail",
    "timestamp": "2026-08-01T19:13:53+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 70.252251
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.783951,
      "auc": 0.961286,
      "balanced_accuracy": 0.794405,
      "ci95": [
       0.745603,
       0.834072
      ],
      "n_test": 486,
      "score": 76.01388
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.783951,
      "brier": 0.310343,
      "ece": 0.090324,
      "mean_confidence": 0.694573,
      "overconfidence": -0.089378,
      "score": 54.837963
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.760139,
      "mean_agreement": 0.786237,
      "mean_kappa": 0.697654,
      "mean_retention": 0.724422,
      "n_cases": 486,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.93,
         0.8786,
         0.786
        ],
        "mean_kappa": 0.818673,
        "mean_retention": 0.859632,
        "retention_by_severity": [
         0.9394,
         0.8884,
         0.7511
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.7695,
         0.7078,
         0.6029
        ],
        "mean_kappa": 0.564809,
        "mean_retention": 0.599009,
        "retention_by_severity": [
         0.7185,
         0.6474,
         0.4311
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9835,
         0.8272,
         0.5905
        ],
        "mean_kappa": 0.709479,
        "mean_retention": 0.714626,
        "retention_by_severity": [
         0.9864,
         0.7915,
         0.366
        ],
        "scored": true
       }
      },
      "score": 72.442213,
      "worst_retention": 0.365976
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.760139,
      "mean_retention": 0.896793,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.766537,
        "n_labels": 571,
        "retention": 1
       },
       "n20": {
        "chance_corrected": 0.693,
        "n_labels": 140,
        "retention": 0.911676
       },
       "n5": {
        "chance_corrected": 0.591922,
        "n_labels": 35,
        "retention": 0.778702
       }
      },
      "score": 89.679257
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {
       "steel_type": {
        "gap": 0.190301,
        "per_group": {
         "A300": 0.555961,
         "A400": 0.746262
        },
        "worst": 0.555961
       }
      },
      "basis": "worst demographic subgroup (real metadata)",
      "has_real_metadata": true,
      "score": 48.195501,
      "worst_class_recall": 0.538462
     },
     "uncertainty": {
      "full_accuracy": 0.783951,
      "risk_coverage_auc": 0.930371,
      "score": 91.876635,
      "selective_acc_at_50": 0.950617,
      "selective_acc_at_80": 0.861183
     }
    },
    "domain": "industrial",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": true,
     "n_train_available": 1164,
     "notes": null,
     "subgroup_basis": "worst demographic subgroup (real metadata)",
     "train_capped": false
    },
    "modality": "Tabular / imaging-derived",
    "n_classes": 7,
    "n_test": 486,
    "runtime_s": 1.6,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. The label is the argmax of the seven trailing one-hot fault columns, which are dropped from the features.",
    "task": "steel_plate_faults",
    "task_name": "Steel plate surface faults (7-class)",
    "timestamp": "2026-08-01T19:13:43+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 67.604273
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.8025,
      "auc": 0.869485,
      "balanced_accuracy": 0.804869,
      "ci95": [
       0.769938,
       0.841755
      ],
      "n_test": 400,
      "score": 60.973771
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.8025,
      "brier": 0.291134,
      "ece": 0.053544,
      "mean_confidence": 0.770944,
      "overconfidence": -0.031556,
      "score": 73.228125
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.609738,
      "mean_agreement": 0.885556,
      "mean_kappa": 0.771678,
      "mean_retention": 0.802664,
      "n_cases": 400,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9875,
         0.96,
         0.91
        ],
        "mean_kappa": 0.904894,
        "mean_retention": 0.92718,
        "retention_by_severity": [
         0.9759,
         0.9506,
         0.8551
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.915,
         0.8075,
         0.745
        ],
        "mean_kappa": 0.647395,
        "mean_retention": 0.69864,
        "retention_by_severity": [
         0.8543,
         0.681,
         0.5606
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.965,
         0.8975,
         0.7825
        ],
        "mean_kappa": 0.762744,
        "mean_retention": 0.782173,
        "retention_by_severity": [
         0.9613,
         0.7876,
         0.5976
        ],
        "scored": true
       }
      },
      "score": 80.266447,
      "worst_retention": 0.56061
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.609738,
      "mean_retention": 0.646395,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.481208,
        "n_labels": 200,
        "retention": 0.789205
       },
       "n20": {
        "chance_corrected": 0.322028,
        "n_labels": 40,
        "retention": 0.528142
       },
       "n5": {
        "chance_corrected": 0.379158,
        "n_labels": 10,
        "retention": 0.621838
       }
      },
      "score": 64.639473
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 54.205607,
      "worst_class_recall": 0.771028
     },
     "uncertainty": {
      "full_accuracy": 0.8025,
      "risk_coverage_auc": 0.891217,
      "score": 78.243386,
      "selective_acc_at_50": 0.895,
      "selective_acc_at_80": 0.853125
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 959,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 400,
    "runtime_s": 1.5,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (red wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_red",
    "task_name": "Wine quality, red (binarised)",
    "timestamp": "2026-08-01T19:13:41+00:00"
   },
   {
    "code_fingerprint": "b444c02112cc",
    "composite": {
     "coverage": 0.72,
     "dimensions_scored": [
      "accuracy",
      "calibration",
      "corruption",
      "limited_data",
      "subgroup",
      "uncertainty"
     ],
     "index": 63.660968
    },
    "device": "mps",
    "dimensions": {
     "accuracy": {
      "accuracy": 0.850612,
      "auc": 0.913064,
      "balanced_accuracy": 0.813796,
      "ci95": [
       0.787287,
       0.838595
      ],
      "n_test": 1225,
      "score": 62.75924
     },
     "acquisition": null,
     "calibration": {
      "accuracy": 0.850612,
      "brier": 0.223919,
      "ece": 0.051031,
      "mean_confidence": 0.799582,
      "overconfidence": -0.051031,
      "score": 74.484694
     },
     "corruption": {
      "basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
      "clean_reference_cc": 0.623862,
      "mean_agreement": 0.846019,
      "mean_kappa": 0.594344,
      "mean_retention": 0.638942,
      "n_cases": 1200,
      "per_perturbation": {
       "coding_drift": {
        "agreement_by_severity": [
         1,
         1,
         1
        ],
        "mean_kappa": 1,
        "mean_retention": 1,
        "not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
        "retention_by_severity": [
         1,
         1,
         1
        ],
        "scored": false
       },
       "entry_error": {
        "agreement_by_severity": [
         0.9875,
         0.9558,
         0.9017
        ],
        "mean_kappa": 0.875738,
        "mean_retention": 0.911383,
        "retention_by_severity": [
         0.986,
         0.9182,
         0.83
        ],
        "scored": true
       },
       "field_dropout": {
        "agreement_by_severity": [
         0.9258,
         0.7475,
         0.725
        ],
        "mean_kappa": 0.351449,
        "mean_retention": 0.35686,
        "retention_by_severity": [
         0.9006,
         0.1306,
         0.0394
        ],
        "scored": true
       },
       "unit_scale": {
        "agreement_by_severity": [
         0.9592,
         0.8133,
         0.5983
        ],
        "mean_kappa": 0.555844,
        "mean_retention": 0.648583,
        "retention_by_severity": [
         0.925,
         0.609,
         0.4118
        ],
        "scored": true
       }
      },
      "score": 63.894206,
      "worst_retention": 0.039437
     },
     "cost": null,
     "limited_data": {
      "full_reference_cc": 0.627592,
      "mean_retention": 0.613625,
      "per_budget": {
       "n100": {
        "chance_corrected": 0.509202,
        "n_labels": 200,
        "retention": 0.811359
       },
       "n20": {
        "chance_corrected": 0.442122,
        "n_labels": 40,
        "retention": 0.704473
       },
       "n5": {
        "chance_corrected": 0.203995,
        "n_labels": 10,
        "retention": 0.325044
       }
      },
      "score": 61.362516
     },
     "spec_sensitivity": null,
     "subgroup": {
      "attributes": {},
      "basis": "worst diagnostic class (no demographic metadata in source)",
      "has_real_metadata": false,
      "score": 40.487805,
      "worst_class_recall": 0.702439
     },
     "uncertainty": {
      "full_accuracy": 0.850612,
      "risk_coverage_auc": 0.944288,
      "score": 88.857663,
      "selective_acc_at_50": 0.96732,
      "selective_acc_at_80": 0.908163
     }
    },
    "domain": "agriculture",
    "kind": "tabular",
    "limitations": {
     "has_real_subgroup_metadata": false,
     "n_train_available": 2938,
     "notes": null,
     "subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
     "train_capped": false
    },
    "modality": "Tabular / physicochemical",
    "n_classes": 2,
    "n_test": 1225,
    "runtime_s": 1.7,
    "seed": 20260727,
    "spec_version": "MedEval-1 v0.1",
    "split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (white wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
    "task": "wine_quality_white",
    "task_name": "Wine quality, white (binarised)",
    "timestamp": "2026-08-01T19:14:08+00:00"
   }
  ],
  "heldout": null,
  "sources": [
   {
    "path": "results/cross_site/occupancy__tab_rf.json",
    "sha256": "a405d2c58de366fe1d1affa19d29be533b62395b8de360dd8e1b9d63f0ee1713"
   },
   {
    "path": "results/heldout/audit_report.json",
    "sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
   },
   {
    "path": "results/heldout/cycle.json",
    "sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
   },
   {
    "path": "results/heldout/manifests/gen1.seal.json",
    "sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
   },
   {
    "path": "results/heldout/manifests/gen2.seal.json",
    "sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
   },
   {
    "path": "results/heldout/manifests/public.seal.json",
    "sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
   },
   {
    "path": "results/records/adult_income__tab_rf.json",
    "sha256": "0c72895bc452f76c94279742836ab08f6a3b90a995b96b35d1fc1ba390a5e3da"
   },
   {
    "path": "results/records/anuran_mfcc__tab_rf.json",
    "sha256": "2aa16d343d4791780b6688e3cbdec9935e8bc1207243f0c4a3d814767a88dc3e"
   },
   {
    "path": "results/records/breast_wdbc__tab_rf.json",
    "sha256": "b7fc1d6a57687ecba17348b2d8fbbea4f67ee453bb423b3148d82f67be2ffb52"
   },
   {
    "path": "results/records/coil2000_caravan__tab_rf.json",
    "sha256": "64ae58d98289e58681d35676afc9effd95649504396200c830fae76b85b36f91"
   },
   {
    "path": "results/records/diabetes_pima__tab_rf.json",
    "sha256": "aa7459b59b46fc8a1d1624f4d7c8ea2b1e7c4acd75206c45c5e3ab218477ef11"
   },
   {
    "path": "results/records/dry_bean__tab_rf.json",
    "sha256": "6b7012e82fb4d774eb7443cf3e5b6b335484b897b5be188ec49212af3ed727a4"
   },
   {
    "path": "results/records/german_credit__tab_rf.json",
    "sha256": "24ff6f5c6838f28b152a67cef79296d839893be12124bb44b60f2624dcb22c65"
   },
   {
    "path": "results/records/heart_cleveland__tab_rf.json",
    "sha256": "4e3b6e895abda7ad0c9377c4fdf3c85e91e061d7fb2705de140ae2b116a2fb4f"
   },
   {
    "path": "results/records/landsat_statlog__tab_rf.json",
    "sha256": "62fa9186b3120e32280b31f0b83e1a03855fe9020b74ac04bd2bdf8d6d2d67f9"
   },
   {
    "path": "results/records/occupancy_detection_2__tab_rf.json",
    "sha256": "7e81c304359ec6370e7e9464502c0f66c2368e2b5b036e558c8c5ffd35d7316e"
   },
   {
    "path": "results/records/occupancy_detection__tab_rf.json",
    "sha256": "fe8a6aea38b07daf1ae9f259f3a6e3eb1783d625afd1bfe80b95c96ea988872f"
   },
   {
    "path": "results/records/secom__tab_rf.json",
    "sha256": "6450800d5380b41fabcb5512ad4df69fee970347b54814ad5793ca20f71b3947"
   },
   {
    "path": "results/records/steel_plate_faults__tab_rf.json",
    "sha256": "34a1824eb2b617a095e35ef2f305eea48af27d9386eb7cb39713abaf85847083"
   },
   {
    "path": "results/records/wine_quality_red__tab_rf.json",
    "sha256": "f02c6fee1004ba93bcc07253fb83e41177515660541d7f7b5a2e3687131473c2"
   },
   {
    "path": "results/records/wine_quality_white__tab_rf.json",
    "sha256": "16bcf923dd4d658453df5e051aa97fb67046704ce9379fb9e22c380c88045ba6"
   }
  ]
 },
 "declared": null,
 "declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
 "expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
 "expiry_utc": "2026-09-20T00:00:00Z",
 "issued_utc": "2026-08-03T16:03:42Z",
 "issuer": {
  "algo": "Ed25519",
  "key_id": "ns-passport-2026-07",
  "name": "NakedSignal",
  "public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
 },
 "observed": null,
 "observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
 "passport_version": "0.1",
 "signature": "Glju72u6iRXGMk0+GctJ7k09PqTYZ3Sr19NL5eS8aMR8MUcIS4Rze/LU+YpxQb/3ZFmm50L8lFPqPpT1ukSoCg==",
 "spec_version": "MedEval-1 v0.1",
 "subject": {
  "behaviour_version": "1.0.0",
  "code_fingerprint": null,
  "code_fingerprints": [
   "b444c02112cc",
   "de32bca72d56"
  ],
  "family": "classical",
  "model_id": "tab_rf",
  "model_name": "Random forest",
  "params": "400 trees"
 }
}

Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.