Logistic regression
LIVEThe signed evidence record for tab_logreg. Everything below was read out of the passport file; the signature was checked when this page was built, by the same code the Verify button runs.
tab_logregThe check runs against the public key published on the governance page, using your browser's own Ed25519 implementation. The exit code shown is the exit code verify_passport.py returns for the same document.
Computed in this repo by the MedEval-1 harness on data the model had never seen. Every number below was copied out of a file the harness wrote — 21 of them, each listed with its SHA-256 at the foot of this panel — and is printed exactly as it was signed, not re-rounded for display.
| Task | Domain | n test | accuracy | calibration | uncertainty | subgroup | corruption | acquisition | limited data | cost | spec sensitivity | Index |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| adult_income | labour | 16281 | 53.286619 | 97.092269 | 91.086494 | 29.43609 | 90.644038 | n/a | 79.10956 | n/a | n/a | 70.641549 |
| anuran_mfcc | ecology | 995 | 72.156942 | 0 | 58.575159 | 76.089222 | 64.782891 | n/a | 87.030642 | n/a | n/a | 58.950366 |
| breast_wdbc | medicine | 143 | 98.113208 | 86.140641 | 99.450692 | 96.226415 | 79.717474 | n/a | 92.827635 | n/a | n/a | 91.618086 |
| coil2000_caravan | financial | 4000 | 1.574346 | 94.048802 | 94.400346 | 0 | 81.516422 | n/a | 66.666667 | n/a | n/a | 47.275486 |
| diabetes_pima | medicine | 193 | 35.572139 | 67.027738 | 68.833292 | 32.727273 | 90.583491 | n/a | 96.614497 | n/a | n/a | 60.64544 |
| dry_bean | agriculture | 3404 | 93.098832 | 90.474555 | 98.807007 | 87.961558 | 75.224967 | n/a | 96.569948 | n/a | n/a | 89.218101 |
| german_credit | financial | 250 | 32.380952 | 62.532703 | 75.100291 | 27.629234 | 96.732026 | n/a | 62.941176 | n/a | n/a | 56.039968 |
| heart_cleveland | medicine | 75 | 70.714286 | 61.290501 | 83.936481 | 60 | 84.624018 | n/a | 83.838384 | n/a | n/a | 72.605785 |
| landsat_statlog | earth | 2000 | 74.631754 | 86.991913 | 95.539596 | 14.691943 | 67.25434 | n/a | 97.095518 | n/a | n/a | 71.051886 |
| occupancy_detection_2 | energy | 9752 | 79.953726 | 83.575984 | 98.412349 | 74.914592 | 88.425888 | n/a | 96.840405 | n/a | n/a | 84.713377 |
| occupancy_detection | energy | 2665 | 95.472947 | 89.92598 | 99.438032 | 93.620791 | 88.191024 | n/a | 92.965901 | n/a | n/a | 92.79503 |
| secom | industrial | 392 | 0 | 61.814278 | 85.230543 | 0 | not scored | n/a | 0 | n/a | n/a | 21.202385 |
| steel_plate_faults | industrial | 486 | 70.434925 | 66.004218 | 82.429163 | 34.175968 | 66.145126 | n/a | 92.579436 | n/a | n/a | 67.058277 |
| wine_quality_red | agriculture | 400 | 50.386896 | 69.318061 | 68.821262 | 44.859813 | 70.833887 | n/a | 63.096663 | n/a | n/a | 59.705537 |
| wine_quality_white | agriculture | 1225 | 38.226844 | 85.631979 | 76.669329 | 0 | 63.772944 | n/a | 66.666667 | n/a | n/a | 52.273715 |
Cross-site retention · occupancy
retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column
Private held-out track
this subject was not submitted to a sealed generation.Known limitations
Audit chain
python src/heldout/audit.py 1f66c7719ab3943c6fcc21 source files, with hashes
| results/cross_site/occupancy__tab_logreg.json | 39f6dd624d9723016056f3deb5caf4f5be5c852ab6ec756c45afd28732dd3d52 |
| results/heldout/audit_report.json | df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2 |
| results/heldout/cycle.json | e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502 |
| results/heldout/manifests/gen1.seal.json | 91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81 |
| results/heldout/manifests/gen2.seal.json | f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547 |
| results/heldout/manifests/public.seal.json | 1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c |
| results/records/adult_income__tab_logreg.json | 9d49f868e37c94bdcbde73d8ea4f043ccd8d8fc63c89550f8e2d95605e67deb0 |
| results/records/anuran_mfcc__tab_logreg.json | c1271e3ccaa18a559772a12ca02fb9265413e70bb06f299c2c87b78c9f5ac6f0 |
| results/records/breast_wdbc__tab_logreg.json | 9756b728ad8a36e7bd2806f46f3fb6b9c2bd8590a97bcf5c22fc66f28fb165be |
| results/records/coil2000_caravan__tab_logreg.json | 252f17a68d072d2cb05ee5514df2f21bdf1d6caa45904c07260ed5c59778507b |
| results/records/diabetes_pima__tab_logreg.json | d8c33c2b3ecb291523db6e86a5fb15de72eab3df9c9f8050abf5a97a862441d3 |
| results/records/dry_bean__tab_logreg.json | 746ef76d9275ef54925dba9afd291e7dcc4986601dfbf86802325ca2bffa87ee |
| results/records/german_credit__tab_logreg.json | ef0656154fcaed4e3ec40dfe3f52387aed341910b42f970cb4feb3ea9a0dcdd9 |
| results/records/heart_cleveland__tab_logreg.json | ee6774fabcb5b6f03727dfba90c2b3081f6f98b06172ecfc2f9e4241b3432ee2 |
| results/records/landsat_statlog__tab_logreg.json | 2fe08d5aceaf2cfc8e5eb9cbdf58169f99876093ad34aea9e16de93890ee949f |
| results/records/occupancy_detection_2__tab_logreg.json | 33859c9106ad2179d4b479a3b53c112f6e1e36dfd7324b915658e8ab4f751bce |
| results/records/occupancy_detection__tab_logreg.json | 7a7adae2920572b0569780fdcc65883daabc66360174c0989ebfb3727800090d |
| results/records/secom__tab_logreg.json | 5e2c3e5e4a5dea390dd9968e2121956ee3e704ff150bd3a93301bd52d299e797 |
| results/records/steel_plate_faults__tab_logreg.json | bcb35954c6c21b3a1a223d12a924f3ff3ca1cebc8016823872b1c305aa3b4e6b |
| results/records/wine_quality_red__tab_logreg.json | 528c7eedc18d0baf7620c8f7ba6e36991a41f7675d52c150ca92ce4ae23452fe |
| results/records/wine_quality_white__tab_logreg.json | a1caccca58fe53897a0124cf62fc019fe3b4bc6b3a15aee3d5c22da9d3460e68 |
Five fields a vendor asserts about its own product — intended use, forbidden use, training cutoff, regulatory clearances, and who is personally attesting. NakedSignal never fills these in on a vendor's behalf.
No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.
Production history and drift — how the model has actually behaved since it was deployed, on real traffic.
Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.
Raw JSON — the whole signed document, 76,592 characters
{
"canonicalisation": {
"float_rounding": "every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact",
"form": "RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting",
"signed_over": "the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8"
},
"computed": {
"audit": {
"audit_command": "python src/heldout/audit.py 1f66c7719ab3943c6fcc",
"bundles": 48,
"ledger_entries": 100,
"ledger_head": "6b7b371ff650a8edaf477490f2a5f9b043433f3ba9941989ee1962ed8f1b1d88",
"ledger_intact": true,
"rederived": 48
},
"cross_site": {
"code_fingerprint": "b444c02112cc",
"device": "mps",
"family": "occupancy",
"matrix": {
"occupancy_detection": {
"occupancy_detection": {
"balanced_accuracy": 0.977365,
"chance_corrected": 0.954729,
"ece": 0.020148
},
"occupancy_detection_2": {
"balanced_accuracy": 0.899769,
"chance_corrected": 0.799537,
"ece": 0.032848,
"retention": 0.837449
}
},
"occupancy_detection_2": {
"occupancy_detection": {
"balanced_accuracy": 0.977365,
"chance_corrected": 0.954729,
"ece": 0.020148,
"retention": 1
},
"occupancy_detection_2": {
"balanced_accuracy": 0.899769,
"chance_corrected": 0.799537,
"ece": 0.032848
}
}
},
"mean_retention": 0.918725,
"note": "retention is absolute retained skill (chance-corrected), per spec 2.2.1; the diagonal is same-site, every off-diagonal cell is train-on-row, test-on-column",
"score": 91.872451,
"site_labels": {
"occupancy_detection": "test period 1",
"occupancy_detection_2": "test period 2"
},
"sites": [
"occupancy_detection",
"occupancy_detection_2"
],
"spec_version": "MedEval-1 v0.1",
"worst_retention": 0.837449
},
"evaluations": [
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 70.641549
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.851299,
"auc": 0.900916,
"balanced_accuracy": 0.766433,
"ci95": [
0.758858,
0.775213
],
"n_test": 16281,
"score": 53.286619
},
"acquisition": null,
"calibration": {
"accuracy": 0.851299,
"brier": 0.206377,
"ece": 0.005815,
"mean_confidence": 0.856261,
"overconfidence": 0.004961,
"score": 97.092269
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.547023,
"mean_agreement": 0.950926,
"mean_kappa": 0.834848,
"mean_retention": 0.90644,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9975,
0.9833,
0.97
],
"mean_kappa": 0.95156,
"mean_retention": 0.999333,
"retention_by_severity": [
0.998,
1,
1
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9692,
0.9492,
0.8283
],
"mean_kappa": 0.66839,
"mean_retention": 0.725449,
"retention_by_severity": [
0.9936,
0.9041,
0.2786
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9992,
0.9942,
0.8675
],
"mean_kappa": 0.884593,
"mean_retention": 0.99454,
"retention_by_severity": [
0.998,
0.9856,
1
],
"scored": true
}
},
"score": 90.644038,
"worst_retention": 0.278623
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.532866,
"mean_retention": 0.791096,
"per_budget": {
"n100": {
"chance_corrected": 0.516932,
"n_labels": 200,
"retention": 0.970098
},
"n20": {
"chance_corrected": 0.437943,
"n_labels": 40,
"retention": 0.821862
},
"n5": {
"chance_corrected": 0.309769,
"n_labels": 10,
"retention": 0.581327
}
},
"score": 79.10956
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"age_band": {
"gap": 0.00671,
"per_group": {
"45-59": 0.758929,
"60+": 0.765639,
"<45": 0.759156
},
"worst": 0.758929
},
"race": {
"gap": 0.120628,
"per_group": {
"Amer-Indian-Eskimo": 0.64718,
"Asian-Pac-Islander": 0.766549,
"Black": 0.721149,
"Other": 0.706364,
"White": 0.767809
},
"worst": 0.64718
},
"sex": {
"gap": 0.000022,
"per_group": {
"Female": 0.757175,
"Male": 0.757197
},
"worst": 0.757175
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 29.43609,
"worst_class_recall": 0.605564
},
"uncertainty": {
"full_accuracy": 0.851299,
"risk_coverage_auc": 0.955432,
"score": 91.086494,
"selective_acc_at_50": 0.978378,
"selective_acc_at_80": 0.913321
}
},
"domain": "labour",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 27676,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": true
},
"modality": "Tabular / socioeconomic",
"n_classes": 2,
"n_test": 16281,
"runtime_s": 0.2,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split (UCI adult.test scored whole); validation carved from adult.data only (stratified 15%, seed 20260727)",
"task": "adult_income",
"task_name": "Census income (>$50K)",
"timestamp": "2026-08-01T19:16:09+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 58.950366
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.620101,
"auc": 0.92757,
"balanced_accuracy": 0.749412,
"ci95": [
0.707507,
0.788672
],
"n_test": 995,
"score": 72.156942
},
"acquisition": null,
"calibration": {
"accuracy": 0.620101,
"brier": 0.65531,
"ece": 0.260751,
"mean_confidence": 0.878081,
"overconfidence": 0.257981,
"score": 0
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.721569,
"mean_agreement": 0.736683,
"mean_kappa": 0.697545,
"mean_retention": 0.647829,
"n_cases": 995,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9497,
0.8693,
0.7538
],
"mean_kappa": 0.834531,
"mean_retention": 0.875213,
"retention_by_severity": [
0.9282,
0.8991,
0.7983
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.7598,
0.4935,
0.3487
],
"mean_kappa": 0.468494,
"mean_retention": 0.311539,
"retention_by_severity": [
0.6965,
0.1654,
0.0727
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9779,
0.8774,
0.6
],
"mean_kappa": 0.789612,
"mean_retention": 0.756735,
"retention_by_severity": [
0.9967,
0.8264,
0.4471
],
"scored": true
}
},
"score": 64.782891,
"worst_retention": 0.07272
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.721569,
"mean_retention": 0.870306,
"per_budget": {
"n100": {
"chance_corrected": 0.776993,
"n_labels": 794,
"retention": 1
},
"n20": {
"chance_corrected": 0.727142,
"n_labels": 200,
"retention": 1
},
"n5": {
"chance_corrected": 0.440821,
"n_labels": 50,
"retention": 0.610919
}
},
"score": 87.030642
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"family": {
"gap": 0.065627,
"per_group": {
"Hylidae": 0.85043,
"Leptodactylidae": 0.784803
},
"worst": 0.784803
},
"genus": {
"gap": null,
"per_group": {
"Hypsiboas": 0.925172
},
"worst": null
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 76.089222,
"worst_class_recall": 0.223975
},
"uncertainty": {
"full_accuracy": 0.620101,
"risk_coverage_auc": 0.627176,
"score": 58.575159,
"selective_acc_at_50": 0.648594,
"selective_acc_at_80": 0.644472
}
},
"domain": "ecology",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 5073,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": false
},
"modality": "Bioacoustic / MFCC",
"n_classes": 10,
"n_test": 995,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "no official split; minted at the standard seed (20260727) by partitioning the 60 RecordID recording groups 60/15/15/25 and taking whole recordings, so no recording spans two splits. A row-level random split would place near-duplicate syllables from one recording on both sides and score memorisation (~99%); class balance is therefore uncontrolled across splits and is not stratified",
"task": "anuran_mfcc",
"task_name": "Anuran call species (10-class)",
"timestamp": "2026-08-01T19:14:27+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 91.618086
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.993007,
"auc": 0.992662,
"balanced_accuracy": 0.990566,
"ci95": [
0.968077,
1
],
"n_test": 143,
"score": 98.113208
},
"acquisition": null,
"calibration": {
"accuracy": 0.993007,
"brier": 0.026788,
"ece": 0.027719,
"mean_confidence": 0.971885,
"overconfidence": -0.021122,
"score": 86.140641
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.981132,
"mean_agreement": 0.874903,
"mean_kappa": 0.78041,
"mean_retention": 0.797175,
"n_cases": 143,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.986,
0.965,
0.8951
],
"mean_kappa": 0.894183,
"mean_retention": 0.927137,
"retention_by_severity": [
0.9774,
0.9434,
0.8607
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.993,
0.993,
0.9371
],
"mean_kappa": 0.943292,
"mean_retention": 0.937393,
"retention_by_severity": [
0.9887,
0.9887,
0.8348
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.993,
0.7483,
0.3636
],
"mean_kappa": 0.503756,
"mean_retention": 0.526994,
"retention_by_severity": [
0.9887,
0.5923,
0
],
"scored": true
}
},
"score": 79.717474,
"worst_retention": 0
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.981132,
"mean_retention": 0.928276,
"per_budget": {
"n100": {
"chance_corrected": 0.970021,
"n_labels": 200,
"retention": 0.988675
},
"n20": {
"chance_corrected": 0.814465,
"n_labels": 40,
"retention": 0.830128
},
"n5": {
"chance_corrected": 0.947799,
"n_labels": 10,
"retention": 0.966026
}
},
"score": 92.827635
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 96.226415,
"worst_class_recall": 0.981132
},
"uncertainty": {
"full_accuracy": 0.993007,
"risk_coverage_auc": 0.997253,
"score": 99.450692,
"selective_acc_at_50": 1,
"selective_acc_at_80": 0.991228
}
},
"domain": "medicine",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 341,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / imaging-derived",
"n_classes": 2,
"n_test": 143,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
"task": "breast_wdbc",
"task_name": "Breast cytology (WDBC)",
"timestamp": "2026-08-01T19:13:37+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 47.275486
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.9405,
"auc": 0.726572,
"balanced_accuracy": 0.507872,
"ci95": [
0.501587,
0.516474
],
"n_test": 4000,
"score": 1.574346
},
"acquisition": null,
"calibration": {
"accuracy": 0.9405,
"brier": 0.107726,
"ece": 0.011902,
"mean_confidence": 0.943005,
"overconfidence": 0.002505,
"score": 94.048802
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.040202,
"mean_agreement": 0.951296,
"mean_kappa": 0.483838,
"mean_retention": 0.815164,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9983,
0.9958,
0.9883
],
"mean_kappa": 0.644412,
"mean_retention": 0.985325,
"retention_by_severity": [
0.956,
1,
1
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9967,
0.9917,
0.9858
],
"mean_kappa": 0.234038,
"mean_retention": 0.460168,
"retention_by_severity": [
0.956,
0.2893,
0.1352
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
1,
0.9958,
0.6092
],
"mean_kappa": 0.573063,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
}
},
"score": 81.516422,
"worst_retention": 0.13522
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.015743,
"mean_retention": 0.666667,
"per_budget": {
"n100": {
"chance_corrected": 0.274787,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0,
"n_labels": 40,
"retention": 0
},
"n5": {
"chance_corrected": 0.10187,
"n_labels": 10,
"retention": 1
}
},
"score": 66.666667
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.016807
},
"uncertainty": {
"full_accuracy": 0.9405,
"risk_coverage_auc": 0.972002,
"score": 94.400346,
"selective_acc_at_50": 0.9735,
"selective_acc_at_80": 0.961875
}
},
"domain": "financial",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 4948,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / insurance",
"n_classes": 2,
"n_test": 4000,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split (CoIL Challenge 2000: ticdata2000.txt 5822 train / ticeval2000.txt 4000 test). The test labels were released separately, in tictgts2000.txt, only after the challenge closed -- they were not available to anyone tuning on this benchmark at the time. Test scored whole; validation carved from ticdata2000.txt only (stratified 15%, seed 20260727).",
"task": "coil2000_caravan",
"task_name": "Caravan insurance purchase (CoIL 2000)",
"timestamp": "2026-08-01T19:15:22+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 60.64544
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.725389,
"auc": 0.783345,
"balanced_accuracy": 0.677861,
"ci95": [
0.613723,
0.751398
],
"n_test": 193,
"score": 35.572139
},
"acquisition": null,
"calibration": {
"accuracy": 0.725389,
"brier": 0.361033,
"ece": 0.065945,
"mean_confidence": 0.790776,
"overconfidence": 0.065387,
"score": 67.027738
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.355721,
"mean_agreement": 0.893495,
"mean_kappa": 0.805428,
"mean_retention": 0.905835,
"n_cases": 193,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9896,
0.9741,
0.9171
],
"mean_kappa": 0.906041,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9741,
0.9378,
0.9482
],
"mean_kappa": 0.882132,
"mean_retention": 0.931846,
"retention_by_severity": [
0.9134,
0.8821,
1
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9845,
0.9275,
0.3886
],
"mean_kappa": 0.628111,
"mean_retention": 0.785659,
"retention_by_severity": [
0.9973,
1,
0.3596
],
"scored": true
}
},
"score": 90.583491,
"worst_retention": 0.35964
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.355721,
"mean_retention": 0.966145,
"per_budget": {
"n100": {
"chance_corrected": 0.401801,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.429756,
"n_labels": 40,
"retention": 1
},
"n5": {
"chance_corrected": 0.319593,
"n_labels": 10,
"retention": 0.898435
}
},
"score": 96.614497
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"age_band": {
"gap": 0.009091,
"per_group": {
"45-59": 0.663636,
"<45": 0.672727
},
"worst": 0.663636
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 32.727273,
"worst_class_recall": 0.522388
},
"uncertainty": {
"full_accuracy": 0.725389,
"risk_coverage_auc": 0.844166,
"score": 68.833292,
"selective_acc_at_50": 0.854167,
"selective_acc_at_80": 0.785714
}
},
"domain": "medicine",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 460,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": false
},
"modality": "Tabular / clinical",
"n_classes": 2,
"n_test": 193,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
"task": "diabetes_pima",
"task_name": "Diabetes onset (Pima)",
"timestamp": "2026-08-01T19:13:35+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 89.218101
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.93067,
"auc": 0.995574,
"balanced_accuracy": 0.940847,
"ci95": [
0.933445,
0.949987
],
"n_test": 3404,
"score": 93.098832
},
"acquisition": null,
"calibration": {
"accuracy": 0.93067,
"brier": 0.105137,
"ece": 0.019051,
"mean_confidence": 0.915137,
"overconfidence": -0.015533,
"score": 90.474555
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.93135,
"mean_agreement": 0.778889,
"mean_kappa": 0.7413,
"mean_retention": 0.75225,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9775,
0.9025,
0.7675
],
"mean_kappa": 0.858214,
"mean_retention": 0.898208,
"retention_by_severity": [
0.9843,
0.9204,
0.79
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9575,
0.8825,
0.5983
],
"mean_kappa": 0.768035,
"mean_retention": 0.720179,
"retention_by_severity": [
0.9221,
0.8503,
0.3881
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9817,
0.8467,
0.0958
],
"mean_kappa": 0.597651,
"mean_retention": 0.638363,
"retention_by_severity": [
1,
0.9151,
0
],
"scored": true
}
},
"score": 75.224967,
"worst_retention": 0
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.930988,
"mean_retention": 0.965699,
"per_budget": {
"n100": {
"chance_corrected": 0.925595,
"n_labels": 700,
"retention": 0.994207
},
"n20": {
"chance_corrected": 0.905913,
"n_labels": 140,
"retention": 0.973066
},
"n5": {
"chance_corrected": 0.865657,
"n_labels": 35,
"retention": 0.929825
}
},
"score": 96.569948
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 87.961558,
"worst_class_recall": 0.896813
},
"uncertainty": {
"full_accuracy": 0.93067,
"risk_coverage_auc": 0.989774,
"score": 98.807007,
"selective_acc_at_50": 0.99765,
"selective_acc_at_80": 0.983841
}
},
"domain": "agriculture",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 8166,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / imaging-derived",
"n_classes": 7,
"n_test": 3404,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. Parsed from the ARFF in the donor's zip, whose nominal class order is the class order used here.",
"task": "dry_bean",
"task_name": "Dry bean variety (7-class)",
"timestamp": "2026-08-01T19:15:26+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 56.039968
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.756,
"auc": 0.791771,
"balanced_accuracy": 0.661905,
"ci95": [
0.611306,
0.722672
],
"n_test": 250,
"score": 32.380952
},
"acquisition": null,
"calibration": {
"accuracy": 0.756,
"brier": 0.332951,
"ece": 0.074935,
"mean_confidence": 0.79844,
"overconfidence": 0.04244,
"score": 62.532703
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.32381,
"mean_agreement": 0.953333,
"mean_kappa": 0.858539,
"mean_retention": 0.96732,
"n_cases": 250,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.996,
0.984,
0.964
],
"mean_kappa": 0.940772,
"mean_retention": 0.958824,
"retention_by_severity": [
1,
1,
0.8765
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.936,
0.92,
0.88
],
"mean_kappa": 0.736584,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.988,
0.98,
0.932
],
"mean_kappa": 0.898261,
"mean_retention": 0.943137,
"retention_by_severity": [
0.9353,
0.9824,
0.9118
],
"scored": true
}
},
"score": 96.732026,
"worst_retention": 0.876471
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.32381,
"mean_retention": 0.629412,
"per_budget": {
"n100": {
"chance_corrected": 0.466667,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.213333,
"n_labels": 40,
"retention": 0.658824
},
"n5": {
"chance_corrected": 0.074286,
"n_labels": 10,
"retention": 0.229412
}
},
"score": 62.941176
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"age_band": {
"gap": 0.155835,
"per_group": {
"45-59": 0.793981,
"<45": 0.638146
},
"worst": 0.638146
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 27.629234,
"worst_class_recall": 0.426667
},
"uncertainty": {
"full_accuracy": 0.756,
"risk_coverage_auc": 0.875501,
"score": 75.100291,
"selective_acc_at_50": 0.896,
"selective_acc_at_80": 0.805
}
},
"domain": "financial",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 600,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": false
},
"modality": "Tabular / financial",
"n_classes": 2,
"n_test": 250,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
"task": "german_credit",
"task_name": "Consumer credit risk (Statlog)",
"timestamp": "2026-08-01T19:13:39+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 72.605785
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.853333,
"auc": 0.903571,
"balanced_accuracy": 0.853571,
"ci95": [
0.773817,
0.928504
],
"n_test": 75,
"score": 70.714286
},
"acquisition": null,
"calibration": {
"accuracy": 0.853333,
"brier": 0.246061,
"ece": 0.077419,
"mean_confidence": 0.829689,
"overconfidence": -0.023644,
"score": 61.290501
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.707143,
"mean_agreement": 0.89037,
"mean_kappa": 0.777977,
"mean_retention": 0.84624,
"n_cases": 75,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.96,
0.96,
0.8667
],
"mean_kappa": 0.857986,
"mean_retention": 0.90404,
"retention_by_severity": [
0.9646,
0.9646,
0.7828
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.96,
0.9333,
0.6933
],
"mean_kappa": 0.71917,
"mean_retention": 0.789562,
"retention_by_severity": [
0.9545,
0.8838,
0.5303
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9467,
0.7733,
0.92
],
"mean_kappa": 0.756777,
"mean_retention": 0.845118,
"retention_by_severity": [
0.9141,
0.6919,
0.9293
],
"scored": true
}
},
"score": 84.624018,
"worst_retention": 0.530303
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.707143,
"mean_retention": 0.838384,
"per_budget": {
"n100": {
"chance_corrected": 0.707143,
"n_labels": 178,
"retention": 1
},
"n20": {
"chance_corrected": 0.578571,
"n_labels": 40,
"retention": 0.818182
},
"n5": {
"chance_corrected": 0.492857,
"n_labels": 10,
"retention": 0.69697
}
},
"score": 83.838384
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"age_band": {
"gap": null,
"per_group": {
"45-59": 0.85
},
"worst": null
},
"sex": {
"gap": 0.076552,
"per_group": {
"female": 0.8,
"male": 0.876552
},
"worst": 0.8
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 60,
"worst_class_recall": 0.85
},
"uncertainty": {
"full_accuracy": 0.853333,
"risk_coverage_auc": 0.919682,
"score": 83.936481,
"selective_acc_at_50": 0.921053,
"selective_acc_at_80": 0.883333
}
},
"domain": "medicine",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 178,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": false
},
"modality": "Tabular / clinical",
"n_classes": 2,
"n_test": 75,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727)",
"task": "heart_cleveland",
"task_name": "Coronary artery disease (Cleveland)",
"timestamp": "2026-08-01T19:13:33+00:00"
},
{
"code_fingerprint": "de32bca72d56",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 71.051886
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.8355,
"auc": 0.975552,
"balanced_accuracy": 0.788598,
"ci95": [
0.774709,
0.805025
],
"n_test": 2000,
"score": 74.631754
},
"acquisition": null,
"calibration": {
"accuracy": 0.8355,
"brier": 0.212281,
"ece": 0.026016,
"mean_confidence": 0.858246,
"overconfidence": 0.022746,
"score": 86.991913
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.742217,
"mean_agreement": 0.695,
"mean_kappa": 0.635778,
"mean_retention": 0.672543,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.975,
0.9033,
0.805
],
"mean_kappa": 0.868651,
"mean_retention": 0.93369,
"retention_by_severity": [
0.9803,
0.9627,
0.8581
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9108,
0.675,
0.2592
],
"mean_kappa": 0.536669,
"mean_retention": 0.547975,
"retention_by_severity": [
0.9316,
0.6109,
0.1015
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9683,
0.63,
0.1283
],
"mean_kappa": 0.502014,
"mean_retention": 0.535965,
"retention_by_severity": [
0.9889,
0.5961,
0.0229
],
"scored": true
}
},
"score": 67.25434,
"worst_retention": 0.022933
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.746318,
"mean_retention": 0.970955,
"per_budget": {
"n100": {
"chance_corrected": 0.775168,
"n_labels": 600,
"retention": 1
},
"n20": {
"chance_corrected": 0.732551,
"n_labels": 120,
"retention": 0.981554
},
"n5": {
"chance_corrected": 0.695054,
"n_labels": 30,
"retention": 0.931311
}
},
"score": 97.095518
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 14.691943,
"worst_class_recall": 0.2891
},
"uncertainty": {
"full_accuracy": 0.8355,
"risk_coverage_auc": 0.96283,
"score": 95.539596,
"selective_acc_at_50": 0.987,
"selective_acc_at_80": 0.92375
}
},
"domain": "earth",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 3769,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / multispectral",
"n_classes": 6,
"n_test": 2000,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split (Statlog sat.trn 4435 / sat.tst 2000), scored whole, with validation carved from sat.trn only (stratified 15%, seed 20260727). LEAKS, AND THE ACCURACIES HERE ARE OPTIMISTIC BECAUSE OF IT: each row is a 3x3 ground window, and 1,939 of the 2,000 test windows (97%) have an immediate neighbour in the training file sharing six of their nine ground pixels -- the two files are interleaved windows of one 82x100 scene, not separated regions. No 36-vector appears in both files, so byte-level deduplication passes while the test set overlaps training data on the ground; byte disjointness is not independence. Measured in variation/modalities/multispectral.md section 8. Raw class codes are 1,2,3,4,5,7 -- code 6 was withdrawn by the donors -- and are mapped positionally onto 0..5, not by subtracting one.",
"task": "landsat_statlog",
"task_name": "Landsat satellite land cover (6-class)",
"timestamp": "2026-08-03T16:02:36+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 84.713377
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.914377,
"auc": 0.980323,
"balanced_accuracy": 0.899769,
"ci95": [
0.892369,
0.908635
],
"n_test": 9752,
"score": 79.953726
},
"acquisition": null,
"calibration": {
"accuracy": 0.914377,
"brier": 0.095536,
"ece": 0.032848,
"mean_confidence": 0.936639,
"overconfidence": 0.022263,
"score": 83.575984
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.828612,
"mean_agreement": 0.945648,
"mean_kappa": 0.805737,
"mean_retention": 0.884259,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.99,
0.9808,
0.9608
],
"mean_kappa": 0.934114,
"mean_retention": 0.98611,
"retention_by_severity": [
1,
1,
0.9583
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
1,
0.9492,
0.9375
],
"mean_kappa": 0.892998,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9767,
0.9375,
0.7783
],
"mean_kappa": 0.590098,
"mean_retention": 0.666667,
"retention_by_severity": [
1,
1,
0
],
"scored": true
}
},
"score": 88.425888,
"worst_retention": 0
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.799537,
"mean_retention": 0.968404,
"per_budget": {
"n100": {
"chance_corrected": 0.971108,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.820101,
"n_labels": 40,
"retention": 1
},
"n5": {
"chance_corrected": 0.723751,
"n_labels": 10,
"retention": 0.905212
}
},
"score": 96.840405
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 74.914592,
"worst_class_recall": 0.874573
},
"uncertainty": {
"full_accuracy": 0.914377,
"risk_coverage_auc": 0.992062,
"score": 98.412349,
"selective_acc_at_50": 0.99959,
"selective_acc_at_80": 0.995642
}
},
"domain": "energy",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 6922,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / environmental sensors",
"n_classes": 2,
"n_test": 9752,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 2 (datatest2.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
"task": "occupancy_detection_2",
"task_name": "Room occupancy from sensors (period 2)",
"timestamp": "2026-08-01T19:15:19+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 92.79503
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.974859,
"auc": 0.992121,
"balanced_accuracy": 0.977365,
"ci95": [
0.971977,
0.982285
],
"n_test": 2665,
"score": 95.472947
},
"acquisition": null,
"calibration": {
"accuracy": 0.974859,
"brier": 0.041825,
"ece": 0.020148,
"mean_confidence": 0.960381,
"overconfidence": -0.014478,
"score": 89.92598
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.948521,
"mean_agreement": 0.951574,
"mean_kappa": 0.873249,
"mean_retention": 0.88191,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.995,
0.9883,
0.9675
],
"mean_kappa": 0.964616,
"mean_retention": 0.978229,
"retention_by_severity": [
1,
0.9911,
0.9436
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
1,
0.9942,
0.9967
],
"mean_kappa": 0.993464,
"mean_retention": 1,
"retention_by_severity": [
1,
1,
1
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9967,
0.995,
0.6308
],
"mean_kappa": 0.661668,
"mean_retention": 0.667501,
"retention_by_severity": [
1,
1,
0.0025
],
"scored": true
}
},
"score": 88.191024,
"worst_retention": 0.002504
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.954729,
"mean_retention": 0.929659,
"per_budget": {
"n100": {
"chance_corrected": 0.96294,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.895604,
"n_labels": 40,
"retention": 0.938071
},
"n5": {
"chance_corrected": 0.812385,
"n_labels": 10,
"retention": 0.850906
}
},
"score": 92.965901
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 93.620791,
"worst_class_recall": 0.968104
},
"uncertainty": {
"full_accuracy": 0.974859,
"risk_coverage_auc": 0.99719,
"score": 99.438032,
"selective_acc_at_50": 1,
"selective_acc_at_80": 0.99531
}
},
"domain": "energy",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 6922,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / environmental sensors",
"n_classes": 2,
"n_test": 2665,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "official released split: the donor released one training period (datatraining.txt, 8143 rows) and two separate later test periods (datatest.txt 2665 rows, datatest2.txt 9752 rows); THIS task scores test period 1 (datatest.txt) whole and never fits on it. Validation is the last 15% of the training period in time order, not a random carve -- consecutive rows are one minute apart, so a random validation set would be scored on rows whose own neighbours are in train.",
"task": "occupancy_detection",
"task_name": "Room occupancy from sensors (period 1)",
"timestamp": "2026-08-01T19:15:16+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.58,
"dimensions_scored": [
"accuracy",
"calibration",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 21.202385
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.92602,
"auc": 0.608696,
"balanced_accuracy": 0.493207,
"ci95": [
0.486554,
0.49863
],
"n_test": 392,
"score": 0
},
"acquisition": null,
"calibration": {
"accuracy": 0.92602,
"brier": 0.149519,
"ece": 0.076371,
"mean_confidence": 0.996369,
"overconfidence": 0.070348,
"score": 61.814278
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0,
"mean_agreement": null,
"mean_kappa": null,
"mean_retention": null,
"n_cases": 392,
"not_scored_reason": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 0,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9872,
0.9668,
0.9388
],
"mean_kappa": 0.081303,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"field_dropout": {
"agreement_by_severity": [
0.9949,
0.9949,
0.9643
],
"mean_kappa": 0.579204,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
},
"unit_scale": {
"agreement_by_severity": [
0.9974,
0.9974,
0.0051
],
"mean_kappa": 0.584876,
"mean_retention": 0,
"not_applicable": "the clean model scores at or below chance on this corpus (chance-corrected balanced accuracy 0.0000), so there is no retained discrimination to measure a ratio against; retention is reported as measured but not scored, because a 0.00 here would be the metric's arithmetic rather than the model's robustness",
"retention_by_severity": [
0,
0,
0
],
"scored": false
}
},
"score": null,
"worst_retention": null
},
"cost": null,
"limited_data": {
"full_reference_cc": 0,
"mean_retention": 0,
"per_budget": {
"n100": {
"chance_corrected": 0.124094,
"n_labels": 176,
"retention": 0
},
"n20": {
"chance_corrected": 0,
"n_labels": 40,
"retention": 0
},
"n5": {
"chance_corrected": 0,
"n_labels": 10,
"retention": 0
}
},
"score": 0
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0
},
"uncertainty": {
"full_accuracy": 0.92602,
"risk_coverage_auc": 0.926153,
"score": 85.230543,
"selective_acc_at_50": 0.943878,
"selective_acc_at_80": 0.933121
}
},
"domain": "industrial",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 940,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / process sensors",
"n_classes": 2,
"n_test": 392,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "chronological split minted by MedEval-1 on the production timestamp shipped in secom_labels.data: earliest 60% train, next 15% validation, latest 25% test. A random split is not defensible on a process-monitoring corpus -- tools drift and are recalibrated, so shuffling lets the model interpolate a tool's state from wafers measured minutes on either side, while production always extrapolates forward. Missing values imputed with the TRAIN median only; channels all-missing or constant on the training period dropped (468 of 590 channels kept).",
"task": "secom",
"task_name": "Semiconductor fabrication pass/fail",
"timestamp": "2026-08-01T19:13:53+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 67.058277
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.713992,
"auc": 0.932632,
"balanced_accuracy": 0.746585,
"ci95": [
0.700104,
0.79283
],
"n_test": 486,
"score": 70.434925
},
"acquisition": null,
"calibration": {
"accuracy": 0.713992,
"brier": 0.405677,
"ece": 0.067992,
"mean_confidence": 0.70528,
"overconfidence": -0.008712,
"score": 66.004218
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.704349,
"mean_agreement": 0.758116,
"mean_kappa": 0.671006,
"mean_retention": 0.661451,
"n_cases": 486,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9753,
0.9033,
0.7551
],
"mean_kappa": 0.843901,
"mean_retention": 0.868935,
"retention_by_severity": [
0.9531,
0.9134,
0.7403
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.7366,
0.6667,
0.535
],
"mean_kappa": 0.514781,
"mean_retention": 0.486918,
"retention_by_severity": [
0.7054,
0.5143,
0.241
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9774,
0.7305,
0.5432
],
"mean_kappa": 0.654336,
"mean_retention": 0.628501,
"retention_by_severity": [
0.9859,
0.6761,
0.2235
],
"scored": true
}
},
"score": 66.145126,
"worst_retention": 0.223531
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.704349,
"mean_retention": 0.925794,
"per_budget": {
"n100": {
"chance_corrected": 0.723243,
"n_labels": 571,
"retention": 1
},
"n20": {
"chance_corrected": 0.635584,
"n_labels": 140,
"retention": 0.902371
},
"n5": {
"chance_corrected": 0.616314,
"n_labels": 35,
"retention": 0.875012
}
},
"score": 92.579436
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {
"steel_type": {
"gap": 0.182203,
"per_group": {
"A300": 0.435794,
"A400": 0.617997
},
"worst": 0.435794
}
},
"basis": "worst demographic subgroup (real metadata)",
"has_real_metadata": true,
"score": 34.175968,
"worst_class_recall": 0.564356
},
"uncertainty": {
"full_accuracy": 0.713992,
"risk_coverage_auc": 0.849393,
"score": 82.429163,
"selective_acc_at_50": 0.82716,
"selective_acc_at_80": 0.771208
}
},
"domain": "industrial",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": true,
"n_train_available": 1164,
"notes": null,
"subgroup_basis": "worst demographic subgroup (real metadata)",
"train_capped": false
},
"modality": "Tabular / imaging-derived",
"n_classes": 7,
"n_test": 486,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus. The label is the argmax of the seven trailing one-hot fault columns, which are dropped from the features.",
"task": "steel_plate_faults",
"task_name": "Steel plate surface faults (7-class)",
"timestamp": "2026-08-01T19:13:43+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 59.705537
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.75,
"auc": 0.815923,
"balanced_accuracy": 0.751934,
"ci95": [
0.709575,
0.799232
],
"n_test": 400,
"score": 50.386896
},
"acquisition": null,
"calibration": {
"accuracy": 0.75,
"brier": 0.346947,
"ece": 0.061364,
"mean_confidence": 0.75821,
"overconfidence": 0.00821,
"score": 69.318061
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.503869,
"mean_agreement": 0.819722,
"mean_kappa": 0.637728,
"mean_retention": 0.708339,
"n_cases": 400,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9775,
0.9525,
0.875
],
"mean_kappa": 0.869658,
"mean_retention": 0.970815,
"retention_by_severity": [
0.9949,
0.9963,
0.9212
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.9025,
0.73,
0.6975
],
"mean_kappa": 0.553853,
"mean_retention": 0.598458,
"retention_by_severity": [
0.8669,
0.5088,
0.4197
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.96,
0.7725,
0.51
],
"mean_kappa": 0.489673,
"mean_retention": 0.555744,
"retention_by_severity": [
1,
0.6672,
0
],
"scored": true
}
},
"score": 70.833887,
"worst_retention": 0
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.503869,
"mean_retention": 0.630967,
"per_budget": {
"n100": {
"chance_corrected": 0.518842,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.449905,
"n_labels": 40,
"retention": 0.8929
},
"n5": {
"chance_corrected": 0,
"n_labels": 10,
"retention": 0
}
},
"score": 63.096663
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 44.859813,
"worst_class_recall": 0.724299
},
"uncertainty": {
"full_accuracy": 0.75,
"risk_coverage_auc": 0.844106,
"score": 68.821262,
"selective_acc_at_50": 0.825,
"selective_acc_at_80": 0.796875
}
},
"domain": "agriculture",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 959,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / physicochemical",
"n_classes": 2,
"n_test": 400,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (red wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
"task": "wine_quality_red",
"task_name": "Wine quality, red (binarised)",
"timestamp": "2026-08-01T19:13:41+00:00"
},
{
"code_fingerprint": "b444c02112cc",
"composite": {
"coverage": 0.72,
"dimensions_scored": [
"accuracy",
"calibration",
"corruption",
"limited_data",
"subgroup",
"uncertainty"
],
"index": 52.273715
},
"device": "mps",
"dimensions": {
"accuracy": {
"accuracy": 0.757551,
"auc": 0.827081,
"balanced_accuracy": 0.691134,
"ci95": [
0.664891,
0.716687
],
"n_test": 1225,
"score": 38.226844
},
"acquisition": null,
"calibration": {
"accuracy": 0.757551,
"brier": 0.315442,
"ece": 0.028736,
"mean_confidence": 0.752467,
"overconfidence": -0.005084,
"score": 85.631979
},
"corruption": {
"basis": "tabular degradation (field dropout, coding drift, unit/scale change, entry error), each at three severities",
"clean_reference_cc": 0.38084,
"mean_agreement": 0.803611,
"mean_kappa": 0.538934,
"mean_retention": 0.637729,
"n_cases": 1200,
"per_perturbation": {
"coding_drift": {
"agreement_by_severity": [
1,
1,
1
],
"mean_kappa": 1,
"mean_retention": 1,
"not_applicable": "this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result",
"retention_by_severity": [
1,
1,
1
],
"scored": false
},
"entry_error": {
"agreement_by_severity": [
0.9875,
0.935,
0.845
],
"mean_kappa": 0.805296,
"mean_retention": 0.938348,
"retention_by_severity": [
0.9867,
0.9268,
0.9015
],
"scored": true
},
"field_dropout": {
"agreement_by_severity": [
0.89,
0.7867,
0.7808
],
"mean_kappa": 0.311528,
"mean_retention": 0.318974,
"retention_by_severity": [
0.7204,
0.0906,
0.1459
],
"scored": true
},
"unit_scale": {
"agreement_by_severity": [
0.9783,
0.7942,
0.235
],
"mean_kappa": 0.499976,
"mean_retention": 0.655867,
"retention_by_severity": [
0.9676,
1,
0
],
"scored": true
}
},
"score": 63.772944,
"worst_retention": 0
},
"cost": null,
"limited_data": {
"full_reference_cc": 0.382268,
"mean_retention": 0.666667,
"per_budget": {
"n100": {
"chance_corrected": 0.410205,
"n_labels": 200,
"retention": 1
},
"n20": {
"chance_corrected": 0.454601,
"n_labels": 40,
"retention": 1
},
"n5": {
"chance_corrected": 0,
"n_labels": 10,
"retention": 0
}
},
"score": 66.666667
},
"spec_sensitivity": null,
"subgroup": {
"attributes": {},
"basis": "worst diagnostic class (no demographic metadata in source)",
"has_real_metadata": false,
"score": 0,
"worst_class_recall": 0.490244
},
"uncertainty": {
"full_accuracy": 0.757551,
"risk_coverage_auc": 0.883347,
"score": 76.669329,
"selective_acc_at_50": 0.897059,
"selective_acc_at_80": 0.811224
}
},
"domain": "agriculture",
"kind": "tabular",
"limitations": {
"has_real_subgroup_metadata": false,
"n_train_available": 2938,
"notes": null,
"subgroup_basis": "worst diagnostic class (no demographic metadata in source)",
"train_capped": false
},
"modality": "Tabular / physicochemical",
"n_classes": 2,
"n_test": 1225,
"runtime_s": 0.1,
"seed": 20260727,
"spec_version": "MedEval-1 v0.1",
"split_origin": "minted by MedEval-1 (stratified 60/15/25, seed 20260727); no official split is published for this corpus (white wine, Vinho Verde). The binary target is an ANALYST CHOICE made here, not a property of the release: the donors ship an ordinal 0-10 sensory score and this task binarises it at quality >= 6. A different cut would give a different positive rate and a different board. The raw quality column is dropped from the features.",
"task": "wine_quality_white",
"task_name": "Wine quality, white (binarised)",
"timestamp": "2026-08-01T19:14:08+00:00"
}
],
"heldout": null,
"sources": [
{
"path": "results/cross_site/occupancy__tab_logreg.json",
"sha256": "39f6dd624d9723016056f3deb5caf4f5be5c852ab6ec756c45afd28732dd3d52"
},
{
"path": "results/heldout/audit_report.json",
"sha256": "df460b6d91a6512dfa1e7dc650637fce2b4dd3a1ec32e03446bf6e11f9c597d2"
},
{
"path": "results/heldout/cycle.json",
"sha256": "e012b8978457b51b5c31d18d14d2cc92f100731897523306acae19934f943502"
},
{
"path": "results/heldout/manifests/gen1.seal.json",
"sha256": "91a7adfa29870a0719125cdd48f6dcb149f1f6485ac1cfa104c6e00057ad9b81"
},
{
"path": "results/heldout/manifests/gen2.seal.json",
"sha256": "f61470918a09cf39e074ef14fb7ade4ac3a13d36d9385f18cf9b92e4b5329547"
},
{
"path": "results/heldout/manifests/public.seal.json",
"sha256": "1a4fd930d2e781ac7163483990f15d6694e28e507d1b3ca009ff09395e47497c"
},
{
"path": "results/records/adult_income__tab_logreg.json",
"sha256": "9d49f868e37c94bdcbde73d8ea4f043ccd8d8fc63c89550f8e2d95605e67deb0"
},
{
"path": "results/records/anuran_mfcc__tab_logreg.json",
"sha256": "c1271e3ccaa18a559772a12ca02fb9265413e70bb06f299c2c87b78c9f5ac6f0"
},
{
"path": "results/records/breast_wdbc__tab_logreg.json",
"sha256": "9756b728ad8a36e7bd2806f46f3fb6b9c2bd8590a97bcf5c22fc66f28fb165be"
},
{
"path": "results/records/coil2000_caravan__tab_logreg.json",
"sha256": "252f17a68d072d2cb05ee5514df2f21bdf1d6caa45904c07260ed5c59778507b"
},
{
"path": "results/records/diabetes_pima__tab_logreg.json",
"sha256": "d8c33c2b3ecb291523db6e86a5fb15de72eab3df9c9f8050abf5a97a862441d3"
},
{
"path": "results/records/dry_bean__tab_logreg.json",
"sha256": "746ef76d9275ef54925dba9afd291e7dcc4986601dfbf86802325ca2bffa87ee"
},
{
"path": "results/records/german_credit__tab_logreg.json",
"sha256": "ef0656154fcaed4e3ec40dfe3f52387aed341910b42f970cb4feb3ea9a0dcdd9"
},
{
"path": "results/records/heart_cleveland__tab_logreg.json",
"sha256": "ee6774fabcb5b6f03727dfba90c2b3081f6f98b06172ecfc2f9e4241b3432ee2"
},
{
"path": "results/records/landsat_statlog__tab_logreg.json",
"sha256": "2fe08d5aceaf2cfc8e5eb9cbdf58169f99876093ad34aea9e16de93890ee949f"
},
{
"path": "results/records/occupancy_detection_2__tab_logreg.json",
"sha256": "33859c9106ad2179d4b479a3b53c112f6e1e36dfd7324b915658e8ab4f751bce"
},
{
"path": "results/records/occupancy_detection__tab_logreg.json",
"sha256": "7a7adae2920572b0569780fdcc65883daabc66360174c0989ebfb3727800090d"
},
{
"path": "results/records/secom__tab_logreg.json",
"sha256": "5e2c3e5e4a5dea390dd9968e2121956ee3e704ff150bd3a93301bd52d299e797"
},
{
"path": "results/records/steel_plate_faults__tab_logreg.json",
"sha256": "bcb35954c6c21b3a1a223d12a924f3ff3ca1cebc8016823872b1c305aa3b4e6b"
},
{
"path": "results/records/wine_quality_red__tab_logreg.json",
"sha256": "528c7eedc18d0baf7620c8f7ba6e36991a41f7675d52c150ca92ce4ae23452fe"
},
{
"path": "results/records/wine_quality_white__tab_logreg.json",
"sha256": "a1caccca58fe53897a0124cf62fc019fe3b4bc6b3a15aee3d5c22da9d3460e68"
}
]
},
"declared": null,
"declared_note": "No vendor declaration. This is an open reference baseline implemented by NakedSignal, not a commercial product: there is no vendor to assert an intended use or to sign an attestation, so the declared block is null rather than filled in on the vendor's behalf.",
"expiry_basis": "the sealed held-out generation this evidence cycle is anchored to (gen2, set hash ad30d214daa285c5...) is published to expire 2026-09-20; per spec, a passport lapses when its generation retires or its spec version is superseded, whichever comes first",
"expiry_utc": "2026-09-20T00:00:00Z",
"issued_utc": "2026-08-03T16:03:42Z",
"issuer": {
"algo": "Ed25519",
"key_id": "ns-passport-2026-07",
"name": "NakedSignal",
"public_key_b64": "0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw="
},
"observed": null,
"observed_note": "Not available: this requires continuous monitoring. Production history and drift are observed fields, and nothing is deployed, so there is nothing to observe. The object is present and null on purpose. The shape is the roadmap, not a claim.",
"passport_version": "0.1",
"signature": "tn7ly4md5fR34HyiyAR06s1iIi2FMO5XyDGxSdbQo7oNEkv7iVBFGdtTBpbHgSUJukFh0VOYCZoZLBxLlsDxCA==",
"spec_version": "MedEval-1 v0.1",
"subject": {
"behaviour_version": "1.0.0",
"code_fingerprint": null,
"code_fingerprints": [
"b444c02112cc",
"de32bca72d56"
],
"family": "classical",
"model_id": "tab_logreg",
"model_name": "Logistic regression",
"params": "standardised, L2, C=1"
}
}Canonicalisation: RFC 8785 (JCS)-compatible: UTF-8, object keys sorted by code point, no insignificant whitespace, ECMAScript number formatting. every float is rounded once at emit to 6 decimal places (Python round(), banker's rounding); integers are left exact. Signed over the whole document with the `signature` member removed, canonicalised as above and encoded UTF-8. The public key for ns-passport-2026-07 is 0PBCC6mjzcppR5QdPcK7vP/AamB6MP8l1T4ypMYbsMw= — see Governance.