Census income
LIVEThe same instrument on census income prediction — the strongest split provenance in the suite, and the corpus where dimension five is measured against real demographic subgroups rather than a class proxy. No separate specification document has been published for this benchmark — the MedEval-1 spec is the document, and this page is that instrument run on another corpus. Everything below is read from the same payload.json the medical leaderboard renders from.
The index is the weighted mean of the dimensions that were scored, on the weights published in the spec. Dimensions this corpus cannot support are excluded from it rather than counted as zero, which is why an index here is not comparable to a medical one model-for-model — it is a mean over a different set of questions, and the table below says which.
Every model, on all 9 dimensions
Ranked by index. An em-dash is a dimension the harness left null.| # | Model | Index | 1 Accuracy | 2 Cross-site | 3 Calibration | 4 Limited data | 5 Subgroup | 6 Corruption | 7 Uncertainty | Specification sensitivity | Cost, latency & footprint |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GTE-base frozen + linear probe109M frozen, mean-pooled 768-d, logistic head; contrastive sentence objective, general domain | 80.3 | 70.6 | — | 92.2 | 79.5 | 51.3 | 94.6 | 99.1 | — | 99.6 |
| 2 | TF-IDF + linear SVM (Platt-scaled)squared hinge, C=1, 5-fold Platt calibration | 78.4 | 76.3 | — | 98.3 | 72.5 | 25.0 | 96.8 | 99.8 | — | — |
| 3 | TF-IDF + MLP (256)SVD-256 then 1 hidden layer of 256 | 75.3 | 73.0 | — | 97.8 | 35.7 | 40.0 | 97.2 | 99.8 | — | — |
| 4 | PubMedBERT-sentence frozen + linear probe109M frozen, mean-pooled 768-d, logistic head; biomedical pretraining PLUS a sentence objective | 75.2 | 71.6 | — | 95.3 | 63.5 | 22.4 | 93.1 | 98.6 | — | 99.8 |
| 5 | Gradient boosting200 stages, depth 3, lr 0.05 | 72.5 | 55.9 | — | 91.6 | 80.6 | 49.3 | 85.7 | 93.1 | — | — |
| 6 | Logistic regressionstandardised, L2, C=1 | 70.6 | 53.3 | — | 97.1 | 79.1 | 29.4 | 90.6 | 91.1 | — | — |
| 7 | BiomedBERT frozen + linear probe109M frozen, mean-pooled 768-d, logistic head | 69.9 | 66.1 | — | 92.2 | 73.2 | 0.0 | 86.4 | 98.7 | — | 98.9 |
| 8 | TF-IDF + logistic regression1-2 grams, 200k features, sublinear tf | 69.6 | 59.4 | — | 93.8 | 84.8 | 0.0 | 93.5 | 99.7 | — | — |
| 9 | Random forest400 trees | 69.4 | 52.8 | — | 85.0 | 90.0 | 39.2 | 83.8 | 89.6 | — | — |
| 10 | MLP (128, 64)2 hidden layers, adam, early stopping when n>=200 | 69.1 | 50.8 | — | 92.3 | 87.7 | 24.9 | 89.7 | 90.7 | — | — |
| 11 | TF-IDF + complement naive Bayes1-2 grams, alpha=0.3 | 67.7 | 52.5 | — | 88.7 | 87.0 | 0.0 | 98.1 | 99.3 | — | — |
- 2. Cross-site robustness — No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- . — Not computed on any record in this benchmark, and not approximated.
The measurements the scores came from
Raw metrics, not the 0–100 rescaling.| Model | Balanced accuracy (95% CI) | AUC | ECE | Risk–coverage AUC | Mean label-budget retention |
|---|---|---|---|---|---|
| GTE-base frozen + linear probe | 0.853[0.824–0.882] | 0.949 | 0.016 | 0.995 | 0.795 |
| TF-IDF + linear SVM (Platt-scaled) | 0.882[0.855–0.909] | 0.986 | 0.003 | 0.999 | 0.725 |
| TF-IDF + MLP (256) | 0.865[0.838–0.893] | 0.983 | 0.004 | 0.999 | 0.357 |
| PubMedBERT-sentence frozen + linear probe | 0.858[0.825–0.884] | 0.931 | 0.009 | 0.993 | 0.635 |
| Gradient boosting | 0.779[0.773–0.789] | 0.919 | 0.017 | 0.965 | 0.806 |
| Logistic regression | 0.766[0.759–0.775] | 0.901 | 0.006 | 0.955 | 0.791 |
| BiomedBERT frozen + linear probe | 0.831[0.798–0.859] | 0.936 | 0.016 | 0.994 | 0.732 |
| TF-IDF + logistic regression | 0.797[0.762–0.829] | 0.984 | 0.012 | 0.998 | 0.848 |
| Random forest | 0.764[0.757–0.772] | 0.889 | 0.030 | 0.948 | 0.900 |
| MLP (128, 64) | 0.754[0.747–0.763] | 0.897 | 0.015 | 0.953 | 0.877 |
| TF-IDF + complement naive Bayes | 0.763[0.730–0.794] | 0.949 | 0.023 | 0.996 | 0.870 |
Balanced accuracy carries a bootstrap 95% interval; read it before reading the rank. ECE is expected calibration error, so lower is better and the 0–100 dimension score inverts it. Every value here is on the record; hover any figure for the unrounded number.
Who the best model fails
LIVEDimension five, on GTE-base frozen + linear probe — the top-ranked model above. Basis, in the evaluator’s own words: worst demographic subgroup (real metadata). This is the strong reading — the corpus ships the attributes, so the harness did not have to fall back to the worst-class proxy every medical imaging task uses.
| AUworst | 0.756 | |
| unknown | 0.800 | |
| US | 0.852 | |
| CA | 0.875 | |
| GB | 0.912 | |
| IN | 0.986 |
| High School or equivalentworst | 0.794 | |
| not specified | 0.852 | |
| Unspecified | 0.859 | |
| Bachelor's Degree | 0.876 | |
| Certification | 1.000 | |
| Master's Degree | 1.000 | |
| Professional | 1.000 | |
| Some College Coursework Completed | 1.000 |
Per-group recall, and the gap between the best and worst group. These numbers measure who the model fails on this corpus. They are not an audit of any deployed system, and the subgroup attributes are model inputs here, as they are in the standard formulation of the task.
The corpus
Census income (>$50K)
adult_income · Tabular / socioeconomic- Source
- UCI Adult / Census IncomeKohavi 1996 / UCI ML Repository, Adult
- Split
- official released split (UCI adult.test scored whole); validation carved from adult.data only (stratified 15%, seed 20260727)
- Sizes
- 12,000 train · 16,281 held out · 2 classesTraining capped at the standard budget from 27,676 available rows, so every model in the suite sees the same amount of data.
- Classes
- <=50K · >50K
- Registry note
- Official released train/test split -- the test file is used whole and validation is carved from the official train file only. Carries sex, race and age, so dimension 5 is measured against real demographic subgroups rather than the worst-class fallback. Dimension 2 (cross-site robustness) is undefined: there is no re-acquisition of the same subjects to measure. Known critique: Ding et al., "Retiring Adult: New Datasets for Fair Machine Learning" (NeurIPS 2021) argues this 1994 census extract is outdated as a fairness benchmark and proposes ACS/folktables as the replacement; that task is on the roadmap and this one is published with the criticism attached, not without it.
Fraudulent job posting detection
jobpost_fraud · Recruitment text- Source
- EMSCAD (Employment Scam Aegean Dataset)Vidros, Kolias, Kambourakis & Akoglu, Future Internet 9(1):6, 2017
- Split
- no official split published; 60/15/25 train/val/test minted at the standard seed (20260727), stratified on the 4.8% positive class
- Sizes
- 10,728 train · 4,470 held out · 2 classes
- Classes
- genuine · fraudulent
- Registry note
- 4.8% of 17,880 postings are fraudulent, so a 25% held-out split carries roughly 215 positives and the intervals are wide by construction. LICENCE: the HuggingFace mirror tags this CC0, but the underlying EMSCAD release from the University of the Aegean is CC BY-NC-SA 4.0. It is treated as non-commercial here and cited to the original. Subgroups are recorded fields (country, required education), not inferred attributes.
Labour market
The same dimensions, pointed at census income prediction.
The census income task is scored by the identical harness that scores the medical suite — same evaluators, same seed, same chance-correction, same record format. Nothing in the scoring code knows what a patient is, which is why adding a domain is a registry row and an adapter branch rather than a second product.
It is also the strongest split provenance in the whole suite. The corpus ships an official train/test split, so the held-out file is scored whole and validation is carved from the training file alone — no minted split, no seed to trust. Every task card states which of the two it is.
And because the corpus carries sex, race and age, dimension five is measured against real demographic subgroups here rather than the worst-class proxy the imaging tasks fall back to.
- The corpus is a 1994 census extract and has a published critique attached: Ding et al., "Retiring Adult: New Datasets for Fair Machine Learning" (NeurIPS 2021) argues it is outdated as a fairness benchmark and proposes ACS/folktables as the replacement. That task is on the roadmap. This one ships with the criticism cited on its card, which is the only version of shipping it that a standard can defend.
- Dimension two is undefined here — there is no re-acquisition of the same subjects to measure — and renders as an em-dash rather than a zero. Dimension six is measured: field dropout, coding drift, unit change and entry error, at three severities. Coding drift is reported unscored on this corpus, because its categoricals arrive pre-encoded as ordinals and the transform has no one-hot block to act on.
- The subgroup attributes are also model inputs, as they are in the standard formulation of this task. The dimension-five numbers measure who the model fails; they are not an audit of a deployed lending or hiring system.
Re-derive every figure on this page with python src/build_site.py --emit-json in plan_d/ — the same run writes the medical leaderboard. All benchmarks reads the same 9 dimensions per domain, and the trust graph traces any figure here back to the run that produced it.