Platform overview
LIVEWhat the console is built on: what we measure, what we issue, what we enforce, and why the whole thing is a standing function rather than a project.
NakedSignal builds the layer that decides whether a model’s output can be acted on.
We started in labour-market data, where the question was whether a signal was reliable enough to underwrite on.
We now run the same instrument in medicine, where the question is whether it is reliable enough to treat on.
Every public benchmark rots. It gets published, becomes the reference, gets scraped into the next model’s training data — and from then on a high score measures memorisation.
So evaluation is not a project that finishes. It is a standing function: keep minting sealed sets, keep scoring against them, keep publishing what fails. We run that function continuously, with one instrument, across domains. This site is that function’s output rather than a description of it.
The current run
Read frompayload.json, the same file the MedEval-1 leaderboard renders from.One instrument, three domains
Counted out ofpayload.json at build time.The claim is not that this is a cross-domain company. It is that the instrument is domain-general, and that it is running on three domains at once — the same 9 dimensions, the same evaluators, the same seed and code fingerprint. The empty cells are the point: each one is a question this corpus cannot answer, marked rather than filled.
| Domain | 1 Task accuracy | 2 Cross-site robustness | 3 Calibration | 4 Limited data | 5 Subgroup safety | 6 Corruption / OOD | 7 Uncertainty & deferral |
|---|---|---|---|---|---|---|---|
Medicine 134 records | 134 scored | 111 scored | 134 scored | 122 scored | 134 8 strong | 130 scored | 134 scored |
Agriculture & food 12 records | 12 scored | — not measured | 12 scored | 12 scored | 12 0 strong | 12 scored | 12 scored |
Labour market 11 records | 11 scored | — not measured | 11 scored | 11 scored | 11 scored | 11 scored | 11 scored |
Energy & buildings 8 records | 8 scored | — not measured | 8 scored | 8 scored | 8 0 strong | 8 scored | 8 scored |
Financial 8 records | 8 scored | — not measured | 8 scored | 8 scored | 8 4 strong | 8 scored | 8 scored |
Industrial inspection 8 records | 8 scored | — not measured | 8 scored | 8 scored | 8 4 strong | 4 scored | 8 scored |
Legal 7 records | 7 scored | — not measured | 7 scored | 7 scored | 7 0 strong | 7 scored | 7 scored |
Earth observation 4 records | 4 scored | — not measured | 4 scored | 4 scored | 4 0 strong | 4 scored | 4 scored |
Ecology & bioacoustics 4 records | 4 scored | — not measured | 4 scored | 4 scored | 4 scored | 4 scored | 4 scored |
- Medicine · 5 Subgroup safety 126 of 134 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 8 had the inputs the strong reading needs.
- Agriculture & food · 5 Subgroup safety 12 of 12 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
- Energy & buildings · 5 Subgroup safety 8 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
- Financial · 5 Subgroup safety 4 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 4 had the inputs the strong reading needs.
- Industrial inspection · 5 Subgroup safety 4 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 4 had the inputs the strong reading needs.
- Legal · 5 Subgroup safety 7 of 7 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
- Earth observation · 5 Subgroup safety 4 of 4 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
- Agriculture & food · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Labour market · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Energy & buildings · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Financial · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Industrial inspection · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Legal · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Earth observation · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
- Ecology & bioacoustics · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
What a standing function issues
A score is an event. What accumulates from running the function is infrastructure — and it is the accumulation, not any single number, that is the company.
identity — which model, scored how, valid until when
relationships — what shares a corpus, what travels, what failed
enforcement — the point where a certificate stops being paperwork
the standing judgement the other three make possible
Institutions of this kind — index providers, ratings agencies, standards bodies — are paid for continuity rather than for a project, and they only work if the party issuing the evidence is structurally separate from the party being measured. That separation is a governance design, not a promise, and it is written down: the ledger split and what already runs on each side →