NakedSignal OS · Documentation
Back to the consoleDocumentation

Platform overview

LIVE

What the console is built on: what we measure, what we issue, what we enforce, and why the whole thing is a standing function rather than a project.

NakedSignal builds the layer that decides whether a model’s output can be acted on.

We started in labour-market data, where the question was whether a signal was reliable enough to underwrite on.

We now run the same instrument in medicine, where the question is whether it is reliable enough to treat on.

Every public benchmark rots. It gets published, becomes the reference, gets scraped into the next model’s training data — and from then on a high score measures memorisation.

So evaluation is not a project that finishes. It is a standing function: keep minting sealed sets, keep scoring against them, keep publishing what fails. We run that function continuously, with one instrument, across domains. This site is that function’s output rather than a description of it.

The current run

Read from payload.json, the same file the MedEval-1 leaderboard renders from.
196
evaluations
22
models
35
tasks
21
modalities
1,058,247
held-out cases scored
4,640,688
perturbed evaluations
spec MedEval-1 v0.1code fingerprint de32bca72d56seed 20260727device mpscompute 430.6 mingenerated 2026-08-03T19:03+00:00

One instrument, three domains

Counted out of payload.json at build time.

The claim is not that this is a cross-domain company. It is that the instrument is domain-general, and that it is running on three domains at once — the same 9 dimensions, the same evaluators, the same seed and code fingerprint. The empty cells are the point: each one is a question this corpus cannot answer, marked rather than filled.

Domain
1
Task accuracy
2
Cross-site robustness
3
Calibration
4
Limited data
5
Subgroup safety
6
Corruption / OOD
7
Uncertainty & deferral
Medicine
134 records
134
scored
111
scored
134
scored
122
scored
134
8 strong
130
scored
134
scored
Agriculture & food
12 records
12
scored
not measured
12
scored
12
scored
12
0 strong
12
scored
12
scored
Labour market
11 records
11
scored
not measured
11
scored
11
scored
11
scored
11
scored
11
scored
Energy & buildings
8 records
8
scored
not measured
8
scored
8
scored
8
0 strong
8
scored
8
scored
Financial
8 records
8
scored
not measured
8
scored
8
scored
8
4 strong
8
scored
8
scored
Industrial inspection
8 records
8
scored
not measured
8
scored
8
scored
8
4 strong
4
scored
8
scored
Legal
7 records
7
scored
not measured
7
scored
7
scored
7
0 strong
7
scored
7
scored
Earth observation
4 records
4
scored
not measured
4
scored
4
scored
4
0 strong
4
scored
4
scored
Ecology & bioacoustics
4 records
4
scored
not measured
4
scored
4
scored
4
scored
4
scored
4
scored
n records scored on this dimensionn scored, some on a weaker basis the record declares not defined for this corpus, and not approximated
  • Medicine · 5 Subgroup safety 126 of 134 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 8 had the inputs the strong reading needs.
  • Agriculture & food · 5 Subgroup safety 12 of 12 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
  • Energy & buildings · 5 Subgroup safety 8 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
  • Financial · 5 Subgroup safety 4 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 4 had the inputs the strong reading needs.
  • Industrial inspection · 5 Subgroup safety 4 of 8 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 4 had the inputs the strong reading needs.
  • Legal · 5 Subgroup safety 7 of 7 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
  • Earth observation · 5 Subgroup safety 4 of 4 scored on a weaker basis the evaluator declares in the record: worst diagnostic class (no demographic metadata in source). 0 had the inputs the strong reading needs.
  • Agriculture & food · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Labour market · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Energy & buildings · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Financial · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Industrial inspection · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Legal · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Earth observation · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.
  • Ecology & bioacoustics · 2 Cross-site robustness No re-acquisition of the same subjects exists in this corpus, so there is nothing to measure. Widening the definition to fill the cell would be a spec change, not a demo.

What a standing function issues

A score is an event. What accumulates from running the function is infrastructure — and it is the accumulation, not any single number, that is the company.

Institutions of this kind — index providers, ratings agencies, standards bodies — are paid for continuity rather than for a project, and they only work if the party issuing the evidence is structurally separate from the party being measured. That separation is a governance design, not a promise, and it is written down: the ledger split and what already runs on each side →