Synthetic variation
PROTOTYPEA model is built in one place and used in another. The gap between those two places is measurable, and most of it is boring: lighting, sensor age, a different convention at the console. This page measures how much a model changes its mind when only those things change, and specifies how that measurement becomes a generator.
1. The gap nobody prices
A model is validated where it was built. It is then deployed somewhere that differs in a hundred small ways, and the headline number does not travel with it. This is not a medical problem, an exotic problem, or a problem of bad models. It is the ordinary condition of putting software that learned from one sample into a world that keeps producing new ones.
The industry response is usually to re-collect and re-label at the new site, which is slow, expensive, and has to be repeated at the site after that. The alternative is to treat the difference between sites as a measurable quantity rather than an accident — to characterise how much a deployment can vary, and then make the model meet that variation before it ships.
The chest film that travels badly
Built on — A detector trained on one hospital's images, validated on held-out patients from that same hospital, and reported at a single headline number.
Deployed into — A second hospital, a different make of unit, an older tube, a warmer room, and a radiographer who positions to a different convention.
- detector response curve and noise floor
- window/level chosen at the console
- sensor resolution and field of view
- ambient light where the study is read
Instrumented here, on simulated re-acquisition over public data. This is the column the figures below are computed from.
The screener that learned one hiring cycle
Built on — A CV screening model fitted to one country's applications during one part of the cycle, scored against the hires that were actually made.
Deployed into — Another region a year later: different résumé conventions, different job-title vocabulary, different tenure norms, and a labour market that has itself moved.
- document formatting and parsing conventions
- job-title and qualification vocabulary
- tenure and gap norms between regions
- the base rate the model was calibrated against
Degradation is instrumented here, on UCI Adult: field dropout, coding drift, unit change and entry error, at three severities, in §5. Re-acquisition is not — the census corpus has no construction for it, so dimension 2 keeps an explicit null rather than a widened definition.
The credit model that met a different bureau
Built on — A default model fitted on one bureau's field coverage across a calm period, with every feature present and every field populated.
Deployed into — A portfolio where three fields are coded differently, one is absent entirely, and the macro conditions that set the base rate have changed.
- field coverage and missingness patterns
- coding conventions between providers
- the macro regime behind the base rate
- entry error and late-arriving corrections
Instrumented here, on UCI German Credit. The tabular degradation family — field dropout, coding drift, unit change, entry error — is implemented and run; the figures are in §5. Re-acquisition is not measured: a re-coded credit file is not a second observation of the borrower, so dimension 2 keeps an explicit null.
The inspection rig that only worked in the lab
Built on — A visual quality model trained under a fixed lighting rig, one camera, one working distance, one bench.
Deployed into — A plant floor: daylight through a roof panel, a replacement camera two revisions newer, vibration, and a housing that runs fifteen degrees hotter by the afternoon shift.
- illumination spectrum and intensity over a shift
- camera revision, gain and exposure defaults
- working distance and mounting drift
- thermal noise as the housing warms
Not measured here. It is the closest analogue to the imaging construction and the cheapest place to acquire genuine paired data, because both conditions can be staged deliberately.
2. What actually varies, written down
For images the harness carries 12 transforms in two families, kept separate because they answer different questions. The re-acquisition family changes how the input was made — a different machine, a different setting, a different room. The degradation family changes how badly it was made. The distinction matters: in the first family the correct answer is unchanged by construction, so any change of mind is pure site sensitivity and not a harder case.
Tables and documents carry degradation families of their own — 4 and 5 respectively, listed in §5. They have no re-acquisition family at all: nothing in a credit file or a discharge note re-measures the subject, so dimension 2 stays an explicit null on every one of those corpora rather than being widened to cover paperwork. Their figures are kept in a separate section for the same reason, because a retention number for a gamma shift and one for a unit swap are not two samples of one quantity.
- Gamma
- stands in for a different detector's response curve
- Window / level
- stands in for a different reading protocol at the console
- Resolution
- stands in for an older or newer sensor at the same site
- Field of view
- stands in for a different positioning convention
- Detector noise floor
- stands in for a hotter room, an ageing tube, a cheaper unit
- Gaussian noise
- stands in for electrical noise in the capture chain
- Shot noise
- stands in for a lower dose or a shorter exposure
- Defocus blur
- stands in for motion, or a focus the operator did not catch
- Contrast
- stands in for a mis-set display or export pipeline
- Brightness
- stands in for ambient light in the room where it was captured
- Pixelate
- stands in for a downscaled copy pasted between systems
- Quantise
- stands in for bit-depth lost to a lossy archive format
3. How much a model changes its mind
Each transform is applied at three severities to every model and task, and the model’s answers are compared against its own answers on the clean input. Retention is that agreement, chance-corrected: 1.00 means the shift cost nothing, 0.00 means none of the original agreement survived. It is not accuracy — a model can be mediocre and stable, or excellent and brittle, and those are different products.
Retention under each shift
1.00 = the shift cost nothing · mean 0.70bar = mean retention across every model × task that measured it · ticks = severity 1 / 2 / 3 · hairline = mean across the 12 scored
This was pre-registered to be published whichever way it came out, and it came out this way. The measurement, its limits and its confounds are in
variation/ docs/ REPORT- S2- realism- audit. md. Nothing above is withdrawn — every figure is still the figure that ran. What changed is what it means.The ordering is the interesting part. The shifts that cost the most are not the dramatic ones. Brightness and contrast — the ambient light in the room and the setting on a display — sit at the bottom, below sensor noise and below blur. A model can be robust to a genuinely bad image and still change its mind because someone dimmed the lights. That is not a story about model capacity; it is a story about which variation was present in the training sample and which was not.
4. Versatility is a property you can measure
Clean accuracy
before any shift · 0-100
Retention, re-acquired
same subject, different machine
Retention, degraded
same machine, worse capture
5. The same measurement, on tables and documents
Everything above is imaging. The identical construction runs on the tabular and text corpora, and it is the reason the labour-market and financial cards in §1 are no longer empty. What degrades a table is not a lens or a lamp — it is the paperwork: a field the receiving site does not collect, a category coded under a different scheme, a measurement in different units, a human mistyping a digit. What degrades a document is the pipeline it travelled through.
These are dimension 6 only. There is no dimension 2 here and the harness refuses to invent one: re-acquisition means the same subject measured again elsewhere, and a re-coded credit file is not a second measurement of the borrower. Dimension 2 is an explicit null on all 18 of these corpora.
Retention under each shift
1.00 = the shift cost nothing · mean 0.79this corpus has no one-hot encoded categorical block, so there is no coding scheme to drift; the transform is a no-op here and its retention of 1.00 is an artefact, not a result
bar = mean retention across every model × task that measured it · ticks = severity 1 / 2 / 3 · hairline = mean across the 3 scored
The ordering carries the finding. A missing field and a unit swap cost more than a transcription error, and by a wide margin — a model that is perfectly stable under noisy data entry will still change its mind when a column arrives in different units, because nothing in training told it which unit it was reading. This is the tabular restatement of the imaging finding: the cheap, boring, administrative differences are the expensive ones.
Retention under each shift
1.00 = the shift cost nothing · mean 0.96† Note cut short at severity 3: the transform may remove the evidence the label depends on, not merely obscure it; retention at these severities is not a clean robustness measurement. The bar above averages all three severities, so it includes that one.
bar = mean retention across every model × task that measured it · ticks = severity 1 / 2 / 3 · hairline = mean across the 5 scored
6. From measuring to closing PROTOTYPE
Everything above is a measurement. The measurement is also, already, a generator — the transforms that probe a model are the same transforms that could train it. What is missing is the discipline that makes that honest rather than circular.
Training on our published transforms scored 0.1384 at the held-out condition, against 0.1394 for no augmentation at all — it did not help, and it cost 0.045 of clean accuracy to get there. An ordinary off-the-shelf augmentation pipeline scored 0.1610, beating both. The fitted envelope scored 0.1542, which is approximately the null its own fit predicted.
The gap between how robust each arm looks on our own generator and how it performs at a real acquisition change is 0.66–0.68, and no arm meaningfully closes it despite differing by sixteenfold in training data. That gap is a property of the corpus, not of the augmentation. Step 5 below — verify on a site the model has never seen — is the step that produced this, and it is the reason it is written as load-bearing.
- Measure the envelope, from real sites. Collect the same subjects, or the same units, through the machines and conditions the model will actually meet. Fit the distribution of each parameter — how much gamma, how much noise floor, how much positioning drift — rather than assuming a range. This step needs paired multi-site data and is the only expensive part.
- Publish the envelope as part of the specification. A range of conditions a model claims to cover is a falsifiable statement. It belongs in the passport beside the scores, so a buyer can check their own site against it before purchase rather than after.
- Generate inside the measured envelope, seeded. Sample the fitted parameters to synthesise variation that spans the range the deployment can produce. Determinism is not optional: the generator version and seed go in the passport, so the training distribution is reproducible by someone who does not trust us.
- Train against it, and re-measure on the same harness. Retention is re-computed with the identical transforms and severities, so before and after are comparable numbers rather than two different experiments.
- Verify on a real site the model has never seen. This is the load-bearing step. Improvement measured on synthetic variation drawn from the same generator that produced the training data is not evidence — it is the model learning the generator. The claim is only admissible if it survives a held-out physical site, held to the same separation rules as the rotating held-out track.
7. What separates this from data augmentation, after testing it
Augmentation perturbs training data to make a model less fragile, and it has been standard practice for a decade. This section used to claim three things separated the protocol above from turning on a default augmentation pipeline. We tested that claim against an off-the-shelf pipeline at a held-out real condition and one of the three did not survive. All three are still listed, with the failure first and marked, because a differentiator that failed its test is more informative than one that was quietly dropped.
- Calibration was supposed to be the difference. We tested that, and it is not. Standard augmentation applies a range someone chose; the design in section 6 fits the range to measured variation between real sites, which makes it a claim about the world that can be shown to be wrong. It was, on both halves. The fitted range came back as essentially the identity — the real acquisition difference is not in the parameters our forward model exposes, so there was nothing to calibrate to. And at a held-out real condition an ordinary off-the-shelf augmentation pipeline beat our published transforms (0.1610 against 0.1384 balanced accuracy) and beat the fitted envelope (0.1542) too.
So calibration is not currently the differentiator, and this bullet no longer claims it is. What survives the test is the two bullets below — off-generator verification and a signed record — and those are the reason we could run this experiment against ourselves and report it. A vendor whose pipeline is not verified off-generator would not have found this out, which is the argument now, and it is a smaller and truer one. - The result is verified off-generator. The improvement must appear at a physical site that contributed no training variation. Anything else measures the generator.
- The whole loop is recorded in a signed document. Envelope, generator version, seed, before-and-after retention, and the site the verification ran at — all inside the evidence passport, so the claim travels with the model instead of living in a slide.
8. What this cannot do
The limits are not marginal and are worth stating before anyone buys against them.
- It cannot invent a shift nobody measured. Synthetic variation covers the parameters in the envelope. A site that differs in a way no one thought to parameterise is outside it, and the method is silent there rather than protective.
- It assumes the label survives the transform. That holds by construction for re-acquisition — changing how a scan was made does not change what it shows. It stops holding if a transform is pushed far enough to destroy the finding, which is why severity is bounded and graded rather than open-ended.
- It does not make a weak model strong. It narrows the gap between where a model was built and where it runs. A model that was never good enough at the source site does not become good enough by being made stable.
- Today’s numbers are a stand-in, and the stand-in has been tested. The shift on this page is simulated over single-source public data. Every figure here describes a model’s sensitivity to our transforms, which was a hypothesis about its sensitivity to real sites. That hypothesis has now been measured against a real acquisition change and it does not hold — see section 3. So this bullet is no longer a caution about something unknown; it is a statement about something known. Retention under these transforms should not be read as evidence about a change of scanner until the envelope is fitted and the loop is closed off-generator.
9. What it would take
The expensive input is not compute and not modelling. It is paired data: the same subjects or the same units, captured under two or more genuinely different conditions, with the conditions recorded. Everything downstream of that — fitting the envelope, generating inside it, retraining, re-measuring, and signing the result — is already built or specified in this repository.
The cheapest place to acquire it is not medicine. It is any setting where both conditions can be staged deliberately: an inspection rig moved between two benches, one device revision against its successor, one room in the morning and the same room in the afternoon. A single such pairing is enough to replace the simulated envelope with a measured one, and to turn every figure on this page from a hypothesis into a result.
Related: the console’s Variation studio puts the fitted envelope on the sliders, so the gap between a published severity and real variation is something you move rather than something you read; dimensions 2 and 6 are where these numbers enter the scorecard; the passport is where an envelope and a generator seed would be recorded; live evaluation is the separation rule step 5 borrows.