A model scores 94% on a benchmark in January. After a retrain, it scores 97% on the same benchmark in April. The release notes call that progress. But did a paraphrase of the test questions enter the training corpus between those runs? Without checking, we cannot tell how much of the improvement came from learning and how much came from leakage.
A held-out test measures generalisation only while it stays outside training. Most teams know that. What is easier to miss is how little leakage can move a leaderboard score.
A test only tests what it hasn’t seen
The logic of a held-out benchmark is the same logic behind a train/test split in any machine learning workflow: you can only claim a model generalises if you check it against examples it never encountered while learning. Score a model on the training data itself and a model that has simply memorised its inputs will look indistinguishable from one that has actually learned the underlying task. The split exists specifically to make that distinction visible.
A public benchmark gives different models the same questions and scoring rules. That comparison depends on every model being tested on material it has not already encountered during training.
How little leakage it takes to move the number
Contamination is not an all-or-nothing failure. Even a small leak can inflate a score enough to change a ranking.
Take a 500-question benchmark. A model scores 70% on questions it has never seen. Suppose 15% of the benchmark has entered its pretraining corpus in a recognisable form, and it scores 98% on those leaked items, partly through recall.
In this example, leaking 15% of the questions raises the reported score from 70% to 74.2%. The 4.2-point gain comes entirely from contamination. A routine audit could miss the leak while a leaderboard reports it as improved capability.
The score rises further as more questions leak. An audit that searches only for exact copies will miss paraphrases, translations and near-duplicates that a model can recognise.
Why the leak is rarely deliberate
Deliberately feeding a model the answer key is one route to contamination. Accidental reuse is harder to prevent with a policy statement alone.
A test question can appear in a blog post, forum thread or paper appendix, then enter a web-scale training corpus. A vendor using the same contributor pool for training and evaluation may collect closely related material for both. A lab that builds its own models and benchmarks also needs someone independent to check that the two pools stayed separate.
There’s a fourth pathway that’s easy to miss because it looks like ordinary iteration rather than a leak at all: a model provider runs their own model against a public benchmark to check progress, logs the transcript for debugging, and that debugging log later gets swept into a future training run along with everything else the team produced that quarter. Nobody copied the benchmark on purpose. The benchmark simply passed through a system that eventually feeds a training corpus, the same way any other document would.
Contamination is not the same problem as realistic overlap
A benchmark should resemble the conditions where the model will be used. Noisy retail audio belongs in a test of noisy retail speech. Leakage occurs when a specific test item enters training; overlap in the kinds of shops, speakers or microphones does not by itself invalidate the test.
The question is whether the model could have encountered this particular item. Training on a thousand other noisy-shop clips can help it generalise to an unseen clip. Training on a near-duplicate of that clip can let it recall the answer. Distinguishing the two requires records of where both the training and benchmark material came from.
Why benchmark ownership matters
A written policy that says “we don’t contaminate our own benchmarks” is a promise. Structural independence is a different kind of guarantee, because it removes the channel through which contamination would happen even if someone wanted it to, or simply made a mistake.
Separating benchmark work from training-data delivery removes a direct route to leakage. If the benchmark custodian does not train frontier models, its held-out set cannot enter its own pretraining run. If it also keeps that material out of deliveries to model providers, the set cannot become training data through an ordinary delivery.
That separation still needs auditing. Web scraping, paraphrases and overlapping contributor pools can introduce leakage elsewhere.
Keeping test data out of training
Independence of ownership is the starting condition, not the whole mechanism. Underneath it, a held-out set stays held out because of a handful of concrete operational choices, each closing one specific leak pathway.
| Mechanism | Leak pathway it closes |
|---|---|
| Separate contributor pools for benchmark and delivery work | The same clip, photo or dialogue ending up in both a training order and an eval set |
| Region-pinned, access-logged storage; clips streamed rather than downloaded | Bulk copies of held-out material leaving the environment where they can be tracked |
| Reviewer chain recorded per batch | An untraceable point where held-out material changed hands or purpose |
| No training of frontier models by the entity holding the benchmark | The benchmark’s own custodian accidentally or deliberately training on it |
| Periodic refresh of held-out material | Slow diffusion of older benchmark items into the public web over time |
Access logs and separate storage are familiar controls. For benchmark work, they need to cover the held-out pool specifically and be checked separately from the controls used for ordinary training-data delivery.
Repeated discussion can expose a benchmark even when its original storage and delivery were handled carefully. Consider an illustrative refresh schedule: if items are replaced annually, and incidental web exposure takes eighteen months to become a concern, most items retire first. Keep the same set for three or four years and more of it may have appeared in material a model could train on. These timings are assumptions for the example, not a measured safe interval.
Reading a benchmark claim like a skeptic
None of this requires a buyer or reader to audit a vendor’s infrastructure directly. It does mean a benchmark number is worth a short list of questions before it gets cited as evidence of anything.
- Q1Who owns the held-out set, and does that entity also train the models being scored against it?
- Q2Has the benchmark material ever been provided, in whole or in part, to a model provider, including the one whose model is being scored?
- Q3Is the benchmark pool drawn from a separate contributor group than the one producing ordinary training-data deliveries, or could the same clip plausibly appear in both?
- Q4When was the benchmark last refreshed, and is there a policy for retiring items that have been publicly discussed or reproduced?
Those four answers give a score some context. Without them, a precise-looking result can combine unseen-task performance with recall of leaked material, and the reader cannot tell how much each contributed.
Ask before you cite the number
Perit’s benchmark audio and evaluation data are held out from the same pool used for training-data delivery, and are never provided to a model provider. See how the quality and calibration system works, or check the FAQ for how benchmark material is separated from ordinary delivery.