{"id":919,"date":"2026-09-30T18:11:09","date_gmt":"2026-09-30T18:11:09","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=919"},"modified":"2026-09-30T18:42:33","modified_gmt":"2026-09-30T18:42:33","slug":"the-benchmark-nobody-sees","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/the-benchmark-nobody-sees\/","title":{"rendered":"The Benchmark Nobody Sees"},"content":{"rendered":"\n\n<div class=\"pp-post\" style=\"max-width: 50rem;margin: 0 auto;font-family: Verdana, Geneva, Tahoma, sans-serif;font-size: 1rem;line-height: 1.68;color: #000\">\n<div class=\"meta\" style=\"font-size: 0.92rem;color: #333;margin-bottom: 1.6rem;padding-bottom: 0.65rem;border-bottom: 1px solid #aaa\">A benchmark score that quietly climbed, a leak nobody put there on purpose, and the arithmetic of how little contamination it takes to notice<\/div>\n<p class=\"lede\" style=\"margin: 0 0 1rem;font-size: 1.08rem\">A model scores 94% on a benchmark in January. After a retrain, it scores 97% on the same benchmark in April. The release notes call that progress. But did a paraphrase of the test questions enter the training corpus between those runs? Without checking, we cannot tell how much of the improvement came from learning and how much came from leakage.<\/p>\n<p style=\"margin: 0 0 1rem\">A held-out test measures generalisation only while it stays outside training. Most teams know that. What is easier to miss is how little leakage can move a leaderboard score.<\/p>\n<h2 id=\"a-test-only-tests-what-it-hasnt-seen\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">A test only tests what it hasn&#8217;t seen<\/h2>\n<p style=\"margin: 0 0 1rem\">The logic of a held-out benchmark is the same logic behind a train\/test split in any machine learning workflow: you can only claim a model generalises if you check it against examples it never encountered while learning. Score a model on the training data itself and a model that has simply memorised its inputs will look indistinguishable from one that has actually learned the underlying task. The split exists specifically to make that distinction visible.<\/p>\n<p style=\"margin: 0 0 1rem\">A public benchmark gives different models the same questions and scoring rules. That comparison depends on every model being tested on material it has not already encountered during training.<\/p>\n<h2 id=\"how-little-leakage-it-takes-to-move-the-number\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">How little leakage it takes to move the number<\/h2>\n<p style=\"margin: 0 0 1rem\">Contamination is not an all-or-nothing failure. Even a small leak can inflate a score enough to change a ranking.<\/p>\n<p style=\"margin: 0 0 1rem\">Take a 500-question benchmark. A model scores 70% on questions it has never seen. Suppose 15% of the benchmark has entered its pretraining corpus in a recognisable form, and it scores 98% on those leaked items, partly through recall.<\/p>\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"> <span class=\"eq\" style=\"display: block;text-align: center;font-family: &quot;Times New Roman&quot;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">observed score = (1 \u2212 c) \u00d7 a_clean + c \u00d7 a_leaked<\/span> <span class=\"eq-label\" style=\"display: block;text-align: center;font-size: 0.85rem;color: #333;margin-top: 0.3rem\">c = contaminated fraction, a_clean = accuracy on genuinely unseen items, a_leaked = accuracy on leaked items<\/span> <\/div>\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"> <span class=\"eq\" style=\"display: block;text-align: center;font-family: &quot;Times New Roman&quot;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">observed score = 0.85 \u00d7 70% + 0.15 \u00d7 98% = 59.5% + 14.7% = 74.2%<\/span> <\/div>\n<p style=\"margin: 0 0 1rem\">In this example, leaking 15% of the questions raises the reported score from 70% to 74.2%. The 4.2-point gain comes entirely from contamination. A routine audit could miss the leak while a leaderboard reports it as improved capability.<\/p>\n<figure class=\"viz\" style=\"margin: 2.2rem 0\"> <div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\"> <img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/perit.ai\/blogs\/wp-content\/uploads\/2026\/09\/the-benchmark-nobody-sees-test-leakage.webp\" alt=\"Illustrative chart: 70% clean accuracy rises to 74.2% when 15% of test items leak into training, assuming 98% accuracy on leaked items.\" width=\"1520\" height=\"688\" style=\"display:block;width:100%;min-width:560px;height:auto\"> <\/div> <figcaption class=\"viz-caption\" style=\"margin-top: 0.65rem;font-size: 0.9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 1.<\/strong> A linear blend of clean and leaked accuracy, using the illustrative figures above. Even modest contamination produces a headline number well above the model&#8217;s genuine unseen-data performance.<\/figcaption> <\/figure>\n<p style=\"margin: 0 0 1rem\">The score rises further as more questions leak. An audit that searches only for exact copies will miss paraphrases, translations and near-duplicates that a model can recognise.<\/p>\n<h2 id=\"why-the-leak-is-rarely-deliberate\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">Why the leak is rarely deliberate<\/h2>\n<p style=\"margin: 0 0 1rem\">Deliberately feeding a model the answer key is one route to contamination. Accidental reuse is harder to prevent with a policy statement alone.<\/p>\n<p style=\"margin: 0 0 1rem\">A test question can appear in a blog post, forum thread or paper appendix, then enter a web-scale training corpus. A vendor using the same contributor pool for training and evaluation may collect closely related material for both. A lab that builds its own models and benchmarks also needs someone independent to check that the two pools stayed separate.<\/p>\n<p style=\"margin: 0 0 1rem\">There&#8217;s a fourth pathway that&#8217;s easy to miss because it looks like ordinary iteration rather than a leak at all: a model provider runs their own model against a public benchmark to check progress, logs the transcript for debugging, and that debugging log later gets swept into a future training run along with everything else the team produced that quarter. Nobody copied the benchmark on purpose. The benchmark simply passed through a system that eventually feeds a training corpus, the same way any other document would.<\/p>\n<h2 id=\"contamination-is-not-the-same-problem-as-realistic-overlap\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">Contamination is not the same problem as realistic overlap<\/h2>\n<p style=\"margin: 0 0 1rem\">A benchmark should resemble the conditions where the model will be used. Noisy retail audio belongs in a test of noisy retail speech. Leakage occurs when a specific test item enters training; overlap in the kinds of shops, speakers or microphones does not by itself invalidate the test.<\/p>\n<p style=\"margin: 0 0 1rem\">The question is whether the model could have encountered this particular item. Training on a thousand other noisy-shop clips can help it generalise to an unseen clip. Training on a near-duplicate of that clip can let it recall the answer. Distinguishing the two requires records of where both the training and benchmark material came from.<\/p>\n<div class=\"quote-block\" style=\"margin: 1.6rem 0;padding: 0.15rem 0 0.15rem 1rem;border-left: 3px solid #000;font-size: 1.05rem;font-style: italic\">Data can leak through ordinary collection, logging and delivery workflows. Preventing it requires checks in those workflows.<\/div>\n<h2 id=\"why-who-owns-the-benchmark-changes-the-incentive-not-just-the-paperwor\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">Why benchmark ownership matters<\/h2>\n<p style=\"margin: 0 0 1rem\">A written policy that says &#8220;we don&#8217;t contaminate our own benchmarks&#8221; is a promise. Structural independence is a different kind of guarantee, because it removes the channel through which contamination would happen even if someone wanted it to, or simply made a mistake.<\/p>\n<p style=\"margin: 0 0 1rem\">Separating benchmark work from training-data delivery removes a direct route to leakage. If the benchmark custodian does not train frontier models, its held-out set cannot enter its own pretraining run. If it also keeps that material out of deliveries to model providers, the set cannot become training data through an ordinary delivery.<\/p>\n<p style=\"margin: 0 0 1rem\">That separation still needs auditing. Web scraping, paraphrases and overlapping contributor pools can introduce leakage elsewhere.<\/p>\n<h2 id=\"what-actually-keeps-a-held-out-set-held-out\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">Keeping test data out of training<\/h2>\n<p style=\"margin: 0 0 1rem\">Independence of ownership is the starting condition, not the whole mechanism. Underneath it, a held-out set stays held out because of a handful of concrete operational choices, each closing one specific leak pathway.<\/p>\n<div class=\"table-wrap\" style=\"max-width: 100%;overflow: auto;margin: 1.5rem 0\"> <table style=\"width: 100%;min-width: 640px;border-collapse: collapse;font-size: 0.9rem;line-height: 1.5\"> <thead><tr><th style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: 700;color: #0b0b0b\">Mechanism<\/th><th style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: 700;color: #0b0b0b\">Leak pathway it closes<\/th><\/tr><\/thead> <tbody> <tr><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Separate contributor pools for benchmark and delivery work<\/td><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">The same clip, photo or dialogue ending up in both a training order and an eval set<\/td><\/tr> <tr><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Region-pinned, access-logged storage; clips streamed rather than downloaded<\/td><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Bulk copies of held-out material leaving the environment where they can be tracked<\/td><\/tr> <tr><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Reviewer chain recorded per batch<\/td><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">An untraceable point where held-out material changed hands or purpose<\/td><\/tr> <tr><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">No training of frontier models by the entity holding the benchmark<\/td><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">The benchmark&#8217;s own custodian accidentally or deliberately training on it<\/td><\/tr> <tr><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Periodic refresh of held-out material<\/td><td style=\"text-align: left;vertical-align: top;padding: 0.6rem;border-bottom: 1px solid #e6e5e0\">Slow diffusion of older benchmark items into the public web over time<\/td><\/tr> <\/tbody> <\/table> <\/div>\n<p style=\"margin: 0 0 1rem\">Access logs and separate storage are familiar controls. For benchmark work, they need to cover the held-out pool specifically and be checked separately from the controls used for ordinary training-data delivery.<\/p>\n<p style=\"margin: 0 0 1rem\">Repeated discussion can expose a benchmark even when its original storage and delivery were handled carefully. Consider an illustrative refresh schedule: if items are replaced annually, and incidental web exposure takes eighteen months to become a concern, most items retire first. Keep the same set for three or four years and more of it may have appeared in material a model could train on. These timings are assumptions for the example, not a measured safe interval.<\/p>\n<h2 id=\"reading-a-benchmark-claim-like-a-skeptic\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: 700\">Reading a benchmark claim like a skeptic<\/h2>\n<p style=\"margin: 0 0 1rem\">None of this requires a buyer or reader to audit a vendor&#8217;s infrastructure directly. It does mean a benchmark number is worth a short list of questions before it gets cited as evidence of anything.<\/p>\n<ul class=\"checklist\" style=\"margin: 1.6rem 0;padding: 0;list-style-type: none\"> <li style=\"padding-left: 0.15rem;margin: 0 0 0.6rem;display: flex;gap: 0.85rem;align-items: baseline;padding: 0.75rem 0.95rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;font-size: 0.94rem;line-height: 1.55\"><span class=\"check-no\" style=\"flex: 0 0 auto;min-width: 1.6rem;color: #2a78d6;font-size: 0.78rem;font-weight: 700;letter-spacing: 0.08em;text-transform: uppercase\">Q1<\/span><span class=\"check-text\" style=\"flex: 1 1 auto\">Who owns the held-out set, and does that entity also train the models being scored against it?<\/span><\/li> <li style=\"padding-left: 0.15rem;margin: 0 0 0.6rem;display: flex;gap: 0.85rem;align-items: baseline;padding: 0.75rem 0.95rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;font-size: 0.94rem;line-height: 1.55\"><span class=\"check-no\" style=\"flex: 0 0 auto;min-width: 1.6rem;color: #2a78d6;font-size: 0.78rem;font-weight: 700;letter-spacing: 0.08em;text-transform: uppercase\">Q2<\/span><span class=\"check-text\" style=\"flex: 1 1 auto\">Has the benchmark material ever been provided, in whole or in part, to a model provider, including the one whose model is being scored?<\/span><\/li> <li style=\"padding-left: 0.15rem;margin: 0 0 0.6rem;display: flex;gap: 0.85rem;align-items: baseline;padding: 0.75rem 0.95rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;font-size: 0.94rem;line-height: 1.55\"><span class=\"check-no\" style=\"flex: 0 0 auto;min-width: 1.6rem;color: #2a78d6;font-size: 0.78rem;font-weight: 700;letter-spacing: 0.08em;text-transform: uppercase\">Q3<\/span><span class=\"check-text\" style=\"flex: 1 1 auto\">Is the benchmark pool drawn from a separate contributor group than the one producing ordinary training-data deliveries, or could the same clip plausibly appear in both?<\/span><\/li> <li style=\"padding-left: 0.15rem;margin: 0 0 0.6rem;display: flex;gap: 0.85rem;align-items: baseline;padding: 0.75rem 0.95rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;font-size: 0.94rem;line-height: 1.55\"><span class=\"check-no\" style=\"flex: 0 0 auto;min-width: 1.6rem;color: #2a78d6;font-size: 0.78rem;font-weight: 700;letter-spacing: 0.08em;text-transform: uppercase\">Q4<\/span><span class=\"check-text\" style=\"flex: 1 1 auto\">When was the benchmark last refreshed, and is there a policy for retiring items that have been publicly discussed or reproduced?<\/span><\/li> <\/ul>\n<p style=\"margin: 0 0 1rem\">Those four answers give a score some context. Without them, a precise-looking result can combine unseen-task performance with recall of leaked material, and the reader cannot tell how much each contributed.<\/p>\n<div class=\"note\" style=\"margin: 1.5rem 0;padding: 0.85rem 1rem;border: 1px solid #999;font-size: 0.92rem\"> <strong style=\"display: block;margin-bottom: 0.25rem\">The short version<\/strong> A benchmark is useful while its questions stay outside training. Even a small leak can move the score. Separate storage, delivery controls and independent checks reduce that risk.<\/div>\n<div class=\"cta\" style=\"margin: 2.2rem 0 0;padding: 1.15rem 1.2rem;border: 2px solid #000\"> <h2 id=\"ask-before-you-cite-the-number\" style=\"font-size: 1.28rem;line-height: 1.35;font-weight: 700;margin: 0 0 0.65rem\">Ask before you cite the number<\/h2> <p style=\"margin: 0 0 1rem;margin-bottom: 0\">Perit&#8217;s benchmark audio and evaluation data are held out from the same pool used for training-data delivery, and are never provided to a model provider. <a href=\"https:\/\/perit.ai\/services\/data-annotation\" rel=\"noopener\" style=\"color: #0000ee;text-decoration: underline\" target=\"_blank\">See how the quality and calibration system works<\/a>, or <a href=\"https:\/\/perit.ai\/faq\" rel=\"noopener\" style=\"color: #0000ee;text-decoration: underline\" target=\"_blank\">check the FAQ<\/a> for how benchmark material is separated from ordinary delivery.<\/p> <\/div>\n<footer aria-label=\"Notes on the examples and figures\" class=\"post-note\" style=\"margin-top: 3rem;padding-top: 1rem;border-top: 1px solid #aaa;font-size: 0.88rem;line-height: 1.5;color: #222\"> <p style=\"margin: 0 0 1rem\">The January-to-April benchmark example and the 500-question, 15%-leak scenario are illustrative constructions used to demonstrate the contamination arithmetic, not measurements from a specific published benchmark. The linear blend model of observed score is a simplification; real contamination effects can be non-linear depending on how directly a model can recall leaked material. Public discussions of benchmark contamination in large language models, including analyses of grade-school-math-style benchmarks, informed the framing here but are not cited as sources for the specific figures used.<\/p> <\/footer>\n<\/div>\n\n\n","protected":false},"excerpt":{"rendered":"<p>A model goes from 94% to 97%. How test leakage can inflate that gain, how contamination happens by accident, and how to keep a benchmark held out.<\/p>\n","protected":false},"author":2,"featured_media":913,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22,19],"tags":[],"class_list":["post-919","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","category-research"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/919","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=919"}],"version-history":[{"count":2,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/919\/revisions"}],"predecessor-version":[{"id":921,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/919\/revisions\/921"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/913"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=919"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=919"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=919"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}