{"id":877,"date":"2026-09-29T14:03:38","date_gmt":"2026-09-29T14:03:38","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=877"},"modified":"2026-09-30T18:42:45","modified_gmt":"2026-09-30T18:42:45","slug":"pilot-economics","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/pilot-economics\/","title":{"rendered":"Order 5% First: How a Data Collection Pilot Works"},"content":{"rendered":"\n\r\n<div class=\"pp-post\" style=\"max-width: 50rem;margin: 0 auto;font-family: Verdana, Geneva, Tahoma, sans-serif;font-size: 1rem;line-height: 1.68;color: #000\">\r\n<div class=\"meta\" style=\"font-size: 0.92rem;color: #333;margin-bottom: 1.6rem;padding-bottom: 0.65rem;border-bottom: 1px solid #aaa\">A spec that read fine on paper, the batch that proved it wasn&#8217;t, and the arithmetic behind ordering five percent first<\/div>\r\n<p class=\"lede\" style=\"margin: 0 0 1rem;font-size: 1.08rem\">The brief said \u201crecord two-party calls in a noisy shop.\u201d Ten contributors followed it, and the pilot arrived with shopping-centre music, till beeps and a PA announcement about detergent. The model was failing on compressor hum: the low drone of a walk-in fridge, thirty decibels quieter. The brief had left \u201cnoisy\u201d open to interpretation. The pilot exposed the gap.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Perit delivers five percent of a collection order, or ten to thirty episodes for embodied data, before the main run. That first batch tests whether the written spec produces the evidence the model needs.<\/p>\r\n\r\n<h2 id=\"a-brief-is-a-hypothesis-until-someone-executes-it\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">Testing the written brief<\/h2>\r\n<p style=\"margin: 0 0 1rem\">A collection or annotation brief describes who to record, under which conditions, doing what, and how to label the result. Those instructions must be specific enough for a contributor to follow without asking the person who wrote them what \u201cnoisy\u201d means.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The writer usually starts with a failed clip, an example of poor lighting or an accent the model misses. Translating that example into capture instructions can lose detail. \u201cCompressor hum\u201d becomes \u201cmall ambience\u201d, and the mismatch may stay hidden until delivery.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The pilot runs a small slice through recruiting, capture or annotation, review and delivery. The team receives that slice and its quality report while changes to the rest of the order are still cheap.<\/p>\r\n\r\n<h2 id=\"the-arithmetic-of-catching-it-early\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">The arithmetic of catching it early<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Perit&#8217;s published contributor rate for approved two-party speech is <strong>$20 per audio hour<\/strong>. Take a mid-sized order: 2,000 hours of accent-matched call recordings for a new locale. Contributor payouts alone come to:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &#039;Times New Roman&#039;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">2,000 hours \u00d7 $20\/hour = $40,000<\/span> <span class=\"eq-label\" style=\"display: block;text-align: center;font-size: 0.85rem;color: #333;margin-top: 0.3rem\">contributor pay only; review, QA and delivery overhead are quoted separately<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">Suppose the team misses the noise-condition gap until all 2,000 hours arrive. It can accept data that does not address the failure, or revise the spec and repeat much of the order. A full repeat means another $40,000 in contributor pay and the time needed to record 2,000 hours again.<\/p>\r\n<p style=\"margin: 0 0 1rem\">A 5% pilot of the same order is 100 hours. The contributor cost of finding the gap is:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &#039;Times New Roman&#039;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">100 hours \u00d7 $20\/hour = $2,000<\/span> <span class=\"eq-label\" style=\"display: block;text-align: center;font-size: 0.85rem;color: #333;margin-top: 0.3rem\">the cost of finding out the spec was wrong, instead of the cost of the spec being wrong at full volume<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">That is $2,000 to discover the flaw before committing a $40,000 batch that might need substantial rework.<\/p>\r\n\r\n<h2 id=\"why-this-is-worth-doing-even-when-most-specs-are-fine\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">When the pilot pays for itself<\/h2>\r\n<p style=\"margin: 0 0 1rem\">A pilot has a cost even when the spec is correct. Its expected value depends on how often it catches a gap and how much rework that gap would cause.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Suppose, illustratively, that a first-time spec for a new locale, device or condition has roughly a one-in-three chance of containing a definition gap serious enough that a meaningful share of the batch would need to be redone if it went undiscovered until full delivery. Compare the expected cost of the two strategies:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &#039;Times New Roman&#039;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">E[no pilot] = p \u00d7 (full redo cost) = 0.33 \u00d7 $40,000 \u2248 $13,300<\/span> <span class=\"eq-label\" style=\"display: block;text-align: center;font-size: 0.85rem;color: #333;margin-top: 0.3rem\">expected cost of skipping the pilot across both flawed and clean specs<\/span><\/div>\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;border: 1px solid #999;border-left: 4px solid #000\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &#039;Times New Roman&#039;, Georgia, serif;font-size: 1.12rem;margin: 0.4rem 0\">E[with pilot] = (pilot cost) + p \u00d7 (cost of respecifying before the other 95% runs)<\/span> <span class=\"eq-label\" style=\"display: block;text-align: center;font-size: 0.85rem;color: #333;margin-top: 0.3rem\">the pilot cost is paid on every project; the correction cost only applies to the flawed fraction, and it is much smaller once caught early<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">The pilot costs $2,000 in either case. If it catches a gap, the team can revise the schema and rerun a small slice before committing the other 1,900 hours. Under the illustrative one-in-three failure assumption, this costs less on average than risking a second $40,000 order. The calculation includes pilot costs on the two-thirds of projects whose specs were already sound.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">A line chart with the probability of a spec flaw on the horizontal axis from 0 to 100 percent, and expected cost in dollars on the vertical axis. The no-pilot line rises steeply and linearly from zero. The with-pilot line starts near two thousand dollars and rises much more slowly, crossing below the no-pilot line almost immediately and staying well under it across the whole range. Expected cost with and without a 5% pilot Illustrative curves for the 2,000-hour example at $20 per audio hour With 5% pilotNo pilot $0 $10k $20k $30k $40k 0% 25% 50% 75% 100% Chance the first-draft spec has a costly gap Expected cost (USD) No pilot With 5% pilot Pilot floor \u2248 $2,000paid on every order<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: 0.65rem;font-size: 0.9rem;line-height: 1.55;color: #52514e\">Figure 1. Illustrative costs for the 2,000-hour order. The pilot costs something even when the spec is correct. When a flaw is caught early, the remedy is a revised spec and a small rerun rather than another full order.<\/figcaption><\/figure>\r\n<h2 id=\"the-same-logic-a-different-denominator\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">Collection pilots and annotation samples<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Perit uses different checks for collection and annotation. A collection pilot is 5% of ordered volume. An annotation sample has a fixed size: ten minutes of audio, 100 images or one page of text, returned within five working days with an agreement report.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The collection pilot checks whether the recruiting and capture instructions produce the right inputs across the required conditions. The annotation sample checks whether the team can apply your guide to files you already recognise. An order that needs both stages uses both checks: first a 5% collection pilot, then annotation sampling on that slice before scaling.<\/p>\r\n\r\n<h2 id=\"why-five-percent-and-not-one-or-twenty\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">Choosing the pilot size<\/h2>\r\n<p style=\"margin: 0 0 1rem\">The size of the pilot is itself a decision worth defending, not a round number chosen for convenience.<\/p>\r\n<p style=\"margin: 0 0 1rem\">On the 2,000-hour order, a 1% pilot is 20 hours. That may involve only a few contributors and one or two recording days. It can miss a phone model, dialect or room type that becomes common in the full run, leaving a gap undiscovered until hour 400.<\/p>\r\n<p style=\"margin: 0 0 1rem\">A 20% pilot is 400 hours and costs $8,000 in contributor pay. It may catch the same flaw, but costs four times as much as the 5% pilot. Once the sample covers the conditions the team needs to inspect, adding volume has diminishing value.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Five percent of a multi-thousand-hour order still includes a hundred-plus hours, several contributors and multiple recording days. For embodied data, the pilot uses ten to thirty episodes across several operators and task attempts. It runs in the second week of a four-phase process that starts with a one-page brief.<\/p>\r\n\r\n<div class=\"gate\" style=\"display: flex;flex-wrap: wrap;gap: 0.75rem;margin: 2rem 0\" aria-label=\"The four-phase pilot process\">\r\n<div class=\"gate-step\" style=\"flex: 1 1 11rem;min-width: 11rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin-bottom: 0.4rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;text-transform: uppercase\">Step 01<\/span><strong style=\"display: block;margin-bottom: 0.35rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Day 0: Brief<\/strong><span class=\"step-desc\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">Task, environment, conditions, device or rig, and volume agreed in one page.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 11rem;min-width: 11rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin-bottom: 0.4rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;text-transform: uppercase\">Step 02<\/span><strong style=\"display: block;margin-bottom: 0.35rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">48 hours: Scope and quote<\/strong><span class=\"step-desc\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">Quantity, locales, conditions and schema translated into a concrete plan.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 11rem;min-width: 11rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin-bottom: 0.4rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;text-transform: uppercase\">Step 03<\/span><strong style=\"display: block;margin-bottom: 0.35rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Two weeks: Pilot<\/strong><span class=\"step-desc\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">5% of volume, or 10\u201330 episodes, delivered with a quality report.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 11rem;min-width: 11rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin-bottom: 0.4rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;text-transform: uppercase\">Step 04<\/span><strong style=\"display: block;margin-bottom: 0.35rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Weekly: Full delivery<\/strong><span class=\"step-desc\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">Protocol refined against pilot findings, then a steady delivery cadence.<\/span><\/div>\r\n<\/div>\r\n<div class=\"case\" style=\"margin: 1.45rem 0;padding: 0.9rem 1rem;border-left: 4px solid #1f5fb8\">\r\n<p style=\"margin: 0 0 1rem;margin-bottom: 0\">An egocentric-video example: the brief asks for chest-mounted footage of kitchen cleanup. The tasks and rooms are correct, but half the operators mount the camera high enough that their hands leave the frame when reaching into overhead cupboards. Across thirty pilot episodes, the team can fix the SOP with a rig-height requirement and a reference frame for camera angle. Across three thousand episodes, the same omission means substantial re-shooting.<\/p>\r\n\r\n<\/div>\r\n<h2 id=\"the-pilot-is-a-test-of-the-whole-pipeline-not-just-the-contributors\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">Testing recruiting, capture and review<\/h2>\r\n<p style=\"margin: 0 0 1rem\">It&#8217;s tempting to think of the pilot as checking whether contributors can follow instructions. It&#8217;s actually checking something broader: whether the brief, the recruiting quota, the recording protocol, the review process and the delivery format all fit together the way the plan assumed they would.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The pilot can reveal that an accent is harder to recruit than expected, or that review takes longer than the delivery schedule allows. A written spec cannot establish those facts. Running the workflow can.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Speech, video, image, text and the eight embodied-data types all use the same four-phase process. The capture instructions vary; the pilot still checks them before the main commitment.<\/p>\r\n\r\n<div class=\"quote-block\" style=\"margin: 1.6rem 0;padding: 0.15rem 0 0.15rem 1rem;border-left: 3px solid #000;font-size: 1.05rem;font-style: italic\">A pilot is not a smaller version of the order. It&#8217;s a test of whether the order, as written, actually means what you think it means.<\/div>\r\n<h2 id=\"what-a-pilot-report-should-actually-tell-you\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold\">What a pilot report should actually tell you<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Ask for the same accuracy, agreement and rework measures that will accompany full delivery. A small batch without those results tells you little about whether the pipeline is ready to scale.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Check which episodes or hours passed QC, why any failed, and whether the cause was contributor error or an ambiguous spec. Would the failure recur if the brief stayed unchanged? A clean pilot is useful too: it shows the instructions worked on the sample tested.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Passing the pilot does not end QA. The full run may introduce new contributors, recruiting pressure and edge cases that the pilot did not contain. Weekly quality reports check whether the approved spec continues to hold as volume grows.<\/p>\r\n\r\n<div class=\"note\" style=\"margin: 1.5rem 0;padding: 0.85rem 1rem;border: 1px solid #999;font-size: 0.92rem\"><strong style=\"display: block;margin-bottom: 0.25rem\">The short version<\/strong> A brief is a guess about what real people, in real conditions, will actually produce. The pilot is the cheapest way to test that guess before the full order depends on it being right. In this example, spending $2,000 on the pilot can avoid a $40,000 repeat.<\/div>\r\n<div class=\"cta\" style=\"margin: 2.2rem 0 0;padding: 1.15rem 1.2rem;border: 2px solid #000\">\r\n<h2 id=\"start-with-the-failure-not-the-full-order\" style=\"font-size: 1.28rem;line-height: 1.35;font-weight: bold;margin: 0 0 0.65rem\">Start with the failure, not the full order<\/h2>\r\n<p style=\"margin: 0 0 1rem;margin-bottom: 0\">Every Perit collection and annotation programme runs a pilot before scaling: 5% of volume for speech, video, image and text, or 10 to 30 episodes for embodied data. <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/request-data\" target=\"_blank\" rel=\"noopener\">Send the failing clip, photo or task<\/a> and get a scope and quote within 48 hours, with the pilot landing two weeks later.<\/p>\r\n\r\n<\/div>\r\n<footer class=\"post-note\" style=\"margin-top: 3rem;padding-top: 1rem;border-top: 1px solid #aaa;font-size: 0.88rem;line-height: 1.5;color: #222\" aria-label=\"Notes on the examples and figures\">\r\n<p style=\"margin: 0 0 1rem\">The opening noise-condition example and the egocentric-camera-height case are illustrative composites built to walk through the pilot mechanism, not transcripts of a specific client engagement. The one-in-three flaw-probability figure and the cost curves are illustrative assumptions used to demonstrate the arithmetic, not a published internal statistic. Contributor pay figures are Perit&#8217;s published rates at the time of writing; total delivered cost, which includes review, QA and delivery overhead, is quoted per engagement.<\/p>\r\n\r\n<\/footer><\/div>\r\n\n","protected":false},"excerpt":{"rendered":"<p>A 5% pilot tests whether a data brief works before the full order runs. The cost arithmetic, what to inspect, and why QA continues after the pilot.<\/p>\n","protected":false},"author":2,"featured_media":878,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2,1],"tags":[],"class_list":["post-877","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-featured","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/877","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=877"}],"version-history":[{"count":3,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/877\/revisions"}],"predecessor-version":[{"id":929,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/877\/revisions\/929"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/878"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=877"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=877"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=877"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}