Order 5% First: How a Data Collection Pilot Works

Order 5% First: How a Data Collection Pilot Works
A spec that read fine on paper, the batch that proved it wasn’t, and the arithmetic behind ordering five percent first

The brief said “record two-party calls in a noisy shop.” Ten contributors followed it, and the pilot arrived with shopping-centre music, till beeps and a PA announcement about detergent. The model was failing on compressor hum: the low drone of a walk-in fridge, thirty decibels quieter. The brief had left “noisy” open to interpretation. The pilot exposed the gap.

Perit delivers five percent of a collection order, or ten to thirty episodes for embodied data, before the main run. That first batch tests whether the written spec produces the evidence the model needs.

Testing the written brief

A collection or annotation brief describes who to record, under which conditions, doing what, and how to label the result. Those instructions must be specific enough for a contributor to follow without asking the person who wrote them what “noisy” means.

The writer usually starts with a failed clip, an example of poor lighting or an accent the model misses. Translating that example into capture instructions can lose detail. “Compressor hum” becomes “mall ambience”, and the mismatch may stay hidden until delivery.

The pilot runs a small slice through recruiting, capture or annotation, review and delivery. The team receives that slice and its quality report while changes to the rest of the order are still cheap.

The arithmetic of catching it early

Perit’s published contributor rate for approved two-party speech is $20 per audio hour. Take a mid-sized order: 2,000 hours of accent-matched call recordings for a new locale. Contributor payouts alone come to:

2,000 hours × $20/hour = $40,000 contributor pay only; review, QA and delivery overhead are quoted separately

Suppose the team misses the noise-condition gap until all 2,000 hours arrive. It can accept data that does not address the failure, or revise the spec and repeat much of the order. A full repeat means another $40,000 in contributor pay and the time needed to record 2,000 hours again.

A 5% pilot of the same order is 100 hours. The contributor cost of finding the gap is:

100 hours × $20/hour = $2,000 the cost of finding out the spec was wrong, instead of the cost of the spec being wrong at full volume

That is $2,000 to discover the flaw before committing a $40,000 batch that might need substantial rework.

When the pilot pays for itself

A pilot has a cost even when the spec is correct. Its expected value depends on how often it catches a gap and how much rework that gap would cause.

Suppose, illustratively, that a first-time spec for a new locale, device or condition has roughly a one-in-three chance of containing a definition gap serious enough that a meaningful share of the batch would need to be redone if it went undiscovered until full delivery. Compare the expected cost of the two strategies:

E[no pilot] = p × (full redo cost) = 0.33 × $40,000 ≈ $13,300 expected cost of skipping the pilot across both flawed and clean specs
E[with pilot] = (pilot cost) + p × (cost of respecifying before the other 95% runs) the pilot cost is paid on every project; the correction cost only applies to the flawed fraction, and it is much smaller once caught early

The pilot costs $2,000 in either case. If it catches a gap, the team can revise the schema and rerun a small slice before committing the other 1,900 hours. Under the illustrative one-in-three failure assumption, this costs less on average than risking a second $40,000 order. The calculation includes pilot costs on the two-thirds of projects whose specs were already sound.

A line chart with the probability of a spec flaw on the horizontal axis from 0 to 100 percent, and expected cost in dollars on the vertical axis. The no-pilot line rises steeply and linearly from zero. The with-pilot line starts near two thousand dollars and rises much more slowly, crossing below the no-pilot line almost immediately and staying well under it across the whole range. Expected cost with and without a 5% pilot Illustrative curves for the 2,000-hour example at $20 per audio hour With 5% pilotNo pilot $0 $10k $20k $30k $40k 0% 25% 50% 75% 100% Chance the first-draft spec has a costly gap Expected cost (USD) No pilot With 5% pilot Pilot floor ≈ $2,000paid on every order
Figure 1. Illustrative costs for the 2,000-hour order. The pilot costs something even when the spec is correct. When a flaw is caught early, the remedy is a revised spec and a small rerun rather than another full order.

Collection pilots and annotation samples

Perit uses different checks for collection and annotation. A collection pilot is 5% of ordered volume. An annotation sample has a fixed size: ten minutes of audio, 100 images or one page of text, returned within five working days with an agreement report.

The collection pilot checks whether the recruiting and capture instructions produce the right inputs across the required conditions. The annotation sample checks whether the team can apply your guide to files you already recognise. An order that needs both stages uses both checks: first a 5% collection pilot, then annotation sampling on that slice before scaling.

Choosing the pilot size

The size of the pilot is itself a decision worth defending, not a round number chosen for convenience.

On the 2,000-hour order, a 1% pilot is 20 hours. That may involve only a few contributors and one or two recording days. It can miss a phone model, dialect or room type that becomes common in the full run, leaving a gap undiscovered until hour 400.

A 20% pilot is 400 hours and costs $8,000 in contributor pay. It may catch the same flaw, but costs four times as much as the 5% pilot. Once the sample covers the conditions the team needs to inspect, adding volume has diminishing value.

Five percent of a multi-thousand-hour order still includes a hundred-plus hours, several contributors and multiple recording days. For embodied data, the pilot uses ten to thirty episodes across several operators and task attempts. It runs in the second week of a four-phase process that starts with a one-page brief.

Step 01Day 0: BriefTask, environment, conditions, device or rig, and volume agreed in one page.
Step 0248 hours: Scope and quoteQuantity, locales, conditions and schema translated into a concrete plan.
Step 03Two weeks: Pilot5% of volume, or 10–30 episodes, delivered with a quality report.
Step 04Weekly: Full deliveryProtocol refined against pilot findings, then a steady delivery cadence.

An egocentric-video example: the brief asks for chest-mounted footage of kitchen cleanup. The tasks and rooms are correct, but half the operators mount the camera high enough that their hands leave the frame when reaching into overhead cupboards. Across thirty pilot episodes, the team can fix the SOP with a rig-height requirement and a reference frame for camera angle. Across three thousand episodes, the same omission means substantial re-shooting.

Testing recruiting, capture and review

It’s tempting to think of the pilot as checking whether contributors can follow instructions. It’s actually checking something broader: whether the brief, the recruiting quota, the recording protocol, the review process and the delivery format all fit together the way the plan assumed they would.

The pilot can reveal that an accent is harder to recruit than expected, or that review takes longer than the delivery schedule allows. A written spec cannot establish those facts. Running the workflow can.

Speech, video, image, text and the eight embodied-data types all use the same four-phase process. The capture instructions vary; the pilot still checks them before the main commitment.

A pilot is not a smaller version of the order. It’s a test of whether the order, as written, actually means what you think it means.

What a pilot report should actually tell you

Ask for the same accuracy, agreement and rework measures that will accompany full delivery. A small batch without those results tells you little about whether the pipeline is ready to scale.

Check which episodes or hours passed QC, why any failed, and whether the cause was contributor error or an ambiguous spec. Would the failure recur if the brief stayed unchanged? A clean pilot is useful too: it shows the instructions worked on the sample tested.

Passing the pilot does not end QA. The full run may introduce new contributors, recruiting pressure and edge cases that the pilot did not contain. Weekly quality reports check whether the approved spec continues to hold as volume grows.

The short version A brief is a guess about what real people, in real conditions, will actually produce. The pilot is the cheapest way to test that guess before the full order depends on it being right. In this example, spending $2,000 on the pilot can avoid a $40,000 repeat.

Start with the failure, not the full order

Every Perit collection and annotation programme runs a pilot before scaling: 5% of volume for speech, video, image and text, or 10 to 30 episodes for embodied data. Send the failing clip, photo or task and get a scope and quote within 48 hours, with the pilot landing two weeks later.

The opening noise-condition example and the egocentric-camera-height case are illustrative composites built to walk through the pilot mechanism, not transcripts of a specific client engagement. The one-in-three flaw-probability figure and the cost curves are illustrative assumptions used to demonstrate the arithmetic, not a published internal statistic. Contributor pay figures are Perit’s published rates at the time of writing; total delivered cost, which includes review, QA and delivery overhead, is quoted per engagement.

Was this article helpful? No ratings yet