The brief said “record two-party calls in a noisy shop.” Ten contributors followed it, and the pilot arrived with shopping-centre music, till beeps and a PA announcement about detergent. The model was failing on compressor hum: the low drone of a walk-in fridge, thirty decibels quieter. The brief had left “noisy” open to interpretation. The pilot exposed the gap.
Perit delivers five percent of a collection order, or ten to thirty episodes for embodied data, before the main run. That first batch tests whether the written spec produces the evidence the model needs.
Testing the written brief
A collection or annotation brief describes who to record, under which conditions, doing what, and how to label the result. Those instructions must be specific enough for a contributor to follow without asking the person who wrote them what “noisy” means.
The writer usually starts with a failed clip, an example of poor lighting or an accent the model misses. Translating that example into capture instructions can lose detail. “Compressor hum” becomes “mall ambience”, and the mismatch may stay hidden until delivery.
The pilot runs a small slice through recruiting, capture or annotation, review and delivery. The team receives that slice and its quality report while changes to the rest of the order are still cheap.
The arithmetic of catching it early
Perit’s published contributor rate for approved two-party speech is $20 per audio hour. Take a mid-sized order: 2,000 hours of accent-matched call recordings for a new locale. Contributor payouts alone come to:
Suppose the team misses the noise-condition gap until all 2,000 hours arrive. It can accept data that does not address the failure, or revise the spec and repeat much of the order. A full repeat means another $40,000 in contributor pay and the time needed to record 2,000 hours again.
A 5% pilot of the same order is 100 hours. The contributor cost of finding the gap is:
That is $2,000 to discover the flaw before committing a $40,000 batch that might need substantial rework.
When the pilot pays for itself
A pilot has a cost even when the spec is correct. Its expected value depends on how often it catches a gap and how much rework that gap would cause.
Suppose, illustratively, that a first-time spec for a new locale, device or condition has roughly a one-in-three chance of containing a definition gap serious enough that a meaningful share of the batch would need to be redone if it went undiscovered until full delivery. Compare the expected cost of the two strategies:
The pilot costs $2,000 in either case. If it catches a gap, the team can revise the schema and rerun a small slice before committing the other 1,900 hours. Under the illustrative one-in-three failure assumption, this costs less on average than risking a second $40,000 order. The calculation includes pilot costs on the two-thirds of projects whose specs were already sound.
Collection pilots and annotation samples
Perit uses different checks for collection and annotation. A collection pilot is 5% of ordered volume. An annotation sample has a fixed size: ten minutes of audio, 100 images or one page of text, returned within five working days with an agreement report.
The collection pilot checks whether the recruiting and capture instructions produce the right inputs across the required conditions. The annotation sample checks whether the team can apply your guide to files you already recognise. An order that needs both stages uses both checks: first a 5% collection pilot, then annotation sampling on that slice before scaling.
Choosing the pilot size
The size of the pilot is itself a decision worth defending, not a round number chosen for convenience.
On the 2,000-hour order, a 1% pilot is 20 hours. That may involve only a few contributors and one or two recording days. It can miss a phone model, dialect or room type that becomes common in the full run, leaving a gap undiscovered until hour 400.
A 20% pilot is 400 hours and costs $8,000 in contributor pay. It may catch the same flaw, but costs four times as much as the 5% pilot. Once the sample covers the conditions the team needs to inspect, adding volume has diminishing value.
Five percent of a multi-thousand-hour order still includes a hundred-plus hours, several contributors and multiple recording days. For embodied data, the pilot uses ten to thirty episodes across several operators and task attempts. It runs in the second week of a four-phase process that starts with a one-page brief.
An egocentric-video example: the brief asks for chest-mounted footage of kitchen cleanup. The tasks and rooms are correct, but half the operators mount the camera high enough that their hands leave the frame when reaching into overhead cupboards. Across thirty pilot episodes, the team can fix the SOP with a rig-height requirement and a reference frame for camera angle. Across three thousand episodes, the same omission means substantial re-shooting.
Testing recruiting, capture and review
It’s tempting to think of the pilot as checking whether contributors can follow instructions. It’s actually checking something broader: whether the brief, the recruiting quota, the recording protocol, the review process and the delivery format all fit together the way the plan assumed they would.
The pilot can reveal that an accent is harder to recruit than expected, or that review takes longer than the delivery schedule allows. A written spec cannot establish those facts. Running the workflow can.
Speech, video, image, text and the eight embodied-data types all use the same four-phase process. The capture instructions vary; the pilot still checks them before the main commitment.
What a pilot report should actually tell you
Ask for the same accuracy, agreement and rework measures that will accompany full delivery. A small batch without those results tells you little about whether the pipeline is ready to scale.
Check which episodes or hours passed QC, why any failed, and whether the cause was contributor error or an ambiguous spec. Would the failure recur if the brief stayed unchanged? A clean pilot is useful too: it shows the instructions worked on the sample tested.
Passing the pilot does not end QA. The full run may introduce new contributors, recruiting pressure and edge cases that the pilot did not contain. Weekly quality reports check whether the approved spec continues to hold as volume grows.
Start with the failure, not the full order
Every Perit collection and annotation programme runs a pilot before scaling: 5% of volume for speech, video, image and text, or 10 to 30 episodes for embodied data. Send the failing clip, photo or task and get a scope and quote within 48 hours, with the pilot landing two weeks later.