Collection, Annotation or Both?

Collection, Annotation or Both?
One failed input, three doors, five lines on the annotation bench, and the gate every annotator must clear

A customer gives an amount over a noisy phone line. The speech model records the wrong number, and the rest of the system processes it correctly. In the meeting that follows, someone asks for several hundred thousand rows of “more data”. The failed clip is twelve seconds long. It is worth examining before choosing the order size.

“More data” can mean three different jobs.

The team may need new recordings because the acoustic condition, accent or interaction that breaks the model is absent from its files. It may already own the right recordings, in which case it needs those files transcribed, aligned or otherwise labelled. Or it may know only the failure, the twelve-second clip, and need someone to turn that failure into a collection protocol, an annotation guide and a licensable delivery.

Those routes are related, but they are not interchangeable. Annotation cannot create the warehouse recording you never captured. Collection does not automatically tell you which word was spoken at 7.4 seconds. And “transcription versus alignment” is not the first fork at all: both are jobs inside annotation.

Before choosing a service, check whether the inputs your model needs already exist.

Do you already have the inputs that contain the failure?
No Data Collection

Record or author the missing situations, people, devices, rooms or language.

Yes Data Annotation

Turn files you own into transcripts, timestamps, masks, tracks, spans or judgments.

I only know the failure Request Data

Start with one bad input; scope collection and annotation together.

Check what is missing before choosing a service

Suppose your model misses product names when a shopkeeper speaks beside a refrigerator. “We need annotated audio” sounds like a reasonable brief. It is not yet enough to choose a service.

Open the files you already have. If they contain the same shops, accents, devices and background hum, the missing ingredient may indeed be labels: a verbatim reference transcript, word-level times, speaker turns or entity spans around the product names. If the files are mostly clean office speech, better labels will make the clean office data more precise. They will not introduce the refrigerator, the handset microphone or the shopkeeper.

Start with the failed input. Does that condition appear in your files? What should a person label? Which measurement would show that the new batch helped?

This is why all three Perit routes use nearly the same opening language: send a clip, photograph, paragraph or task the model handles badly. You do not need to solve the data specification before the first conversation. You do need to bring the failure close enough that somebody can inspect it.

A service name is a conclusion. The failed input is the evidence that gets you there.

Collect the missing situations

Choose Data Collection when the model needs experience your current files do not contain. That missing experience might be a locale, an accent, a room, a camera angle, a handset, a lighting condition, a document type or a way people phrase the same intention.

Collection is not restricted to speech. Perit lists four modalities on the same recruiting and QA machinery: speech and audio; video; image; and text and conversation. A contributor might record a two-party call with each speaker on a separate channel, film an everyday task on a chest camera, photograph the same shelf under two lighting conditions, or write a dialogue in the register of the sector they work in. The protocol and equipment change. Recruiting to quotas, capturing consent, running a pilot, reviewing quality and delivering weekly remain the same.

Speech is currently the deepest line: 4,000+ hours delivered, nine live locales and capacity for 50 hours a day. Those figures establish throughput, but the collection spec must still name the people, conditions, devices and outcomes that belong in your batch.

That is the purpose of the pilot. After the brief, Perit returns a scope and quote; then five percent of the planned volume is delivered before the main run. The pilot exists so a vague instruction becomes an obvious problem on a small batch. Perhaps “noisy shop” produces shopping-centre music when the model actually fails on compressor hum. Perhaps “Hindi speakers” is too broad for a regional failure. It is much cheaper to discover that at five percent than at one hundred.

A collection-shaped failure: Your image model performs well on catalogue photographs but misses products on crowded shelves photographed by customers at night. You have labels for clean product images; you do not have the user device, shelf density or light. The missing thing is not another pass over the old boxes. It is new imagery collected across the conditions in which the model is expected to work.

Collection can include labels where ordered. That does not blur the distinction; it means one work order can cross it. The people and situations are collected first, then the resulting files move through an annotation specification. “Collection” describes how the raw evidence comes into existence, not a requirement that it remain raw.

Annotate the inputs you already have

Choose Data Annotation when the useful inputs already sit in your bucket but the answer is not attached to them. This is also the direct answer to a question the FAQ makes explicit: yes, Perit annotates data customers already own; in fact, most annotation work is performed on customer audio, images or video.

The work runs through Foundry. The published bench has 650 people: 250 transcribers, including 50 on fixed-term contracts, and 400 freelance word aligners. It has delivered more than 3,000 transcribed hours and more than 1,000 aligned hours across nine locales. The service page reports 95%+ QA-sample accuracy.

That last number deserves its full name. It is QA-sample accuracy, checked against the project standard; it is not a promise that every possible label on every possible dataset is “95% accurate”. For audio, word error rate is reported separately against an adjudicated reference. Keeping those measures separate is useful because a batch can follow its guide consistently while the guide itself is incomplete, or produce a good transcript while still placing a timing boundary poorly.

Graders stream the files rather than copy them, and the project is pinned to the agreed region. What remains is to decide what answer should be attached to each input.

Choosing the annotation task

This is where transcription and alignment finally become the right question. Audio happens to split into two distinct crafts, while image, video and text each need their own label geometry. A useful way to choose is to finish this sentence: After annotation, I need to know…

Bench line The question it answers What comes back The common mix-up
Transcription What was actually said? Verbatim text, with fillers, repeats and false starts retained under the written convention. Asking for a “clean” transcript, then trying to evaluate a model on disfluencies that were edited away.
Alignment When did each word begin and end? Word-level start and end timestamps placed on top of a transcript. Treating alignment as a substitute for transcription. A timestamp needs a word to belong to.
Image labels Where is the object, and what is true of it? Boxes, polygons, masks, attributes, classes or extracted document fields. Choosing a box when the model must learn an exact boundary, or a mask when a coarse location is enough.
Video labels What happened, when, and to which object? Task boundaries, tracks across frames, segment captions and episode-level success or failure. Labelling isolated frames when the model needs the continuity of an action.
Text labels Which answer is better, which rule passed, or which span carries the meaning? Preferences, rubric scores, entity and intent spans, taxonomy classes or policy labels. Using a vague rating such as “good” where a stranger needs an applicable criterion.

Transcription writes the words

Use transcription when the target is a reliable textual account of the audio. On Perit’s default verbatim convention, “uh”, repeated words and abandoned starts remain because they happened. Speaker attribution, diarization and entity or intent spans can sit beside the transcript when the model needs to distinguish who said what or which tokens carry the transaction.

Then the edge cases arrive. Is “twenty-one” one token or two? How should a partially audible account number appear? A transcription is not merely somebody typing; it is somebody applying the same written decisions thousands of times.

Alignment puts time under those words

Use alignment when correctness depends on when the system heard something, not only what it heard. Each word receives a start and an end. That makes it possible to inspect latency, subtitle timing, turn-taking or the exact interval around an error.

Alignment presupposes a transcript. If you already have a trusted transcript, it can go directly to the alignment line. If you have only audio, the sensible order may contain both: transcribe the utterance, adjudicate the words, then align those words to the waveform. Precisely timestamping the wrong transcript produces precise wrongness.

Images need geometry; video needs continuity; text needs a rule

For an image, the decision is often about the shape of the answer. A bounding box says roughly where an object is; a polygon or mask traces it more closely; an attribute says something about it. The most detailed label is not automatically the best one. Detail that the model never uses adds cost and creates extra boundary decisions for annotators to disagree about.

Video adds time. A box may need to follow the same object across frames. A long episode may need task and sub-task boundaries, captions for each segment, and a final success label. For embodied data, the difference between “the gripper touched the cup” and “the cup was lifted and remained stable” is the difference between an event and an outcome.

Text usually makes the hidden standard visible. A preference label says which of two replies is better. A rubric score says which specific criteria a response passed. A span label identifies the words that carry an entity or intent. If two careful people cannot apply the criterion the same way, the first suspect should be the criterion, not the people.

Scope collection and annotation together

Request Data is for the team that cannot honestly answer the earlier questions yet. It knows the model fails on one call, one shelf, one form or one style of conversation. It does not know how many examples will expose the pattern, which conditions need quotas, whether the delivery needs transcripts or masks, or how those labels should be judged.

Request Data covers collection and annotation together. Perit turns the failure into a volume, locale and condition plan, collects the inputs, annotates them to the agreed depth and licenses the delivery. The team reviews a scope and quote, checks the five-percent pilot, then receives weekly batches with quality reports and reviewer records.

A request-shaped failure: Your support agent mishandles callers who change the requested amount halfway through a sentence. You have three examples and no stable taxonomy for them. The right opening request is not “50,000 aligned utterances”. It is the three failed calls, the correct behaviour and the question you need the set to answer. Collection volume and annotation depth can then follow from the evidence.

Not knowing the specification is not the same as having no requirements. You still know what failed and what a successful model should do. Request Data is useful precisely in the space between those two facts, where a data programme has to be designed.

Calibrating contributors before production

Once a route has been chosen, the more important trust question begins: how does a stranger learn to make the same decision your team would make?

Each data type uses the same calibration process. The waveform or image changes; the contributor still trains against a guide, passes a test and receives production review.

01Read the guideOne versioned convention defines the labels and edge cases.
02Train on ~30 samplesAnswers are visible, so the written rule becomes concrete.
03Pass a held-back testNo pass means no access to paid production work.
04Produce under samplingQA continues; work below threshold returns to calibration.

The guide comes first because agreement cannot be inspected without an agreed answer. Training uses about thirty examples with answers visible; the test withholds them. Production access follows only after a pass, and production itself is sampled rather than trusted forever.

Four numbers then make the process inspectable. QA-sample accuracy asks whether reviewed items match the guide. WER, where the work is audio, compares the transcript with an adjudicated reference. Inter-annotator agreement asks whether independent people reach the same answer on the double-passed portion. Rework rate records how much fell below threshold and had to be redone. The published policy is to ship these numbers with the batch whether or not they flatter the vendor.

A clean accuracy number after silent corrections tells you less than the same number beside the amount of correction required. Rework is part of the account of how delivered quality was produced.

The gate also explains why the taxonomy is not the deepest decision. Collection, transcription, alignment and masks are different crafts. They can still fail in the same way: an ambiguous guide lets two sensible people create two incompatible datasets. The line on the invoice tells you what work was ordered. Calibration tells you whether the work means the same thing from one item to the next.

Most real projects occupy more than one box

A clean decision tree is helpful at the entrance. Real work often becomes a pipeline.

  • A speech team may collect spontaneous calls in a missing locale, transcribe them verbatim, align every word and mark speaker turns.
  • A retail vision team may collect shelf photographs across devices and lighting, then add boxes, masks and product attributes.
  • An agent team may author sector-specific dialogues, generate candidate replies, and ask practising operators for pairwise preferences and rubric scores.
  • A robotics team may record egocentric episodes, segment the tasks, track objects and label whether each episode succeeded.

“Both” is therefore not an indecisive answer. It is often the technically correct one. The discipline lies in keeping the stages explicit: which evidence must be created, which truth must be attached to it, which rights cover its use, and which report will show that the specification survived production.

 

A sample and a pilot are not the same promise

There are two small-batch moments in this process, and they answer different questions.

The annotation sample tests the bench on data you already recognise. Perit’s published offer is ten minutes of audio, 100 images or one page of text, returned labelled within five working days with an agreement report. It lets you inspect the labels, edge-case decisions and measured agreement before a larger scope exists.

The five-percent pilot tests an agreed programme. By then, the quota, collection conditions, schema and delivery format have been written down. The pilot asks whether that complete operating specification works when real contributors and real files pass through it. A good sample can establish that the bench understands your labels; a good pilot establishes that the whole pipeline can produce the intended batch.

The shortest decision rule If the failed situation is absent, collect it. If the situation is present but the answer is absent, annotate it. If you possess only the failure and the desired behaviour, request the end-to-end set. If the work crosses those boundaries, order both stages and keep their acceptance criteria separate.

Return to the twelve-second clip

The customer says an amount. The model writes the wrong number. We can now ask a better question than “Do we need more data?”

Do our files contain enough calls with this speaker profile, device and noise condition? If not, the first move is collection. Do the calls exist but lack adjudicated verbatim references? That is transcription. Do we need to know whether the model recognised the number late, or heard the wrong number altogether? That requires word-level alignment on top of the transcript. Do we have only this failure and no defensible way to turn it into quotas, volume and labels? Start with Request Data.

The required volume now follows from the failures the batch must cover. Label depth follows from the model decision being evaluated, and the quality report checks the written standard.

You do not have to arrive knowing the name of the dataset. You should arrive with one thing your model gets wrong.

Send the sample before the slide deck

Start with ten minutes of audio, 100 images or one page of text. Send Perit a sample and get it back labelled within five working days with the agreement report. If the needed inputs do not exist yet, send the failed input instead; that is enough to begin the collection and annotation spec.

Was this article helpful? 4/5 1 rating