{"id":775,"date":"2026-09-20T13:24:26","date_gmt":"2026-09-20T13:24:26","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=775"},"modified":"2026-09-30T18:42:42","modified_gmt":"2026-09-30T18:42:42","slug":"collection-annotation-or-both","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/collection-annotation-or-both\/","title":{"rendered":"Collection, Annotation or Both?"},"content":{"rendered":"\n\r\n<div class=\"pp-post\" style=\"max-width: 50rem;margin: 0 auto;font-family: Verdana, Geneva, Tahoma, sans-serif;font-size: 1rem;line-height: 1.68;color: #0b0b0b\">\r\n<div class=\"meta\" style=\"font-size: 0.92rem;color: #52514e;margin-bottom: 1.6rem;padding-bottom: 0.65rem;border-bottom: 1px solid #e6e5e0\">One failed input, three doors, five lines on the annotation bench, and the gate every annotator must clear<\/div>\r\n<p class=\"lede\" style=\"margin: 0 0 1rem;font-size: 1.08rem\">A customer gives an amount over a noisy phone line. The speech model records the wrong number, and the rest of the system processes it correctly. In the meeting that follows, someone asks for several hundred thousand rows of \u201cmore data\u201d. The failed clip is twelve seconds long. It is worth examining before choosing the order size.<\/p>\r\n<p style=\"margin: 0 0 1rem\">\u201cMore data\u201d can mean three different jobs.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The team may need new recordings because the acoustic condition, accent or interaction that breaks the model is absent from its files. It may already own the right recordings, in which case it needs those files transcribed, aligned or otherwise labelled. Or it may know only the failure, the twelve-second clip, and need someone to turn that failure into a collection protocol, an annotation guide and a licensable delivery.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Those routes are related, but they are not interchangeable. Annotation cannot create the warehouse recording you never captured. Collection does not automatically tell you which word was spoken at 7.4 seconds. And \u201ctranscription versus alignment\u201d is not the first fork at all: both are jobs inside annotation.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Before choosing a service, check whether the inputs your model needs already exist.<\/p>\r\n<div class=\"decision\" style=\"margin: 2.2rem 0\" aria-label=\"A three-route decision guide\">\r\n<div class=\"decision-title\" style=\"margin: 0 0 0.85rem;font-weight: bold;font-size: 1rem;line-height: 1.45;color: #0b0b0b\">Do you already have the inputs that contain the failure?<\/div>\r\n<div class=\"route-grid\" style=\"display: flex;flex-wrap: wrap;gap: 0.75rem\">\r\n<div class=\"route\" style=\"flex: 1 1 13rem;min-width: 13rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-top: 3px solid #2a78d6;border-radius: 10px;background-color: #fcfcfb\"><span class=\"answer\" style=\"display: block;margin: 0 0 0.35rem;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4;text-transform: uppercase;color: #2a78d6\">No<\/span> <strong class=\"route-name\" style=\"display: block;margin: 0 0 0.3rem;font-size: 1rem;font-weight: bold;color: #0b0b0b\">Data Collection<\/strong>\r\n<p style=\"margin: 0;font-size: 0.9rem;line-height: 1.55;color: #52514e\">Record or author the missing situations, people, devices, rooms or language.<\/p>\r\n<\/div>\r\n<div class=\"route route-b\" style=\"flex: 1 1 13rem;min-width: 13rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-top: 3px solid #2a78d6;border-radius: 10px;background-color: #fcfcfb;border-top-color: #eb6834\"><span class=\"answer\" style=\"display: block;margin: 0 0 0.35rem;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4;text-transform: uppercase;color: #2a78d6\">Yes<\/span> <strong class=\"route-name\" style=\"display: block;margin: 0 0 0.3rem;font-size: 1rem;font-weight: bold;color: #0b0b0b\">Data Annotation<\/strong>\r\n<p style=\"margin: 0;font-size: 0.9rem;line-height: 1.55;color: #52514e\">Turn files you own into transcripts, timestamps, masks, tracks, spans or judgments.<\/p>\r\n<\/div>\r\n<div class=\"route route-c\" style=\"flex: 1 1 13rem;min-width: 13rem;padding: 0.95rem 1rem 1rem;border: 1px solid #e6e5e0;border-top: 3px solid #2a78d6;border-radius: 10px;background-color: #fcfcfb;border-top-color: #1baf7a\"><span class=\"answer\" style=\"display: block;margin: 0 0 0.35rem;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4;text-transform: uppercase;color: #2a78d6\">I only know the failure<\/span> <strong class=\"route-name\" style=\"display: block;margin: 0 0 0.3rem;font-size: 1rem;font-weight: bold;color: #0b0b0b\">Request Data<\/strong>\r\n<p style=\"margin: 0;font-size: 0.9rem;line-height: 1.55;color: #52514e\">Start with one bad input; scope collection and annotation together.<\/p>\r\n<\/div>\r\n<\/div>\r\n<\/div>\r\n<h2 id=\"start-with-the-failure-because-the-noun-can-wait\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Check what is missing before choosing a service<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Suppose your model misses product names when a shopkeeper speaks beside a refrigerator. \u201cWe need annotated audio\u201d sounds like a reasonable brief. It is not yet enough to choose a service.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Open the files you already have. If they contain the same shops, accents, devices and background hum, the missing ingredient may indeed be labels: a verbatim reference transcript, word-level times, speaker turns or entity spans around the product names. If the files are mostly clean office speech, better labels will make the clean office data more precise. They will not introduce the refrigerator, the handset microphone or the shopkeeper.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Start with the failed input. Does that condition appear in your files? What should a person label? Which measurement would show that the new batch helped?<\/p>\r\n<p style=\"margin: 0 0 1rem\">This is why all three Perit routes use nearly the same opening language: send a clip, photograph, paragraph or task the model handles badly. You do not need to solve the data specification before the first conversation. You do need to bring the failure close enough that somebody can inspect it.<\/p>\r\n<div class=\"quote-block\" style=\"margin: 1.8rem 0;padding: 0.15rem 0 0.15rem 1.1rem;border-left: 3px solid #0b0b0b;font-size: 1.05rem;font-style: italic\">A service name is a conclusion. The failed input is the evidence that gets you there.<\/div>\r\n<h2 id=\"door-one-the-situation-is-missing-so-collect-it\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Collect the missing situations<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Choose <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/services\/data-collection\" target=\"_blank\" rel=\"noopener\">Data Collection<\/a> when the model needs experience your current files do not contain. That missing experience might be a locale, an accent, a room, a camera angle, a handset, a lighting condition, a document type or a way people phrase the same intention.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Collection is not restricted to speech. Perit lists four modalities on the same recruiting and QA machinery: speech and audio; video; image; and text and conversation. A contributor might record a two-party call with each speaker on a separate channel, film an everyday task on a chest camera, photograph the same shelf under two lighting conditions, or write a dialogue in the register of the sector they work in. The protocol and equipment change. Recruiting to quotas, capturing consent, running a pilot, reviewing quality and delivering weekly remain the same.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Speech is currently the deepest line: <strong>4,000+ hours delivered, nine live locales and capacity for 50 hours a day<\/strong>. Those figures establish throughput, but the collection spec must still name the people, conditions, devices and outcomes that belong in your batch.<\/p>\r\n<p style=\"margin: 0 0 1rem\">That is the purpose of the pilot. After the brief, Perit returns a scope and quote; then <strong>five percent of the planned volume<\/strong> is delivered before the main run. The pilot exists so a vague instruction becomes an obvious problem on a small batch. Perhaps \u201cnoisy shop\u201d produces shopping-centre music when the model actually fails on compressor hum. Perhaps \u201cHindi speakers\u201d is too broad for a regional failure. It is much cheaper to discover that at five percent than at one hundred.<\/p>\r\n<div class=\"case\" style=\"margin: 1.6rem 0;padding: 0.95rem 1.1rem;border-left: 3px solid #2a78d6;border-radius: 0 10px 10px 0;background-color: #fcfcfb\">\r\n<p style=\"margin: 0\"><strong>A collection-shaped failure:<\/strong> Your image model performs well on catalogue photographs but misses products on crowded shelves photographed by customers at night. You have labels for clean product images; you do not have the user device, shelf density or light. The missing thing is not another pass over the old boxes. It is new imagery collected across the conditions in which the model is expected to work.<\/p>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">Collection can include labels where ordered. That does not blur the distinction; it means one work order can cross it. The people and situations are collected first, then the resulting files move through an annotation specification. \u201cCollection\u201d describes how the raw evidence comes into existence, not a requirement that it remain raw.<\/p>\r\n<h2 id=\"door-two-the-evidence-exists-so-annotate-it\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Annotate the inputs you already have<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Choose <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/services\/data-annotation\" target=\"_blank\" rel=\"noopener\">Data Annotation<\/a> when the useful inputs already sit in your bucket but the answer is not attached to them. This is also the direct answer to a question the <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/faq\" target=\"_blank\" rel=\"noopener\">FAQ<\/a> makes explicit: yes, Perit annotates data customers already own; in fact, most annotation work is performed on customer audio, images or video.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The work runs through <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/foundry.zyka.ai\" target=\"_blank\" rel=\"noopener\">Foundry<\/a>. The published bench has <strong>650 people<\/strong>: 250 transcribers, including 50 on fixed-term contracts, and 400 freelance word aligners. It has delivered more than 3,000 transcribed hours and more than 1,000 aligned hours across nine locales. The service page reports 95%+ QA-sample accuracy.<\/p>\r\n<p style=\"margin: 0 0 1rem\">That last number deserves its full name. It is QA-sample accuracy, checked against the project standard; it is not a promise that every possible label on every possible dataset is \u201c95% accurate\u201d. For audio, word error rate is reported separately against an adjudicated reference. Keeping those measures separate is useful because a batch can follow its guide consistently while the guide itself is incomplete, or produce a good transcript while still placing a timing boundary poorly.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Graders stream the files rather than copy them, and the project is pinned to the agreed region. What remains is to decide what answer should be attached to each input.<\/p>\r\n<h2 id=\"inside-annotation-there-are-five-practical-lines\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Choosing the annotation task<\/h2>\r\n<p style=\"margin: 0 0 1rem\">This is where transcription and alignment finally become the right question. Audio happens to split into two distinct crafts, while image, video and text each need their own label geometry. A useful way to choose is to finish this sentence: <em>After annotation, I need to know\u2026<\/em><\/p>\r\n<div class=\"table-wrap\" style=\"max-width: 100%;overflow: auto;margin: 1.8rem 0;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\">\r\n<table style=\"width: 100%;min-width: 690px;border-collapse: collapse;font-size: 0.9rem;line-height: 1.5\">\r\n<thead>\r\n<tr>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;font-size: 0.82rem;letter-spacing: 0.02em;color: #0b0b0b\">Bench line<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;font-size: 0.82rem;letter-spacing: 0.02em;color: #0b0b0b\">The question it answers<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;font-size: 0.82rem;letter-spacing: 0.02em;color: #0b0b0b\">What comes back<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;font-size: 0.82rem;letter-spacing: 0.02em;color: #0b0b0b\">The common mix-up<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td class=\"row-name\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;width: 16%;font-weight: bold\">Transcription<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">What was actually said?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Verbatim text, with fillers, repeats and false starts retained under the written convention.<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Asking for a \u201cclean\u201d transcript, then trying to evaluate a model on disfluencies that were edited away.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"row-name\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;width: 16%;font-weight: bold\">Alignment<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">When did each word begin and end?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Word-level start and end timestamps placed on top of a transcript.<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Treating alignment as a substitute for transcription. A timestamp needs a word to belong to.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"row-name\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;width: 16%;font-weight: bold\">Image labels<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Where is the object, and what is true of it?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Boxes, polygons, masks, attributes, classes or extracted document fields.<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Choosing a box when the model must learn an exact boundary, or a mask when a coarse location is enough.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"row-name\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;width: 16%;font-weight: bold\">Video labels<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">What happened, when, and to which object?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Task boundaries, tracks across frames, segment captions and episode-level success or failure.<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;border-bottom: 1px solid #e6e5e0;color: #52514e\">Labelling isolated frames when the model needs the continuity of an action.<\/td>\r\n<\/tr>\r\n<tr>\r\n<td class=\"row-name row-last\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;color: #0b0b0b;width: 16%;font-weight: bold;border-bottom: 0\">Text labels<\/td>\r\n<td class=\"row-last\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;color: #52514e;border-bottom: 0\">Which answer is better, which rule passed, or which span carries the meaning?<\/td>\r\n<td class=\"row-last\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;color: #52514e;border-bottom: 0\">Preferences, rubric scores, entity and intent spans, taxonomy classes or policy labels.<\/td>\r\n<td class=\"row-last\" style=\"text-align: left;vertical-align: top;padding: 0.7rem 0.8rem;color: #52514e;border-bottom: 0\">Using a vague rating such as \u201cgood\u201d where a stranger needs an applicable criterion.<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<h3 style=\"font-size: 1.05rem;line-height: 1.4;margin: 1.65rem 0 0.45rem;font-weight: bold;color: #0b0b0b\">Transcription writes the words<\/h3>\r\n<p style=\"margin: 0 0 1rem\">Use transcription when the target is a reliable textual account of the audio. On Perit&#8217;s default verbatim convention, \u201cuh\u201d, repeated words and abandoned starts remain because they happened. Speaker attribution, diarization and entity or intent spans can sit beside the transcript when the model needs to distinguish who said what or which tokens carry the transaction.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Then the edge cases arrive. Is \u201ctwenty-one\u201d one token or two? How should a partially audible account number appear? A transcription is not merely somebody typing; it is somebody applying the same written decisions thousands of times.<\/p>\r\n<h3 style=\"font-size: 1.05rem;line-height: 1.4;margin: 1.65rem 0 0.45rem;font-weight: bold;color: #0b0b0b\">Alignment puts time under those words<\/h3>\r\n<p style=\"margin: 0 0 1rem\">Use alignment when correctness depends on <em>when<\/em> the system heard something, not only <em>what<\/em> it heard. Each word receives a start and an end. That makes it possible to inspect latency, subtitle timing, turn-taking or the exact interval around an error.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Alignment presupposes a transcript. If you already have a trusted transcript, it can go directly to the alignment line. If you have only audio, the sensible order may contain both: transcribe the utterance, adjudicate the words, then align those words to the waveform. Precisely timestamping the wrong transcript produces precise wrongness.<\/p>\r\n<h3 style=\"font-size: 1.05rem;line-height: 1.4;margin: 1.65rem 0 0.45rem;font-weight: bold;color: #0b0b0b\">Images need geometry; video needs continuity; text needs a rule<\/h3>\r\n<p style=\"margin: 0 0 1rem\">For an image, the decision is often about the shape of the answer. A bounding box says roughly where an object is; a polygon or mask traces it more closely; an attribute says something about it. The most detailed label is not automatically the best one. Detail that the model never uses adds cost and creates extra boundary decisions for annotators to disagree about.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Video adds time. A box may need to follow the same object across frames. A long episode may need task and sub-task boundaries, captions for each segment, and a final success label. For embodied data, the difference between \u201cthe gripper touched the cup\u201d and \u201cthe cup was lifted and remained stable\u201d is the difference between an event and an outcome.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Text usually makes the hidden standard visible. A preference label says which of two replies is better. A rubric score says which specific criteria a response passed. A span label identifies the words that carry an entity or intent. If two careful people cannot apply the criterion the same way, the first suspect should be the criterion, not the people.<\/p>\r\n<h2 id=\"door-three-you-have-a-failure-but-not-yet-a-specification\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Scope collection and annotation together<\/h2>\r\n<p style=\"margin: 0 0 1rem\"><a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/request-data\" target=\"_blank\" rel=\"noopener\">Request Data<\/a> is for the team that cannot honestly answer the earlier questions yet. It knows the model fails on one call, one shelf, one form or one style of conversation. It does not know how many examples will expose the pattern, which conditions need quotas, whether the delivery needs transcripts or masks, or how those labels should be judged.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Request Data covers collection and annotation together. Perit turns the failure into a volume, locale and condition plan, collects the inputs, annotates them to the agreed depth and licenses the delivery. The team reviews a scope and quote, checks the five-percent pilot, then receives weekly batches with quality reports and reviewer records.<\/p>\r\n<div class=\"case\" style=\"margin: 1.6rem 0;padding: 0.95rem 1.1rem;border-left: 3px solid #2a78d6;border-radius: 0 10px 10px 0;background-color: #fcfcfb\">\r\n<p style=\"margin: 0\"><strong>A request-shaped failure:<\/strong> Your support agent mishandles callers who change the requested amount halfway through a sentence. You have three examples and no stable taxonomy for them. The right opening request is not \u201c50,000 aligned utterances\u201d. It is the three failed calls, the correct behaviour and the question you need the set to answer. Collection volume and annotation depth can then follow from the evidence.<\/p>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">Not knowing the specification is not the same as having no requirements. You still know what failed and what a successful model should do. Request Data is useful precisely in the space between those two facts, where a data programme has to be designed.<\/p>\r\n<h2 id=\"the-same-gate-sits-behind-every-answer\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Calibrating contributors before production<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Once a route has been chosen, the more important trust question begins: how does a stranger learn to make the same decision your team would make?<\/p>\r\n<p style=\"margin: 0 0 1rem\">Each data type uses the same calibration process. The waveform or image changes; the contributor still trains against a guide, passes a test and receives production review.<\/p>\r\n<div class=\"gate\" style=\"display: flex;flex-wrap: wrap;gap: 0.75rem;margin: 1.8rem 0\" aria-label=\"The four-stage annotation calibration gate\">\r\n<div class=\"gate-step\" style=\"flex: 1 1 10rem;min-width: 10rem;padding: 0.9rem 0.9rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin: 0 0 0.35rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4\">01<\/span><strong style=\"display: block;margin: 0 0 0.3rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Read the guide<\/strong><span class=\"step-text\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">One versioned convention defines the labels and edge cases.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 10rem;min-width: 10rem;padding: 0.9rem 0.9rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin: 0 0 0.35rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4\">02<\/span><strong style=\"display: block;margin: 0 0 0.3rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Train on ~30 samples<\/strong><span class=\"step-text\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">Answers are visible, so the written rule becomes concrete.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 10rem;min-width: 10rem;padding: 0.9rem 0.9rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin: 0 0 0.35rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4\">03<\/span><strong style=\"display: block;margin: 0 0 0.3rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Pass a held-back test<\/strong><span class=\"step-text\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">No pass means no access to paid production work.<\/span><\/div>\r\n<div class=\"gate-step\" style=\"flex: 1 1 10rem;min-width: 10rem;padding: 0.9rem 0.9rem 1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb\"><span class=\"step-no\" style=\"display: block;margin: 0 0 0.35rem;color: #2a78d6;font-size: 0.78rem;font-weight: bold;letter-spacing: 0.08em;line-height: 1.4\">04<\/span><strong style=\"display: block;margin: 0 0 0.3rem;font-size: 0.92rem;line-height: 1.4;color: #0b0b0b\">Produce under sampling<\/strong><span class=\"step-text\" style=\"display: block;font-size: 0.84rem;line-height: 1.5;color: #52514e\">QA continues; work below threshold returns to calibration.<\/span><\/div>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">The guide comes first because agreement cannot be inspected without an agreed answer. Training uses about thirty examples with answers visible; the test withholds them. Production access follows only after a pass, and production itself is sampled rather than trusted forever.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Four numbers then make the process inspectable. <strong>QA-sample accuracy<\/strong> asks whether reviewed items match the guide. <strong>WER<\/strong>, where the work is audio, compares the transcript with an adjudicated reference. <strong>Inter-annotator agreement<\/strong> asks whether independent people reach the same answer on the double-passed portion. <strong>Rework rate<\/strong> records how much fell below threshold and had to be redone. The published policy is to ship these numbers with the batch whether or not they flatter the vendor.<\/p>\r\n<p style=\"margin: 0 0 1rem\">A clean accuracy number after silent corrections tells you less than the same number beside the amount of correction required. Rework is part of the account of how delivered quality was produced.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The gate also explains why the taxonomy is not the deepest decision. Collection, transcription, alignment and masks are different crafts. They can still fail in the same way: an ambiguous guide lets two sensible people create two incompatible datasets. The line on the invoice tells you what work was ordered. Calibration tells you whether the work means the same thing from one item to the next.<\/p>\r\n<h2 id=\"most-real-projects-occupy-more-than-one-box\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Most real projects occupy more than one box<\/h2>\r\n<p style=\"margin: 0 0 1rem\">A clean decision tree is helpful at the entrance. Real work often becomes a pipeline.<\/p>\r\n<ul style=\"margin: 0.35rem 0 1.2rem 1.35rem;padding: 0\">\r\n<li style=\"margin: 0 0 0.45rem;padding-left: 0.15rem\">A speech team may collect spontaneous calls in a missing locale, transcribe them verbatim, align every word and mark speaker turns.<\/li>\r\n<li style=\"margin: 0 0 0.45rem;padding-left: 0.15rem\">A retail vision team may collect shelf photographs across devices and lighting, then add boxes, masks and product attributes.<\/li>\r\n<li style=\"margin: 0 0 0.45rem;padding-left: 0.15rem\">An agent team may author sector-specific dialogues, generate candidate replies, and ask practising operators for pairwise preferences and rubric scores.<\/li>\r\n<li style=\"margin: 0 0 0.45rem;padding-left: 0.15rem\">A robotics team may record egocentric episodes, segment the tasks, track objects and label whether each episode succeeded.<\/li>\r\n<\/ul>\r\n<p style=\"margin: 0 0 1rem\">\u201cBoth\u201d is therefore not an indecisive answer. It is often the technically correct one. The discipline lies in keeping the stages explicit: which evidence must be created, which truth must be attached to it, which rights cover its use, and which report will show that the specification survived production.<\/p>\r\n<p style=\"margin: 0 0 1rem\">\u00a0<\/p>\r\n<h2 id=\"a-sample-and-a-pilot-are-not-the-same-promise\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">A sample and a pilot are not the same promise<\/h2>\r\n<p style=\"margin: 0 0 1rem\">There are two small-batch moments in this process, and they answer different questions.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The annotation sample tests the bench on data you already recognise. Perit&#8217;s published offer is <strong>ten minutes of audio, 100 images or one page of text<\/strong>, returned labelled within <strong>five working days<\/strong> with an agreement report. It lets you inspect the labels, edge-case decisions and measured agreement before a larger scope exists.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The five-percent pilot tests an agreed programme. By then, the quota, collection conditions, schema and delivery format have been written down. The pilot asks whether that complete operating specification works when real contributors and real files pass through it. A good sample can establish that the bench understands your labels; a good pilot establishes that the whole pipeline can produce the intended batch.<\/p>\r\n<div class=\"note\" style=\"margin: 1.8rem 0;padding: 0.95rem 1.1rem;border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;font-size: 0.92rem\"><strong style=\"display: block;margin: 0 0 0.3rem;color: #0b0b0b\">The shortest decision rule<\/strong> If the failed situation is absent, collect it. If the situation is present but the answer is absent, annotate it. If you possess only the failure and the desired behaviour, request the end-to-end set. If the work crosses those boundaries, order both stages and keep their acceptance criteria separate.<\/div>\r\n<h2 id=\"return-to-the-twelve-second-clip\" style=\"font-size: 1.28rem;line-height: 1.35;margin: 2.5rem 0 0.7rem;font-weight: bold;color: #0b0b0b\">Return to the twelve-second clip<\/h2>\r\n<p style=\"margin: 0 0 1rem\">The customer says an amount. The model writes the wrong number. We can now ask a better question than \u201cDo we need more data?\u201d<\/p>\r\n<p style=\"margin: 0 0 1rem\">Do our files contain enough calls with this speaker profile, device and noise condition? If not, the first move is collection. Do the calls exist but lack adjudicated verbatim references? That is transcription. Do we need to know whether the model recognised the number late, or heard the wrong number altogether? That requires word-level alignment on top of the transcript. Do we have only this failure and no defensible way to turn it into quotas, volume and labels? Start with Request Data.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The required volume now follows from the failures the batch must cover. Label depth follows from the model decision being evaluated, and the quality report checks the written standard.<\/p>\r\n<div class=\"quote-block\" style=\"margin: 1.8rem 0;padding: 0.15rem 0 0.15rem 1.1rem;border-left: 3px solid #0b0b0b;font-size: 1.05rem;font-style: italic\">You do not have to arrive knowing the name of the dataset. You should arrive with one thing your model gets wrong.<\/div>\r\n<div class=\"cta\" style=\"margin: 2.4rem 0 0;padding: 1.2rem 1.25rem;border: 2px solid #0b0b0b;border-radius: 10px;background-color: #fcfcfb\">\r\n<h2 id=\"send-the-sample-before-the-slide-deck\" style=\"font-size: 1.28rem;line-height: 1.35;font-weight: bold;color: #0b0b0b;margin: 0 0 0.65rem\">Send the sample before the slide deck<\/h2>\r\n<p style=\"margin: 0\">Start with ten minutes of audio, 100 images or one page of text. <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/services\/data-annotation\" target=\"_blank\" rel=\"noopener\">Send Perit a sample<\/a> and get it back labelled within five working days with the agreement report. If the needed inputs do not exist yet, <a style=\"color: #0000ee;text-decoration: underline\" href=\"https:\/\/perit.ai\/request-data\" target=\"_blank\" rel=\"noopener\">send the failed input instead<\/a>; that is enough to begin the collection and annotation spec.<\/p>\r\n<\/div>\r\n<\/div>\r\n\n","protected":false},"excerpt":{"rendered":"<p>Does a model failure need new inputs, better labels or a clearer spec? A guide to choosing data collection, annotation or both.<\/p>\n","protected":false},"author":2,"featured_media":780,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[19],"tags":[],"class_list":["post-775","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-research"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/775","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=775"}],"version-history":[{"count":5,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/775\/revisions"}],"predecessor-version":[{"id":927,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/775\/revisions\/927"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/780"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=775"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=775"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=775"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}