A quality report arrives with “95%+ QA-sample accuracy” at the top. Before accepting that number, the team needs to know what was scored. A batch can follow its annotation guide, match a transcription reference or maintain quality in production. Those are separate checks.
Take a thirteen-word sentence from a refund call: “the balance owing on the account is two hundred and forty seven dollars.” A transcriber writes “the balance owing on the account is two forty seven dollars.” The words “hundred” and “and” are missing. To judge that error, we need the transcription reference and the scoring convention.
Three measurements help explain the headline: word error rate, inter-annotator agreement and rework rate. Each catches a different failure.
A single number is doing three jobs
Start with what each measurement actually asks.
Word error rate compares the annotator’s transcript with an adjudicated reference. It measures the edits needed to turn one string into the other. That makes it useful for transcription; it is less suited to judgments without a reference answer.
Inter-annotator agreement measures how often two trained people independently reach the same answer. Product-name spans, preferences between model replies and rubric scores may all require judgments. Agreement helps show whether the guide gives people enough detail to apply those judgments consistently.
Rework rate tracks the share of sampled production work that fails QA and has to be redone. It helps reveal drift after calibration, including a decline in quality from someone who passed the original test.
Distance from an agreed reference. Answers: was the transcript correct?
Distance between two independent annotators. Answers: is the task well-specified?
Share of sampled production work sent back. Answers: is quality holding up over time?
High agreement does not guarantee useful labels. Two annotators can follow the same rigid guide and both miss what the client needed. Likewise, a transcript can match every reference word while reviewers disagree about whether the call was resolved. Read the metric that corresponds to the task.
What WER is actually counting
WER counts the minimum substitutions, deletions and insertions needed to turn a hypothesis transcript into the reference.
Take the refund sentence. The reference has twelve words:
The reference is “the / balance / owing / on / the / account / is / two / hundred / and / forty / seven / dollars”: thirteen words. The hypothesis deletes “hundred” and “and”; nothing else changes.
With S = 0, D = 2, I = 0 and N = 13, WER = 2/13 ≈ 15.4%. The sentence may look plausible on a quick read, but two missing words out of thirteen give it a double-digit error rate.
Now change one thing: suppose the guide does not require filler words to be transcribed, and the only difference between hypothesis and reference is that the caller said “um” once and the annotator correctly dropped it. A guide that scores fillers would count that as a deletion; a guide that excludes fillers scores it as zero error on the same audio. The number moves by a guide decision, not by annotator skill, which is why the report has to name the guide version next to the score.
This is why a WER on its own, without the guide version it was scored against, is close to unreadable. The number is real, but the ruler it was measured with is a choice.
Using disagreement to improve the guide
Span boundaries, response preferences and call-resolution labels often require judgment. The guide defines what counts as a correct label. Agreement then tests whether a second trained annotator can apply that definition the same way.
Take a span-labelling task: mark the product name in “I want the Samsung Galaxy S24 Ultra in the black colour.” One annotator, trained to capture the full retail name, marks the span as “Samsung Galaxy S24 Ultra.” A second annotator, reading the same guide slightly differently, marks only “Galaxy S24 Ultra,” treating “Samsung” as the brand rather than part of the product name.
A simple way to score the overlap is token-level agreement: the number of shared tokens divided by the number of tokens either annotator included:
The first span is {Samsung, Galaxy, S24, Ultra}, four tokens. The second is {Galaxy, S24, Ultra}, three tokens. Their intersection has three tokens and their union has four, so agreement is 3/4 = 75%.
Here the 75% agreement exposes a guide gap. Neither annotator was told whether the brand name belongs in the span. Both decisions are defensible, so the next step is to clarify the guide.
This is the practical use of inter-annotator agreement: it is measured on a double-passed slice of the work specifically so that low agreement can be traced back to guide language before it spreads through the rest of the batch.
Rework rate is the number that catches drift
WER and agreement describe a batch, or a double-passed sample, at a particular time. They do not establish that an annotator who passed calibration in week one still meets the standard in week six.
Rework rate measures sampled production items that fail ongoing QA and return to the annotator or trigger recalibration. These are items produced after the original calibration test.
| Metric | Measured on | Catches | Misses |
|---|---|---|---|
| WER | Every transcribed item, against an adjudicated reference | Transcription errors relative to the agreed guide | Whether the guide itself is well-specified |
| Inter-annotator agreement | A double-passed sample | Ambiguous or under-specified guide language | Drift that develops after calibration passes |
| Rework rate | An ongoing sample of production work | Individual or bench-wide drift over time | Errors too subtle for the sampling rate to surface |
A rising rework rate alongside a stable WER needs investigation. The transcripts checked for WER may still match their references while more production items fail other reviews. Check which annotators, conditions and error types account for the increase before changing sampling or recalibration.
Audio example: a batch reports 96.5% QA-sample accuracy, 4.1% WER against the adjudicated reference, 91% agreement on the double-passed sample and 2% rework. Those figures describe guide compliance, transcript errors, consistency between annotators and production failures. If rework rose to 14%, the team would need to investigate even if WER and agreement stayed stable.
The same checks apply outside audio, with a different reference metric. For bounding boxes, intersection-over-union measures how closely the box matches an adjudicated reference. For video success/failure labels, use accuracy against reviewed outcomes. Agreement and rework remain useful in both cases.
Image example: a shelf-detection batch has 0.88 mean IoU against the reference, 94% agreement on the double-passed sample and 3% rework. The averages still need a condition breakdown. If low-light, crowded shelves account for the weak boxes, that matters to a team whose model already fails on those photos.
Which metric to check first
The three numbers do not deserve equal attention on every project, because they are sensitive to different things going wrong.
- Audio and document transcription: WER first, because a reference exists and the task is close to objectively checkable. Watch the guide version next to it. a WER improving because the guide changed what counts as an error is not the same as annotators getting better.
- Spans, preferences, rubric scores, anything subjective: agreement first. A low WER-equivalent metric here usually just means the guide needs another pass, not that the bench needs replacing.
- Any multi-week or multi-batch engagement: rework rate first, tracked as a trend rather than a single value. A single good batch tells you the bench can do the work; a flat rework rate over eight batches tells you the process holds.
Ask for the reference metric, agreement and rework rate, along with the guide version and reviewer chain. Together they explain more than the headline accuracy figure.
Guide versions matter because edge-case decisions change. A batch reviewed against v3 is not automatically comparable with one reviewed against v4, even if both report 95% accuracy. The reviewer chain records who applied those rules and gives a later audit somewhere to start.
Ask to see all three numbers before the first batch
Perit ships QA-sample accuracy, WER where applicable, inter-annotator agreement and rework rate with every batch, alongside the guide version and reviewer chain that produced them. See the full quality system, or check the FAQ for how the calibration gate and sampling cadence work.