Word Timestamps And Forced Alignment for Voice Agents

Word Timestamps And Forced Alignment for Voice Agents
Every turn-taking model, end-of-turn threshold and simulated caller is built from word timestamps. What the alignment benchmarks actually show, where machine timing fails on real conversation, and why the word nobody bothers to write down can decide when a voice agent speaks.

What the timing data shows

  • Word timestamps are inferred, not measured. A pause is whatever is left between two inferred boundaries, so it inherits the error of both.
  • On spontaneous English, WhisperX’s mean word-boundary error (110.9 ms) is more than half the 208 ms average gap before an answer to a yes/no question, measured across ten languages.
  • On Perit’s word-alignment benchmark, the median transcription system times 60% of words within 100 ms of a human reference, and 26% within 50 ms.
  • Clean transcripts delete fillers. “Um” usually announces a coming delay (61% of the time in a classic corpus study), so deleting it removes a cue a turn-taking model needs.

Say this out loud, at the speed you would say it on the phone: yeah, so, um, I was thinking…

After the “um”, the speaker may pause for half a second. A listener expects them to continue because the filler signals a delay. Herbert Clark and Jean Fox Tree argued that “uh” and “um” are not noise but words, with conventional forms and conventional meanings, and that what they mean is roughly I am starting a delay: a short one for “uh”, a longer one for “um”. Speakers use that meaning to signal, among other things, that they are not done. In the London-Lund corpus they studied, “um” was followed by a delay 61% of the time, against 29% for “uh”.[1] Listeners use that cue to decide whether to wait.

Ask a timing pipeline to process the same sentence. The transcript and alignment method can produce three different accounts of it.

Three versions of word timing: human reference shows a 520 ms pause, a missing filler creates 1,010 ms of silence, and folding the pause into a word leaves 0 ms.
Figure 1. Illustrative durations. Top: what was said, timed by a person against a verbatim transcript. Middle: the filler is missing from the transcript, so its time is booked as silence and a one-second pause appears where the speaker was talking. Bottom: the filler is dropped and the silence is folded into the next word, so the pause disappears. Both mechanisms are discussed below with their sources; how large the effect is varies by system and audio.Open full-size chart

In the first, the pause is where it was. In the second, the “um” never made it into the transcript, the aligner books its time as silence, and the record now says the speaker went quiet for a full second after “so” and then carried on, for no visible reason. In the third, the filler is gone and the silence has been swallowed into the start of the next word, so the pause never happened at all.

Our last post argued that a voice agent learns when to wait from an environment: a simulated caller who hesitates the way real callers hesitate, built from pause statistics mined from recorded calls. This post is about the ruler those pauses are measured with. Word timestamps are inferred, not measured. The tools that infer them are most accurate on the speech that matters least for turn-taking and least accurate on the speech that matters most. And their errors do not simply blur the data. They push it in a direction, the direction becomes a threshold, and the threshold becomes behaviour.

A word timestamp is an inference, not a measurement

Nobody measures when a word starts. A forced aligner is given audio and a transcript and asked for the most probable way to lay one over the other. How it answers depends on what kind of model it is, and the three families in common use answer differently.

Family Common examples How it places a boundary Where it struggles
HMM-GMM with a pronunciation dictionary Montreal Forced Aligner (MFA)[2] Models every phone and its duration explicitly, over 10 ms frames Anything missing from the dictionary or the transcript: omitted fillers, cut-off words, a second language
CTC acoustic models WhisperX’s alignment stage,[3] MMS[4] Finds where each token’s probability spikes, over 20 ms frames A spike says a sound happened, not how long it lasted, so word edges are estimated
Attention-based ASR timing Whisper’s own word timestamps[5] Maps the decoder’s attention back onto audio frames Never trained to place word boundaries; tokens carry a leading space, so pauses fold into words

The CTC row deserves a closer look, because it sits inside many popular pipelines. CTC models learn to emit each token as a spike on a single frame and to fill everything else with a “blank” symbol; Zeyer, Schlüter and Ney showed that this peaky behaviour falls out of the training loss itself.[6] Huang and colleagues measured what it does to timing on Buckeye, a corpus of spontaneous English with hand-corrected boundaries. A standard CTC aligner gave the average phone a duration of 21 ms, about one frame, when the real average there is 82 ms, and its word boundaries were off by 58 ms on average.[7] A model that believes every sound lasts one frame does not know where a word ends. It knows roughly where the word happened and fills in the edges.

The attention row has a subtler problem, and it is the one in the third timeline above. Whisper’s tokenizer attaches the space to the front of each word, and the CrisperWhisper authors traced part of Whisper’s timing error to exactly that: the silence before a word gets counted as part of the word. The same system also leaves many fillers out, which the authors describe as an intended transcription style for contexts where clarity of intent matters more than detailed speech analysis.[5] It is a sensible choice for captions and a damaging one for timing data.

Underneath all three sits one structural fact. None of these systems has a concept of a pause. A pause is whatever is left over between two boundaries the aligner has placed: a residual. Every error in placing the end of one word or the start of the next lands, in full, in the pause beside it. And because a pause has two edges, it collects error from both. If each boundary is off by 90 ms on average, and the two errors are unbiased, independent and roughly normal, the pause between them is off by about 127 ms. Real errors break those assumptions in both directions, but the order of magnitude holds.

What alignment benchmarks show when the speech gets real

The most recent careful comparison comes from the MFA team itself. On Buckeye, their 2026 evaluation puts MFA 3.0’s mean word-boundary error at 21.75 ms, MMS at 49.54 ms and WhisperX at 110.90 ms.[8] Those numbers only mean something next to the thing being measured. Across ten languages, Stivers and colleagues found that answers to yes/no questions began, on average, 208 ms after the question ended. In Japanese the average was 7 ms.[9]

Mean word-boundary error on Buckeye: MFA 3.0 21.75 ms, MMS 49.54 ms and WhisperX 110.90 ms. The reference gap is 208 ms.
Figure 2. Mean word-boundary error on Buckeye, from McAuliffe et al. (2026); MFA is the 3.0 ARPA model.[8] The dashed line is the mean gap before an answer across ten languages, from Stivers et al. (2009).[9] The comparison is of scale, not a like-for-like metric.Open full-size chart

WhisperX’s average error is more than half of that average gap. Run it through the two-edge arithmetic and a pause measured with it is off by something on the order of 150 ms, most of the gap it is supposed to help a model understand.

Rousso and colleagues scored the same three systems against human boundaries on TIMIT, which is read sentences, and on Buckeye, which is conversation. Moving from reading to talking hurts WhisperX most and MFA a little, while MMS barely moves:[10]

Aligner Read: within 25 ms Read: within 50 ms Talk: within 25 ms Talk: within 50 ms Talk: median error Talk: mean error
MFA 72.8% 89.4% 69.9% 84.9% 13.6 ms 976.5 ms
MMS 43.5% 75.7% 52.7% 75.0% 23.1 ms 208.3 ms
WhisperX 52.7% 82.4% 43.1% 67.4% 30.1 ms 11,685.3 ms

Two details in that table matter more than the ranking. The first is the last two columns. WhisperX’s median error on conversation is 30 ms and its mean is nearly twelve seconds. Applying the paper’s own 500 ms cut-off shows why: roughly one WhisperX word in six on Buckeye was more than half a second from where it was said, which the authors attribute to alignment drifting over long inputs. An aligner that is close on most words and lost on a sixth of them is a different instrument from one that is always roughly right, and an average cannot tell you which one you have.

The second detail is what was left out. The comparison was computed only on words the neural systems had recognised correctly. In conversational telephone speech, the words recognisers get wrong more often include turn-initial words, discourse markers and words just before a disfluency: the first copy of a repetition, or the word before a cut-off.[11] Those are the words at the edges of turns and hesitations, which is where timing matters most. The published numbers are the flattering version.

Perit runs its own benchmark on this question. PERIT-ALIGN scores 14 systems on the same real-world audio against word timestamps marked by people, with references produced in two independent passes and a senior reviewer resolving disagreements. A word only counts if both its start and its end land inside the tolerance.[12][13] What it shows most clearly is what happens when the window tightens from “roughly right” to the scale at which turn-taking operates.

PERIT-ALIGN word timing within 100 ms and 50 ms of human references. Median coverage is 60.2% and 25.6%, respectively.
Figure 3. Share of reference words with both start and end within 100 ms (blue) and 50 ms (orange) of the human reference, from the PERIT-ALIGN leaderboard.[12] *Perit’s model is a forced aligner given the reference transcript. Every other system transcribes and times the audio in one pass, so the top row is a reference point, not a like-for-like entry.Open full-size chart
Show the numbers
System Within 100 ms Within 50 ms Boundary MAE (ms)
Perit Golden Data fine-tuned model* 97.85% 83.90% 23.8
Fish Audio transcribe-1 92.33% 67.69% 24.5
Mistral voxtral-mini-transcribe 85.04% 43.96% 45.6
xAI grok-stt-1.0 78.04% 38.89% 52.7
Qwen3-ASR 0.6B 77.41% 59.52% 68.2
Qwen3-ASR 1.7B 76.60% 58.75% 70.0
AssemblyAI universal-3-5-pro 63.45% 22.77% 68.6
Deepgram nova-3 60.19% 13.09% 76.6
OpenAI whisper-large-v3 59.89% 25.79% 90.1
Microsoft mai-transcribe-2 59.85% 25.60% 66.6
OpenAI whisper-large-v3-turbo 46.95% 19.63% 97.8
OpenAI whisper-1 45.86% 15.73% 94.7
NVIDIA parakeet-tdt-0.6b-v3 30.96% 4.16% 116.8
NVIDIA nemotron-3.5-asr-streaming 5.77% 0.65% 258.0

Leave out Perit’s own entry, for reasons explained below, and the median of the other 13 systems times 60% of words correctly at 100 ms. At 50 ms, about a quarter of the 208 ms average gap before an answer, the median falls to 26%. Whisper large-v3 goes from 59.9% to 25.8%, and Deepgram’s nova-3 from 60.2% to 13.1%. The order changes too. The two small Qwen3-ASR models sit fifth and sixth at 100 ms, but at 50 ms only the top two systems beat them. Timestamps that look usable at one tolerance are a different instrument at another, and turn-taking lives at the narrow end.

The top of the board is not a fair fight, and that is what makes it instructive. Perit’s model is handed the correct transcript and only has to place it in time; every other system must work out the words and their timing at once. Knowing exactly which words were spoken, fillers and false starts included, is most of what makes timing possible. And every transcribing system on the board covers about 95% to 98% of the reference words, which leaves a few percent that were spoken and never timed at all. Which words those are is exactly what an average cannot tell you.

The transcript decides the timestamps

An aligner can only place the words it is given. That sounds like a technicality until you ask who decides which words it is given. Somewhere upstream, someone chose a transcription convention: verbatim, which keeps fillers, repetitions and false starts, or clean, which writes down what the speaker meant. The choice is usually made for readability, by someone who never thinks about timing, and it silently decides what the timing data can contain.

Its size is easy to underestimate. Wagner, Zusag and Thallinger compared verbatim and clean references for the same AMI meeting audio and found they differ by 12.1% WER with no content words changed. By their estimate, about 60% of the WER reported on AMI reflects style mismatch rather than lost content.[14] Every one of those stylistic differences is a stretch of sound that one version of the transcript accounts for and the other does not.

When the transcript leaves a sound out, the aligner still has to account for its time. Kouzelis and colleagues showed that aligners degrade sharply when disfluencies (in their tests, repetitions and stuttering) are in the audio but missing from the text, and that MFA degraded most of the systems they compared.[15] That is no coincidence. MFA is the most precise aligner in the benchmarks above because it leans hard on its transcript and its dictionary, and it is brittle for the same reason. Precision and brittleness come from the same place.

Now set this beside what Clark and Fox Tree found. “Um” is a forecast of a delay, and a much better one than “uh”. A clean transcript deletes the forecast and keeps the pause. For a turn-taking model that is the worst trade on offer: the cue that explains the label is removed and the label stays. The model is shown a long silence after “so”, told the speaker kept the turn, and given nothing to explain why. All it can learn is a superstition, that long silences are sometimes holds for no visible reason, and it will apply that superstition to silences that are not holds at all.

Where machine timing breaks on real calls

Fillers are the clearest case, but not the only one. Support calls add conditions that the standard benchmarks do not contain. Buckeye is forty speakers from central Ohio, talking in interviews;[16] MFA’s 2026 evaluation adds Japanese and Korean corpora of spontaneous speech.[8] They are careful, valuable datasets. None of them is a support call.

What happens on the call What the aligner does What the turn-taking data inherits
Fillers: “um”, “uh”, “haan”, “eto” Clean transcripts drop them; their time is smeared into a neighbouring word or booked as silence Holds recorded as bare silence, with the hold cue gone
Cut-off words: “d-“, “Thurs-”, “my acc- my card” No dictionary entry, so they are forced onto the nearest word or marked unknown Restarts look fluent and the repair point vanishes
Backchannels over the other speaker: “mm-hm” while the caller talks On one mixed channel, only one speaker’s words can be placed cleanly at a time Overlap reads as one continuous turn, or as a gap that never happened
Phone-line audio Narrowband phone audio cuts off above about 3.4 kHz, where much of the energy of fricatives like /s/ and /f/ lies Word edges on those sounds blur in ways wideband benchmark audio never tests
Laughter, breaths, a television in the room Unlabelled sound is absorbed into neighbouring words Words and pauses stretched by sounds nobody transcribed
Code-switching: Hinglish, Spanglish, Taglish Monolingual dictionaries and models misplace the other language’s words Timing is likely to be worst around the switch points

How a 40 ms timing error becomes an interruption

Silence is a weak signal that a turn has ended. Heldner and Edlund, working through three corpora of real conversation, found that pauses inside one speaker’s turn are generally longer than the gaps between speakers, and that overlaps make up about 40% of transitions from one speaker to the next.[17] The durations of “I’m still going” and “your turn” overlap heavily, so any rule that decides on silence has a thin margin to work in. Silence rules are still common. OpenAI’s Realtime API defaults to a silence-based mode, and the configuration it documents ends the turn after 500 ms of quiet.[18] Newer systems layer semantic cues on top, as our post on why “yeah” is the hardest word in voice AI describes. But every one of them is tuned, evaluated or trained on timing data, and that is where the ruler’s error gets in. It gets in from two directions.

Two paths from timing error to voice-agent mistakes: swallowed pauses shorten thresholds, while missing fillers remove the cue to keep listening.
Figure 4. The two routes by which timestamp error reaches behaviour. The first link in each chain is documented for common tools; the later links are the mechanism this post argues for, and how much each contributes depends on the pipeline.Open full-size chart

When pauses are swallowed, measured pauses come out shorter than real ones. A threshold tuned on that data to “rarely cut in” is tuned for a world where people pause less than they do, so it lands too short, and in production the agent cuts in on pauses its training data never measured properly.

When pauses are invented, the error is quieter and worse. A simulated caller built from those statistics goes silent where real callers say “um”, so the agent trained against it never meets the filler as a signal. It learns to wait out silence. Then a real caller says “yeah, so, um…”, the silence timer starts after the “um”, and the agent does exactly what it was trained to do.

The two errors push in opposite directions, which makes them dangerous together: a dataset can report a perfectly plausible average pause while both halves of it are wrong. Absolute error hides this as well. An aligner whose boundaries scatter randomly by 40 ms widens the pause distribution. One that systematically stretches words, ending them late or starting the next one early, shortens every pause it measures, and that shift moves every threshold tuned on top of it. The question to ask of timing data is not only how far off it is, but which way.

Questions to ask about any word-timestamp dataset

Whether you are buying timestamped conversational data, training a turn-taking model or tuning an endpointer on your own call logs, these questions separate timing you can train on from timing that will quietly train the wrong thing.

Question Why it matters A good answer
Were the timestamps set by people or by a machine? Machine timing inherits the aligner’s biases on exactly the speech you care about People, against a verbatim transcript, checked in a second pass
Which transcript were they aligned against? Clean transcripts delete fillers and false starts, and the timing inherits the holes Verbatim: fillers, repeats, false starts and cut-off words kept
Is accuracy reported at 50 ms, not just 100 ms? Most systems look usable at 100 ms and not at 50 A tolerance curve from 20 to 200 ms
Is accuracy broken out by word type? An average dominated by fluent words hides fillers and cut-offs Separate figures for fillers, cut-offs, backchannels and overlap
Is the error signed, and are the tails reported? Direction shifts thresholds, and a few lost words wreck a mean Signed bias per word type, for starts and ends separately, plus the median and a high percentile
Which words got no timestamp at all? The words that go untimed are rarely a random sample Coverage reported, with the missing words broken down by type
One channel per speaker, or a mixed recording? Overlapping speech cannot be timed properly on a single channel Separate channels wherever the call allows it

The word nobody wrote down

None of this is an argument against any particular tool. MFA is excellent at what it was built for, and so are the transcription systems on the leaderboard. The failure lives in the joins: a transcription convention chosen for readability, an aligner validated on speech unlike the calls it will time, a pause computed as a leftover, and a threshold tuned on the result.

That chain is what Perit’s annotation bench is built to break. Transcripts are verbatim, with fillers, repeats and false starts kept. Every word gets a start and an end, placed by people: the bench has 400 aligners and more than 1,000 hours aligned so far, in an annotation operation that works across nine locales. Benchmark references are written twice, independently, with a senior reviewer adjudicating every disagreement.[19][13] PERIT-ALIGN exists so the timing tools can be held to the same standard, in public.

In “yeah, so, um, I was thinking…”, the filler tells the listener that the pause is not necessarily the end of the turn. Remove it from the transcript and the timing data loses that cue. Keeping the word is part of keeping the interaction usable for training.

  1. H. H. Clark and J. E. Fox Tree, “Using uh and um in spontaneous speaking”, Cognition 84(1), 2002. Source of the 61% and 29% figures.
  2. M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, M. Sonderegger, “Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi”, Interspeech 2017. Frame settings from the MFA documentation.
  3. M. Bain, J. Huh, T. Han, A. Zisserman, “WhisperX: Time-Accurate Speech Transcription of Long-Form Audio”, Interspeech 2023.
  4. V. Pratap et al., “Scaling Speech Technology to 1,000+ Languages” (MMS), 2023.
  5. L. Wagner, B. Thallinger, M. Zusag, “CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions”, Interspeech 2024.
  6. A. Zeyer, R. Schlüter, H. Ney, “Why does CTC result in peaky behavior?”, 2021.
  7. R. Huang et al., “Less Peaky and More Accurate CTC Forced Alignment by Label Priors”, 2024. The 21 ms and 58 ms figures are for standard CTC on Buckeye; 82 ms is the ground-truth mean phone duration.
  8. M. McAuliffe, K. Gunter, M. Wagner, M. Sonderegger, “Montreal Forced Aligner and the state of speech-to-text alignment in 2026”, 2026. Word-boundary mean errors on Buckeye (MFA 3.0 ARPA model).
  9. T. Stivers et al., “Universals and cultural variation in turn-taking in conversation”, PNAS 106(26), 2009. Mean response offset of 208 ms across ten languages.
  10. R. Rousso, E. Cohen, J. Keshet, E. Chodroff, “Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment”, Interspeech 2024. Word-level results, Tables 1–3; the one-in-six figure is derived from the paper’s 500 ms-thresholded Buckeye results.
  11. S. Goldwater, D. Jurafsky, C. D. Manning, “Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates”, Speech Communication, 2010.
  12. Perit, PERIT-ALIGN word-alignment leaderboard. All 14 systems’ figures as published there; medians in this post exclude Perit’s own entry.
  13. Perit, Benchmark method: two-pass references, adjudication and held-out audio.
  14. L. Wagner, M. Zusag, B. Thallinger, “Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing”, 2026.
  15. T. Kouzelis, G. Paraskevopoulos, A. Katsamanis, V. Katsouros, “Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling”, 2023.
  16. M. A. Pitt et al., “The Buckeye corpus of conversational speech: labeling conventions and a test of transcriber reliability”, Speech Communication, 2005.
  17. M. Heldner and J. Edlund, “Pauses, gaps and overlaps in conversations”, Journal of Phonetics 38(4), 2010.
  18. OpenAI, “Voice activity detection (VAD)”, Realtime API documentation.
  19. Perit, Data annotation: verbatim conventions, word alignment, bench size and QA.

The three timelines at the top are illustrative, not measurements of a specific recording. Benchmark figures are as reported by the cited sources. Pause-error estimates assume unbiased, independent, roughly normal boundary errors and are indicative only. Charts are hand-built SVG from the published numbers.

 

Was this article helpful? No ratings yet