PERIT-STT · Logistics & delivery

Driver and dispatch calls from vehicles, warehouses and doorsteps. Half of it is hands-free.

Audio comes from dispatch radio, driver hotlines, delivery exceptions. Scored on word error rate, with a working logistics & delivery operator as the control.

1,100
Clips in the target held-out set
79h
Audio, never published
6
Languages in the set
1
Working logistics & delivery operator as control

Leaderboard

Illustrative figures. Public runs begin when the held-out sets close.
ModelWord error rate (%)Entity accuracyDiarization error
Human control · sector transcriptionistHuman control
2.8 ± 0.8
99.1%2.5%
Frontier ASR · lab AFrontier
8.3 ± 1.2
92.5%10.9%
Frontier ASR · lab BFrontier
10.3 ± 1.6
89.3%10.9%
Specialist vendor ASRSpecialist
10.8 ± 1.1
90.7%8.3%
Frontier multimodal · lab CFrontier
12.8 ± 0.7
85.9%11.8%
Open weights · 8BOpen weights
17.3 ± 1.2
76.5%17.2%
Task families

What the set actually asks for.

  • Verbatim transcription
  • Entity capture (IDs, amounts, names)
  • Speaker diarization
  • Code-switching
  • Noisy & far-field audio
  • Domain vocabulary
Method & limits

How a run on this set is produced.

  1. 01Audio is held out and never published. A 60-clip development split per sector is open for calibration.
  2. 02Reference transcripts are written by two annotators independently; a senior reviewer adjudicates every disagreement.
  3. 03Each clip is run 3 times per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a working transcriptionist from the same sector, timed and paid at market rate.

Limits, stated up front: 6 languages only, single-channel audio, and no cross-sector transfer. The control is one operator per sector, so the human line carries its own error bar — it is a working standard, not a ceiling.