PERIT-STT

Transcription is solved on clean audio. Nobody calls from clean audio.

PERIT-STT scores models on real support recordings — mobile lines, crosstalk, accents and the domain vocabulary that decides whether a transcript is usable downstream.

Word error rate (%, lower is better)

Errors per hundred words against a double-passed human transcript, adjudicated by a third reviewer on disagreement.

Entity accuracy (%)

Exact match on the tokens that carry the transaction — account numbers, amounts, dates, drug names, SKUs.

Diarization error (%, lower is better)

Speaker attribution error across overlapping speech.

Overall leaderboard

Each model’s mean across all 10 sectors. Open a sector below for the run behind it.

Illustrative figures. Public runs begin when the held-out sets close.
ModelWord error rate (%)Entity accuracyDiarization error
Human control · sector transcriptionistHuman control
2.6 ± 1.6
98.9%2.9%
Frontier ASR · lab AFrontier
7.6 ± 1.3
91.7%9.5%
Specialist vendor ASRSpecialist
9.3 ± 1.5
88.6%8.8%
Frontier ASR · lab BFrontier
9.7 ± 1.6
89.5%10.8%
Frontier multimodal · lab CFrontier
11.5 ± 1.4
85%12.2%
Open weights · 8BOpen weights
19.0 ± 1.5
75%16.8%
Sectors

The same call, ten different rooms.

A refund request in retail and a claim in insurance share a shape. They do not share vocabulary, background noise, or what counts as a wrong answer.

Method

How a run works.

  1. 01Audio is held out and never published. A 60-clip development split per sector is open for calibration.
  2. 02Reference transcripts are written by two annotators independently; a senior reviewer adjudicates every disagreement.
  3. 03Each clip is run 3 times per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a working transcriptionist from the same sector, timed and paid at market rate.