Transcription is solved. The work is not.
PERIT-AUDIO measures what a professional does with what they heard — clinical dictation into a usable note, a deposition worked for admissions, a call graded for the disclosure that was or was not made. Scored against practitioners handling the same recordings.
The PERIT-AUDIO leaderboard
Audio benchmarks report word error rate. Nobody is paid for word error rate. A clinician is paid for the note that results, a compliance officer for the disclosure that was missed, a deposition summarizer for the one admission that matters. Word error rate treats every word as equally important; professional work never does.
Practitioners contributed recordings from their own working context — consented, de-identified, and released with the downstream artifact they produced from each one.
Each recording carries a rubric written against that artifact: what the note had to contain, which admissions the transcript had to surface, which disclosure the call was required to make.
The same recordings were worked by a second credentialed practitioner under normal conditions, giving both the human control and the agreement floor.
210 recordings are held out, spanning accented and code-switched speech, overlapping speakers, and realistic room noise. A 50-clip development split publishes with the harness and the full rubric schema.
Share of recordings where the resulting work product satisfies every rubric criterion. Graded on the artifact, never on the transcript.
- 01
Hearing the words is no longer the bottleneck.
Median word error rate across the held-out set is 2.8%. On the same recordings, 71% of tasks are never all-passed by any model on any run. The gap between those two numbers is the entire argument for grading audio on work product.
- 02
Audio-native models beat cascaded pipelines, and not for the reason expected.
Audio-native systems lead cascaded ASR-plus-reasoning setups by roughly eight points on all-pass despite comparable transcription quality. The advantage shows up in speaker attribution and in hesitation, correction, and emphasis — signal a transcript discards before the reasoning model ever sees it.
- 03
Attribution errors are silent.
A misattributed statement produces a fluent, confident artifact that is wrong about who is liable. Unlike a transcription error, nothing downstream flags it — which is why attribution is reported as its own column rather than folded into the score.
- 04
Accent and code-switching cost more than noise.
Recordings with accented or code-switched speech lose roughly eleven points of all-pass, against four for the noisiest room conditions in the set. The failure is comprehension of the content, not capture of the signal.
Every task authored by someone credentialed to do it.
Clinical dictation → note
A recorded encounter turned into a note another clinician can act on. Graded on what the note contains and on whether anything was asserted the recording did not support.
Deposition & hearing audio
Multi-party legal audio worked for admissions, contradictions, and speaker attribution. Scored against what the examining attorney actually flagged.
Regulated-call compliance
Recorded customer calls graded on whether every required disclosure was made, in the required form. Binary criteria — a disclosure is made or it is not.
Multi-speaker intake
Accented, code-switched, and interpreter-mediated intake with overlapping speech. The hardest family in the suite and the most representative of frontline conditions.
One task, and the rubric behind it.
An 11-minute recorded advisory call between a licensed representative and a retail client, with two interruptions and a brief hold. Determine whether every required disclosure was made, in the required form and sequence. Where a disclosure was made partially or out of order, state which element was deficient and quote the passage it should have appeared in.
- Detects the fee disclosure was made
- Detects the risk disclosure was interrupted and never completed
- Attributes the partial disclosure to the representative, not the client
- Quotes the passage where the omitted element belonged
The risk disclosure begins, is cut off by the client interrupting, and is never resumed after the hold. A transcript reads as though it was delivered. Catching it requires tracking an obligation across a discontinuity — which is exactly what the compliance officer is paid for.
Where the graded failures land.
The model transcribes and summarizes the recording accurately while failing the professional requirement attached to it — a disclosure not completed, a finding not documented.
A rubric-relevant statement assigned to the wrong party. Concentrated in overlapping speech and interpreter-mediated exchanges, and invisible in the output.
An obligation opened before an interruption and never closed after it. The model treats the two sides of the gap as unrelated.
Hesitation, self-correction, and emphasis that changed the professional meaning, discarded before the reasoning step. Almost entirely a cascaded-pipeline failure.
What this suite does not measure
Published up front, because a benchmark that names its own limits is the only kind whose number means anything.
- English-dominant. Code-switched clips involve English plus one other language; fully non-English audio is out of scope in v1.
- No real-time or streaming setting — every task is scored on a complete recording.
- Word error rate is reported for context only and is not part of any score.
- Recording quality reflects professional capture conditions, not consumer far-field microphones.
Run PERIT-AUDIO against your model.
We run the held-out set for any frontier model on request, and return the loss analysis alongside the score. To keep the leaderboard honest, a model is listed from the first run against the prompts — not the best one.