Method

How a run is produced, and what we will not do to make it look better.

Six rules apply to every benchmark on this site. They are the reason a number here is worth arguing with.

Held-out audio stays held out

The scored set is never published. A 60-clip development split per sector is open so anyone can calibrate against the same conventions without seeing the test.

Two passes, then adjudication

Every reference transcript and every rubric score is produced twice, independently. A senior reviewer resolves disagreements and the resolution is recorded with the item.

Runs, not a run

Three runs per clip on transcription, eight calls per persona on conversation. Intervals are bootstrapped at 95% and printed next to every score.

A human control on every board

Someone who does that job takes the same set, blind and timed, at market rate. It is a working standard, not a ceiling — the human line carries its own error bar.

Rubrics come from operators

Each criterion is written by a practitioner and then rewritten until two strangers grading the same call land on the same score.

Provenance travels with the artifact

Annotator, reviewer and adjudicator are attached to every item, so a disputed score can be traced to the people who set it.

What we publish
  • Leaderboards with confidence intervals, per sector and overall
  • The development split, and the conventions the references follow
  • The metric definitions, including what counts as an entity error
  • Notes on what failed and why, when a run is complete
What we do not
  • The held-out audio, or transcripts of it
  • Contributor and reviewer identities
  • Customer-commissioned sets, ever — those belong to the customer
  • Scores for a named model we have not run ourselves
PERIT-STT

Running speech-to-text.

Scored on word error rate — errors per hundred words against a double-passed human transcript, adjudicated by a third reviewer on disagreement.

  1. 01Audio is held out and never published. A 60-clip development split per sector is open for calibration.
  2. 02Reference transcripts are written by two annotators independently; a senior reviewer adjudicates every disagreement.
  3. 03Each clip is run 3 times per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a working transcriptionist from the same sector, timed and paid at market rate.
PERIT-S2S

Running speech-to-speech.

Scored on task success — the caller's stated goal was met, verified against a rubric written by an operator from that sector.

  1. 01Callers are scripted personas played by trained speakers, not synthesised — including the ones who talk over the agent.
  2. 02Every call is scored against the rubric by two reviewers; a third adjudicates splits.
  3. 038 calls per persona per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a licensed agent from that sector handling the same personas blind.

No scores are published yet, and none will be until the held-out sets close. The protocol is here first on purpose: the format should be argued with before the numbers arrive, not after they flatter someone.