The scored set is never published. A 60-clip development split per sector is open so anyone can calibrate against the same conventions without seeing the test.
Every reference transcript and every rubric score is produced twice, independently. A senior reviewer resolves disagreements and the resolution is recorded with the item.
Three runs per clip on transcription, eight calls per persona on conversation. Intervals are bootstrapped at 95% and printed next to every score.
Someone who does that job takes the same set, blind and timed, at market rate. It is a working standard, not a ceiling — the human line carries its own error bar.
Each criterion is written by a practitioner and then rewritten until two strangers grading the same call land on the same score.
Annotator, reviewer and adjudicator are attached to every item, so a disputed score can be traced to the people who set it.
Scored on word error rate — errors per hundred words against a double-passed human transcript, adjudicated by a third reviewer on disagreement.
Scored on task success — the caller's stated goal was met, verified against a rubric written by an operator from that sector.
No scores are published yet, and none will be until the held-out sets close. The protocol is here first on purpose: the format should be argued with before the numbers arrive, not after they flatter someone.