[B]Research · PERIT-BENCH

Frontier benchmarks for work that needs a license.

PERIT-BENCH grades frontier models on tasks a credentialed professional is paid to do — authored by practitioners, graded against the standard a supervisor would actually apply, and scored beside a human control who did the same work.

Benchmarks that frontier models ace on day one measure the exam, not the job. We grade against the ceiling of professional work, publish the number even when it is low, and put a credentialed human in the same table.

4
domain suites
1,030
held-out tasks
8
runs per task
95%
intervals on every score
The suites
PREVIEW · scores are illustrative and models are anonymized by tier. Measured, citable results publish with the suite.
[01]PERIT-LEGAL · Documents · long-context

Grade models on the work a partner would sign

PERIT-LEGAL measures whether frontier models can produce legal work product a supervising partner would put their name on — across six task families drawn from live matters and scored against bar-admitted practitioners doing the same work.

All-pass @1 · top 3
Frontier reasoning · lab A41.2
Frontier reasoning · lab B39.6
Frontier reasoning · lab C36.4
Credentialed human baseline74.8
View full leaderboard →
[02]PERIT-CLINICAL · Text · multi-turn

Grade models on the call a clinician has to defend

PERIT-CLINICAL measures whether frontier models can carry a clinical encounter to a defensible conclusion — scored against physician-authored rubrics across five care settings, with a board-certified control answering the same cases.

Rubric score · top 3
Frontier reasoning · lab A58.7
Frontier reasoning · lab B56.9
Frontier reasoning · lab C53.4
Board-certified control78.2
View full leaderboard →
[03]PERIT-AUDIO · Audio · long-form

Transcription is solved. The work is not

PERIT-AUDIO measures what a professional does with what they heard — clinical dictation into a usable note, a deposition worked for admissions, a call graded for the disclosure that was or was not made. Scored against practitioners handling the same recordings.

All-pass @1 · top 3
Frontier audio-native · lab A34.6
Frontier audio-native · lab B32.1
Frontier audio-native · lab C28.4
Credentialed human baseline69.5
View full leaderboard →
[04]PERIT-VISION · Image · video

Grade models on what an expert eye catches

PERIT-VISION measures whether frontier models can review an image or a video the way a credentialed reviewer does — finding the one thing that should not be there, and stating what follows from it. Scored beside practitioners reviewing the same material.

All-pass @1 · top 3
Frontier multimodal · lab A37.9
Frontier multimodal · lab B35.2
Frontier multimodal · lab C31.8
Credentialed human baseline72.3
View full leaderboard →
How we run them

Hard on purpose

A suite frontier models ace on release measures the exam, not the job. Every task here was authored to sit at the ceiling of professional work, and we ship a hard split that stays unsaturated by design.

A human in the same table

Every suite is scored beside a credentialed practitioner doing the same held-out work under ordinary conditions. Model-versus-model tells you who is ahead. Model-versus-professional tells you whether it can be deployed.

Held out, and named

Held-out sets are never published, and every task carries its author, reviewer, and credential. When you cite the score, you can cite exactly how it was produced — including what it does not measure.

Evaluate your model

Add your model to a leaderboard.

We run any suite against a frontier model on request and return the loss analysis with the score. A model is listed from its first run against the prompts, not its best — that rule is the only thing keeping the number worth publishing.