Frontier benchmarks for work that needs a license.
PERIT-BENCH grades frontier models on tasks a credentialed professional is paid to do — authored by practitioners, graded against the standard a supervisor would actually apply, and scored beside a human control who did the same work.
Benchmarks that frontier models ace on day one measure the exam, not the job. We grade against the ceiling of professional work, publish the number even when it is low, and put a credentialed human in the same table.
Grade models on the work a partner would sign
PERIT-LEGAL measures whether frontier models can produce legal work product a supervising partner would put their name on — across six task families drawn from live matters and scored against bar-admitted practitioners doing the same work.
Grade models on the call a clinician has to defend
PERIT-CLINICAL measures whether frontier models can carry a clinical encounter to a defensible conclusion — scored against physician-authored rubrics across five care settings, with a board-certified control answering the same cases.
Transcription is solved. The work is not
PERIT-AUDIO measures what a professional does with what they heard — clinical dictation into a usable note, a deposition worked for admissions, a call graded for the disclosure that was or was not made. Scored against practitioners handling the same recordings.
Grade models on what an expert eye catches
PERIT-VISION measures whether frontier models can review an image or a video the way a credentialed reviewer does — finding the one thing that should not be there, and stating what follows from it. Scored beside practitioners reviewing the same material.
Hard on purpose
A suite frontier models ace on release measures the exam, not the job. Every task here was authored to sit at the ceiling of professional work, and we ship a hard split that stays unsaturated by design.
A human in the same table
Every suite is scored beside a credentialed practitioner doing the same held-out work under ordinary conditions. Model-versus-model tells you who is ahead. Model-versus-professional tells you whether it can be deployed.
Held out, and named
Held-out sets are never published, and every task carries its author, reviewer, and credential. When you cite the score, you can cite exactly how it was produced — including what it does not measure.
Add your model to a leaderboard.
We run any suite against a frontier model on request and return the loss analysis with the score. A model is listed from its first run against the prompts, not its best — that rule is the only thing keeping the number worth publishing.