[R]Research · PERIT-BENCH

Grade models on real professional work.

PERIT-BENCH measures what a model can actually do in a licensed professional's chair — not what it memorized. Honest, low on purpose, and citable down to the credential.

A preview leaderboard

Illustrative standings while the benchmark is in build. The point isn't the ranking — it's that the gap to the top is wide and honest. Real, citable results publish with the benchmark.

PREVIEW · scores below are illustrative. Real, citable results publish with the benchmark.
01Frontier model · lab A41.7
02Frontier model · lab B38.2
03Frontier model · lab C34.9
04Open weights · 70B27.1
05Open weights · 8B15.4
How we run it

Low scores on purpose

A benchmark frontier models ace on day one measures nothing. We grade against the ceiling of professional work, so the number has somewhere to go.

Held out, always

Tasks are authored and graded by research-grade experts and kept out of any training corpus. A leaked benchmark is a marketing asset, not a measurement.

Citable by construction

Every task carries its author, reviewer, and credential. When you publish the score, you can publish exactly how it was produced.