Grade models on what an expert eye catches.
PERIT-VISION measures whether frontier models can review an image or a video the way a credentialed reviewer does — finding the one thing that should not be there, and stating what follows from it. Scored beside practitioners reviewing the same material.
The PERIT-VISION leaderboard
Vision benchmarks ask whether a model can name what is in the frame. Professionals are not paid to name things. They are paid to notice the one detail that should not be there, to say what it implies, and to be right about it when the finding is subtle and the consequence is expensive.
Credentialed reviewers contributed material from their own caseload — studies, inspection footage, claim evidence — paired with the finding they recorded at the time.
Each item carries a rubric separating detection from interpretation, so a model that sees the finding and misreads it scores differently from one that never saw it.
A second credentialed reviewer worked the same material blind, establishing the human control and the miss rate a professional review actually carries.
260 items are held out, roughly split between still imagery and video. A 60-item development split publishes with the harness, the rubric schema, and the finding taxonomy.
Share of items where detection and interpretation both satisfy every rubric criterion. Seeing the finding is not enough to pass.
- 01
Models describe the frame well and review it poorly.
Partial credit clusters in the high seventies while all-pass tops out at 37.9. Models reliably enumerate what is present and unreliably identify which of it matters — the distinction that separates captioning from review.
- 02
Roughly a quarter of critical findings are never seen.
The leading system recalls 76.2% of findings a reviewer classified as critical, against 94.6% for a blind second human. In every domain here, a missed critical finding is the failure with the largest downstream cost.
- 03
Video is materially harder than stills, and not because of length.
All-pass on video items runs about twelve points below matched still imagery. Failures concentrate on findings that only exist across frames — a change in state, an intermittent fault — which no single frame contains.
- 04
Detection and interpretation fail independently.
Grading them separately shows that about a third of failures are findings the model saw and then misread. Collapsing both into one score would hide the more tractable of the two problems.
Every task authored by someone credentialed to do it.
Diagnostic imaging review
Studies read for findings and their implications, graded against the reviewing radiologist's report. Detection and interpretation are scored as separate criteria.
Industrial & site inspection
Inspection footage reviewed for defects, code violations, and safety conditions. Intermittent and cross-frame faults are the hard cases and are labeled as such.
Damage adjudication
Claim evidence assessed for cause, extent, and whether the documented damage is consistent with the reported event. Graded against the adjuster's determination.
Document & evidence forensics
Scanned documents and evidence imagery examined for alteration, inconsistency, and provenance signals. Scored on what was found and on what was correctly ruled out.
One task, and the rubric behind it.
A 4-minute walkthrough of a commercial rooftop mechanical installation. Identify every condition that would appear on an inspection report, classify each by severity, and state which require immediate remediation. Note any condition that is intermittent across the footage.
- Identifies the corroded mounting bracket
- Identifies the obstructed condensate drain
- Detects the intermittent vibration visible only across frames 0:47–1:12
- Classifies the drain obstruction as requiring immediate remediation
The vibration exists in no single frame — it is only a finding as a change across a sequence. Models sampling frames independently cannot see it in principle, and every system in the table misses it while correctly listing the static defects around it.
Where the graded failures land.
The condition is present and visible and does not appear in the output at all. Concentrated in subtle findings and in the critical class, where it matters most.
The model identifies the finding and draws the wrong conclusion — wrong severity, wrong cause, or wrong remediation. More tractable than a miss, and hidden if the two are scored together.
A finding that only exists as a change over time, missed by systems reasoning over sampled frames. Effectively a video-only failure class.
A condition reported that is not present, at a severity that would trigger action. Costly in adjudication and inspection, where a false finding has its own downstream expense.
What this suite does not measure
Published up front, because a benchmark that names its own limits is the only kind whose number means anything.
- Consented and de-identified material only, which under-represents the rarest findings in every domain.
- No real-time or on-device setting — items are reviewed offline with full resolution available.
- Diagnostic imaging covers four modalities in v1; it is not a substitute for a clinical validation study.
- The human control is a second reviewer, not a multi-reader consensus panel.
Run PERIT-VISION against your model.
We run the held-out set for any frontier model on request, and return the loss analysis alongside the score. To keep the leaderboard honest, a model is listed from the first run against the prompts — not the best one.