[04]Benchmark · PERIT-VISION

Grade models on what an expert eye catches.

PERIT-VISION measures whether frontier models can review an image or a video the way a credentialed reviewer does — finding the one thing that should not be there, and stating what follows from it. Scored beside practitioners reviewing the same material.

4
review domains
260
held-out items
11.6
avg rubric criteria
64%
tasks never all-passed
37.9
best all-pass score

The PERIT-VISION leaderboard

Vision benchmarks ask whether a model can name what is in the frame. Professionals are not paid to name things. They are paid to notice the one detail that should not be there, to say what it implies, and to be right about it when the finding is subtle and the consequence is expensive.

01

Credentialed reviewers contributed material from their own caseload — studies, inspection footage, claim evidence — paired with the finding they recorded at the time.

02

Each item carries a rubric separating detection from interpretation, so a model that sees the finding and misreads it scores differently from one that never saw it.

03

A second credentialed reviewer worked the same material blind, establishing the human control and the miss rate a professional review actually carries.

260 items are held out, roughly split between still imagery and video. A 60-item development split publishes with the harness, the rubric schema, and the finding taxonomy.

PERIT-VISION · leaderboardscore / 100

Share of items where detection and interpretation both satisfy every rubric criterion. Seeing the finding is not enough to pass.

Credentialed human baselinesecond reviewer · blind72.3±3.3
01Frontier multimodal · lab Ahigh37.9±2.9
01Frontier multimodal · lab Bmax35.2±3.0
02Frontier multimodal · lab Chigh31.8±2.9
03Frontier multimodal · lab Astandard28.4±3.0
Rank· one plus the number of models whose lower bound clears this model's upper bound. Models inside each other's intervals share a rank. The human control is not ranked.
PREVIEW · scores are illustrative and models are anonymized by tier. Measured, citable results publish with the suite.
Key takeaways
  1. 01

    Models describe the frame well and review it poorly.

    Partial credit clusters in the high seventies while all-pass tops out at 37.9. Models reliably enumerate what is present and unreliably identify which of it matters — the distinction that separates captioning from review.

  2. 02

    Roughly a quarter of critical findings are never seen.

    The leading system recalls 76.2% of findings a reviewer classified as critical, against 94.6% for a blind second human. In every domain here, a missed critical finding is the failure with the largest downstream cost.

  3. 03

    Video is materially harder than stills, and not because of length.

    All-pass on video items runs about twelve points below matched still imagery. Failures concentrate on findings that only exist across frames — a change in state, an intermittent fault — which no single frame contains.

  4. 04

    Detection and interpretation fail independently.

    Grading them separately shows that about a third of failures are findings the model saw and then misread. Collapsing both into one score would hide the more tractable of the two problems.

What the suite covers

Every task authored by someone credentialed to do it.

Diagnostic imaging review

Studies read for findings and their implications, graded against the reviewing radiologist's report. Detection and interpretation are scored as separate criteria.

Authored by
MD, radiology board-certified
Best score · 76 tasks
33.4

Industrial & site inspection

Inspection footage reviewed for defects, code violations, and safety conditions. Intermittent and cross-frame faults are the hard cases and are labeled as such.

Authored by
Licensed inspector / PE
Best score · 68 tasks
36.8

Damage adjudication

Claim evidence assessed for cause, extent, and whether the documented damage is consistent with the reported event. Graded against the adjuster's determination.

Authored by
Licensed adjuster, 5+ yrs
Best score · 64 tasks
44.2

Document & evidence forensics

Scanned documents and evidence imagery examined for alteration, inconsistency, and provenance signals. Scored on what was found and on what was correctly ruled out.

Authored by
Certified forensic examiner
Best score · 52 tasks
35.1
Sample task

One task, and the rubric behind it.

sample_task · Inspection · videohard
Task prompt

A 4-minute walkthrough of a commercial rooftop mechanical installation. Identify every condition that would appear on an inspection report, classify each by severity, and state which require immediate remediation. Note any condition that is intermittent across the footage.

Environment
4 min 4K videoequipment scheduleprior inspection reportapplicable code section
Grading criteria
  • Identifies the corroded mounting bracket
  • Identifies the obstructed condensate drain
  • Detects the intermittent vibration visible only across frames 0:47–1:12
  • Classifies the drain obstruction as requiring immediate remediation
+ 8 more grading criteria

The vibration exists in no single frame — it is only a finding as a change across a sequence. Models sampling frames independently cannot see it in principle, and every system in the table misses it while correctly listing the static defects around it.

Loss analysis

Where the graded failures land.

35%Finding never detected

The condition is present and visible and does not appear in the output at all. Concentrated in subtle findings and in the critical class, where it matters most.

31%Seen but misinterpreted

The model identifies the finding and draws the wrong conclusion — wrong severity, wrong cause, or wrong remediation. More tractable than a miss, and hidden if the two are scored together.

21%Cross-frame blindness

A finding that only exists as a change over time, missed by systems reasoning over sampled frames. Effectively a video-only failure class.

13%Confident false positive

A condition reported that is not present, at a severity that would trigger action. Costly in adjudication and inspection, where a false finding has its own downstream expense.

Methodology
Grading
Detection and interpretation scored as separate rubric criteria; adjudicated by a second credentialed reviewer.
Video handling
Full-resolution video supplied to models that accept it; frame-sampled fallback recorded as a config field, never mixed into one column.
Runs
8 independent runs per item; all-pass, partial credit, and critical-finding recall reported separately.
Human control
A blind second credentialed reviewer working the same material, so the control carries a realistic human miss rate.
Error bars
95% intervals from 10,000 bootstrap resamples over items.
Dataset
260 held-out items, roughly half video; 60-item development split published with the harness.

What this suite does not measure

Published up front, because a benchmark that names its own limits is the only kind whose number means anything.

  • Consented and de-identified material only, which under-represents the rarest findings in every domain.
  • No real-time or on-device setting — items are reviewed offline with full resolution available.
  • Diagnostic imaging covers four modalities in v1; it is not a substitute for a clinical validation study.
  • The human control is a second reviewer, not a multi-reader consensus panel.
Evaluate your model

Run PERIT-VISION against your model.

We run the held-out set for any frontier model on request, and return the loss analysis alongside the score. To keep the leaderboard honest, a model is listed from the first run against the prompts — not the best one.