[01]Benchmark · PERIT-LEGAL

Grade models on the work a partner would sign.

PERIT-LEGAL measures whether frontier models can produce legal work product a supervising partner would put their name on — across six task families drawn from live matters and scored against bar-admitted practitioners doing the same work.

6
task families
240
held-out tasks
9.4
avg rubric criteria
68%
tasks never all-passed
41.2
best all-pass score

The PERIT-LEGAL leaderboard

Legal benchmarks mostly test recall: pick the right holding from four options, restate the black-letter rule. Firms do not buy recall. They buy work product that survives partner review — and there, a memo that catches eight of ten risks is not eighty percent useful. It is a redraft.

01

Bar-admitted practitioners pulled tasks from matters they had actually run — redlines, chronologies, diligence memos — and reduced each to a form that publishes without touching privilege.

02

Each task carries an all-pass rubric: every element a reviewing partner would require, written by the author and adjudicated by a second bar-admitted reviewer before the task is admitted.

03

A blind panel of practicing lawyers completed the same held-out tasks under ordinary client-work conditions, producing the human control the models are scored beside.

240 tasks are held out and never published. A 60-task development split, built by the same authors through the same pipeline, publishes with the suite alongside the grading harness and the full rubric schema.

PERIT-LEGAL · leaderboardscore / 100

Share of tasks where every rubric criterion passes on a single attempt. No partial credit — this is how a partner reviews a draft.

Credentialed human baselinebar-admitted · 8+ yrs74.8±3.4
01Frontier reasoning · lab Ahigh41.2±2.9
01Frontier reasoning · lab Bmax39.6±3.0
01Frontier reasoning · lab Chigh36.4±2.8
03Frontier general · lab A33.1±2.9
Rank· one plus the number of models whose lower bound clears this model's upper bound. Models inside each other's intervals share a rank. The human control is not ranked.
PREVIEW · scores are illustrative and models are anonymized by tier. Measured, citable results publish with the suite.
Key takeaways
  1. 01

    No model clears the bar for unsupervised legal work.

    The leading system all-passes 41.2% of tasks against 74.8% for a bar-admitted practitioner. Every model in the table would require a lawyer to read the output before it left the building — which is the current state of practice, now with a number attached to it.

  2. 02

    Capability is not consistency.

    The same leading system passes 84.7% of individual criteria but all-passes only 41.2% of tasks, and holds that across eight runs on just 12.4%. Models are close on most of the work and unreliable about which part they will miss — the failure mode least compatible with legal review.

  3. 03

    Authority, not analysis, is where the drafts break.

    Roughly a third of graded failures are citation and authority defects — a real rule applied to the wrong jurisdiction, or a proposition supported by a case that does not stand for it. The reasoning is frequently sound and the support underneath it is not.

  4. 04

    The human control is not the ceiling either.

    Practitioners all-pass 74.8% of tasks, not 100%. We publish that because a benchmark that scores humans at perfect is measuring its own rubric rather than the work. The headroom above both columns is the part worth training against.

What the suite covers

Every task authored by someone credentialed to do it.

M&A diligence & disclosure

Read the data-room set, identify what a buyer must be told, and quantify the exposure. Graded on whether every material item surfaces, not on how well the memo reads.

Authored by
JD, bar-admitted · corporate
Best score · 52 tasks
38.4

Contract redlining

Mark up a counterparty draft against a negotiating position. Every accepted, rejected, and countered term is scored, including the ones a model leaves silently untouched.

Authored by
JD, bar-admitted · 6+ yrs
Best score · 48 tasks
34.1

Litigation chronology

Build a dated chronology from an unordered document set and flag the entries that carry legal significance. Ordering errors and omitted events are graded separately.

Authored by
JD, litigation practice
Best score · 44 tasks
45.7

Regulatory research

Answer a live compliance question across overlapping federal and state regimes, with authority. Multi-jurisdiction questions are scored as a distinct hard split.

Authored by
JD + regulatory speciality
Best score · 46 tasks
40.3

Transcript & deposition analysis

Work a hearing or deposition transcript for admissions, contradictions, and follow-up that was not asked. Graded against what an examining attorney flagged.

Authored by
JD, trial experience
Best score · 28 tasks
43.8

Client-ready drafting

Produce the memo or letter that actually goes out, at the register the recipient expects. Scored on substance first, then on whether it could be sent without a rewrite.

Authored by
JD, bar-admitted · 8+ yrs
Best score · 22 tasks
31.6
Sample task

One task, and the rubric behind it.

sample_task · M&A · indemnitymedium
Task prompt

Buyer's counsel has proposed an indemnity cap of 10% of enterprise value with a 12-month survival period, carved out for fraud and for breach of fundamental representations. The seller is a portfolio company with a disclosed environmental matter in the data room. Advise the seller on cap exposure, identify which carve-outs are uncapped as drafted, and state the governing-law consequence of the survival period as written.

Environment
draft SPA (§8)disclosure schedulesdata-room indexprior-round SPA
Grading criteria
  • Cites governing law correctly
  • Flags the fraud carve-out as uncapped as drafted
  • Quantifies cap exposure against enterprise value
  • Identifies the environmental matter as a disclosed exception
+ 6 more grading criteria

Every model in the table gets the fraud carve-out. Most miss that the disclosed environmental matter changes the exposure calculation entirely — the fact is in the schedules, not the clause, and nothing in the prompt points at it.

Loss analysis

Where the graded failures land.

38%Missed a required element

The analysis is correct as far as it goes and simply omits something the rubric requires — most often a fact that lives in an attachment rather than the operative clause.

31%Unsupported authority

A citation that exists but does not stand for the proposition, or a rule stated accurately for the wrong jurisdiction. Fabricated citations are a small minority of this category.

19%Right rule, wrong application

The governing standard is identified correctly and then applied to the wrong facts, or applied without the exception the disclosure set establishes.

12%Not auditable

The conclusion is defensible but the reasoning cannot be traced back to a source in the record, so a reviewing lawyer cannot sign off without redoing the work.

Methodology
Grading
All-pass rubric adjudicated by a second bar-admitted reviewer; LM judge validated against expert grades before use.
Runs
8 independent runs per task per model; mean reported, all-pass ^8 reported separately.
Error bars
95% intervals from 10,000 bootstrap resamples over tasks.
Human control
practicing lawyers completing the same held-out tasks under ordinary client-work conditions, blind to the study.
Contamination
Held-out set never published; authored post-cutoff from matters that were never public.
Dataset
240 held-out tasks; 60-task development split published under CC-BY with the harness.

What this suite does not measure

Published up front, because a benchmark that names its own limits is the only kind whose number means anything.

  • US law only. No coverage of EU, UK, or cross-border regimes beyond the multi-jurisdiction split.
  • No tax, immigration, or IP prosecution work in v1.
  • Tasks are document-bounded — nothing here measures client counselling, negotiation, or oral advocacy.
  • The human control is a panel of practitioners, not the specific partner who would review a given matter.
Evaluate your model

Run PERIT-LEGAL against your model.

We run the held-out set for any frontier model on request, and return the loss analysis alongside the score. To keep the leaderboard honest, a model is listed from the first run against the prompts — not the best one.