Grade models on the work a partner would sign.
PERIT-LEGAL measures whether frontier models can produce legal work product a supervising partner would put their name on — across six task families drawn from live matters and scored against bar-admitted practitioners doing the same work.
The PERIT-LEGAL leaderboard
Legal benchmarks mostly test recall: pick the right holding from four options, restate the black-letter rule. Firms do not buy recall. They buy work product that survives partner review — and there, a memo that catches eight of ten risks is not eighty percent useful. It is a redraft.
Bar-admitted practitioners pulled tasks from matters they had actually run — redlines, chronologies, diligence memos — and reduced each to a form that publishes without touching privilege.
Each task carries an all-pass rubric: every element a reviewing partner would require, written by the author and adjudicated by a second bar-admitted reviewer before the task is admitted.
A blind panel of practicing lawyers completed the same held-out tasks under ordinary client-work conditions, producing the human control the models are scored beside.
240 tasks are held out and never published. A 60-task development split, built by the same authors through the same pipeline, publishes with the suite alongside the grading harness and the full rubric schema.
Share of tasks where every rubric criterion passes on a single attempt. No partial credit — this is how a partner reviews a draft.
- 01
No model clears the bar for unsupervised legal work.
The leading system all-passes 41.2% of tasks against 74.8% for a bar-admitted practitioner. Every model in the table would require a lawyer to read the output before it left the building — which is the current state of practice, now with a number attached to it.
- 02
Capability is not consistency.
The same leading system passes 84.7% of individual criteria but all-passes only 41.2% of tasks, and holds that across eight runs on just 12.4%. Models are close on most of the work and unreliable about which part they will miss — the failure mode least compatible with legal review.
- 03
Authority, not analysis, is where the drafts break.
Roughly a third of graded failures are citation and authority defects — a real rule applied to the wrong jurisdiction, or a proposition supported by a case that does not stand for it. The reasoning is frequently sound and the support underneath it is not.
- 04
The human control is not the ceiling either.
Practitioners all-pass 74.8% of tasks, not 100%. We publish that because a benchmark that scores humans at perfect is measuring its own rubric rather than the work. The headroom above both columns is the part worth training against.
Every task authored by someone credentialed to do it.
M&A diligence & disclosure
Read the data-room set, identify what a buyer must be told, and quantify the exposure. Graded on whether every material item surfaces, not on how well the memo reads.
Contract redlining
Mark up a counterparty draft against a negotiating position. Every accepted, rejected, and countered term is scored, including the ones a model leaves silently untouched.
Litigation chronology
Build a dated chronology from an unordered document set and flag the entries that carry legal significance. Ordering errors and omitted events are graded separately.
Regulatory research
Answer a live compliance question across overlapping federal and state regimes, with authority. Multi-jurisdiction questions are scored as a distinct hard split.
Transcript & deposition analysis
Work a hearing or deposition transcript for admissions, contradictions, and follow-up that was not asked. Graded against what an examining attorney flagged.
Client-ready drafting
Produce the memo or letter that actually goes out, at the register the recipient expects. Scored on substance first, then on whether it could be sent without a rewrite.
One task, and the rubric behind it.
Buyer's counsel has proposed an indemnity cap of 10% of enterprise value with a 12-month survival period, carved out for fraud and for breach of fundamental representations. The seller is a portfolio company with a disclosed environmental matter in the data room. Advise the seller on cap exposure, identify which carve-outs are uncapped as drafted, and state the governing-law consequence of the survival period as written.
- Cites governing law correctly
- Flags the fraud carve-out as uncapped as drafted
- Quantifies cap exposure against enterprise value
- Identifies the environmental matter as a disclosed exception
Every model in the table gets the fraud carve-out. Most miss that the disclosed environmental matter changes the exposure calculation entirely — the fact is in the schedules, not the clause, and nothing in the prompt points at it.
Where the graded failures land.
The analysis is correct as far as it goes and simply omits something the rubric requires — most often a fact that lives in an attachment rather than the operative clause.
A citation that exists but does not stand for the proposition, or a rule stated accurately for the wrong jurisdiction. Fabricated citations are a small minority of this category.
The governing standard is identified correctly and then applied to the wrong facts, or applied without the exception the disclosure set establishes.
The conclusion is defensible but the reasoning cannot be traced back to a source in the record, so a reviewing lawyer cannot sign off without redoing the work.
What this suite does not measure
Published up front, because a benchmark that names its own limits is the only kind whose number means anything.
- US law only. No coverage of EU, UK, or cross-border regimes beyond the multi-jurisdiction split.
- No tax, immigration, or IP prosecution work in v1.
- Tasks are document-bounded — nothing here measures client counselling, negotiation, or oral advocacy.
- The human control is a panel of practitioners, not the specific partner who would review a given matter.
Run PERIT-LEGAL against your model.
We run the held-out set for any frontier model on request, and return the loss analysis alongside the score. To keep the leaderboard honest, a model is listed from the first run against the prompts — not the best one.