Tasks, rubrics, and RL environments — authored by the experts your model is trying to replace.
Perit connects credentialed physicians, lawyers, bankers, and senior engineers with the labs and enterprises that need human judgment to train and evaluate AI. Pick the use case; we deliver the artifact, the credential behind it, and the price.
62F, new microcytic anemia, ferritin 8 ng/mL, recent NSAID use. Choose the next diagnostic step and justify.
Upper + lower endoscopy before iron repletion. NSAID GI loss can mask a synchronous malignancy; ferritin confirms iron deficiency, not its source.
Your model isn't short on parameters. It's short on people who know what right looks like.
Frontier models plateau where the training data runs out of expertise. More generic data doesn't move the frontier — the judgment of people who do the work for a living does. Perit turns that judgment into artifacts a model can actually learn from and be graded against.
Start with the thing you're trying to buy.
Twelve clusters. Every card names the artifact, the unit you buy it in, the credential behind it, and the rate. No mission statements.
Post-training data
Teach the model what a correct answer actually looks like.
Rubrics & verifiers
Define "good" precisely enough to grade a million outputs.
RL environments
Train agents against a faithful clone of the real tool.
Benchmarks & eval
Publish a score you can defend in a paper.
Professional judgment
Get the call a licensed professional would make.
Agentic coding & SWE
Grade whether the agent shipped working code.
Vision & embodied
Label perception the way an expert eye sees it.
Live agent eval
Catch regressions before your users do.
Enterprise agents
Turn how your best people work into an agent that works.
Data monetization
Turn your workflow exhaust into a licensable asset.
Red-teaming & safety
Find the failure a real adversary would find first.
Everyday-life data
Ground everyday tasks in how people really do them.
An executable replica of the work — not a description of it.
We rebuild the real tool — a CRM, a spreadsheet, an EHR — as an environment the agent can act inside, with a reward function an expert wrote. The agent doesn't answer questions about the work. It does the work.
A number you can audit, not an adjective.
A rubric is what "good" looks like, written down by someone who would know. A verifier turns that rubric into a function that returns a score in [0,1] — so you can grade a million outputs the way one expert would grade one.
The call a board-certified professional would actually make.
Medicine, law, finance, consulting. Adjudicated judgments with the reasoning attached, gated behind a license we verified — not a self-reported bio. Every judgment carries the credential chain that produced it.
Does the cap survive a fraud carve-out under Delaware law as drafted?
A path, not a menu.
Capture the workflow
We sit with the people who do the work and record how it actually happens — the tools, the edge cases, the judgment calls.
Author tasks with credentialed experts
Practitioners write the tasks and demonstrations. The credential that qualifies them is verified and attached.
Build the verifier / rubric
Senior reviewers turn "good" into a scoring function that returns a number in [0,1].
Deliver and measure
You get the artifact plus an inter-annotator agreement report and a delta against your gold set.
Continuously re-grade
Optional: reviewers stay on retainer to grade the deployed agent and catch drift.
Every role priced. Every seat gated.
sample of open roles| Role | Cluster | Credential gate | Rate |
|---|---|---|---|
| Clinical reasoning reviewer | Vertical judgment | MD, board-certified | $140–300 |
| M&A contract adjudicator | Vertical judgment | JD, bar-admitted | $150–260 |
| RevOps environment author | RL environments | Salesforce admin, 5+ yrs | $110–170 |
| SWE task grader | Agentic coding | Staff engineer | $90–180 |
| Systematic trading rubric author | Rubrics & verifiers | Buy-side quant | $130–220 |
| Radiology annotation QA | Vision & embodied | Radiologist + trained team | $120–240 |
We publish the score even when it's low.
PERIT-BENCH grades frontier models on real, expert-authored professional work — the same tasks our specialists are paid to judge. The top score is low on purpose. The headroom is the whole point.
PREVIEW · scores below are illustrative. Real, citable results publish with the benchmark.
See the four suites →We asked board-certified oncologists to grade a model's treatment-plan reasoning. It passed the reading. It failed the medicine.
Fluency is not competence. The gap between an answer that reads well and one that is correct is exactly the gap expert data closes — and exactly the gap a generic labeling crowd cannot see.
You see the price. And the person.
A flat percentage on top of what the expert is paid. No opaque markup, no mystery quote. You approve the profile before the work starts, and the data stays on your platform.
Who touched your data, and when.
A pilot with a price, not a “get in touch.”
Start with one use case, one rubric, one agreement report. Small enough to run in a sprint, real enough to tell you whether expert data moves your metric.
Series A–C teams that need 300 expert-labeled tasks — not 87,000 — start here. So do the labs, on their first cluster.
You can't buy a category you can't name.
Plain definitions first. The vocabulary of expert data, without the pitch.
Tell us the use case. We'll send a sample task set back.
No mission statement, no discovery-call maze. Name the artifact you need and the credential it requires — we reply with a real sample and a price.