Expert data · marketplace

Tasks, rubrics, and RL environments — authored by the experts your model is trying to replace.

Perit connects credentialed physicians, lawyers, bankers, and senior engineers with the labs and enterprises that need human judgment to train and evaluate AI. Pick the use case; we deliver the artifact, the credential behind it, and the price.

Talk to our research team →
task_4417.jsonverifier · pass
Prompt · clinical reasoning

62F, new microcytic anemia, ferritin 8 ng/mL, recent NSAID use. Choose the next diagnostic step and justify.

Expert answer

Upper + lower endoscopy before iron repletion. NSAID GI loss can mask a synchronous malignancy; ferritin confirms iron deficiency, not its source.

verifier → [0,1]
1.00
IAA · n=3
0.94
annotatorr.okafor
reviewerj.ikeda
✓ MDboard-certifiedheme/onc · 11 yrs
12
use-case clusters
5
fixed spec fields / card
3,400+
vetted experts
$140/hr
avg published rate
Figures shown here are illustrative during the build — verified counts replace them before launch.
[01]The failure we fix

Your model isn't short on parameters. It's short on people who know what right looks like.

Frontier models plateau where the training data runs out of expertise. More generic data doesn't move the frontier — the judgment of people who do the work for a living does. Perit turns that judgment into artifacts a model can actually learn from and be graded against.

[02]Use cases · the backbone

Start with the thing you're trying to buy.

Twelve clusters. Every card names the artifact, the unit you buy it in, the credential behind it, and the rate. No mission statements.

[01]Post-training

Post-training data

Teach the model what a correct answer actually looks like.

DeliverableSFT demos + preference pairs
Unitper task
CredentialPractitioners, 5+ yrs
Rate$60–140
[02]Grading

Rubrics & verifiers

Define "good" precisely enough to grade a million outputs.

DeliverableScoring rubric + reference set
Unitper rubric
CredentialSenior domain reviewers
Rate$2k–8k
[03]Agents

RL environments

Train agents against a faithful clone of the real tool.

DeliverableExecutable env + reward fn
Unitper environment
CredentialPower users of the tool
Rate$8k–40k
[04]Measurement

Benchmarks & eval

Publish a score you can defend in a paper.

DeliverableHeld-out benchmark set
Unitper benchmark
CredentialPhD / research-grade
RateProgram
[05]Verticals

Professional judgment

Get the call a licensed professional would make.

DeliverableAdjudicated judgments
Unitper hour
CredentialLicensed / board-cert.
Rate$120–300
[06]Code

Agentic coding & SWE

Grade whether the agent shipped working code.

DeliverableRepo tasks + graded patches
Unitper task
CredentialStaff-level engineers
Rate$90–180
[07]Multimodal

Vision & embodied

Label perception the way an expert eye sees it.

DeliverableAnnotated frames + trajectories
Unitper clip
CredentialTrained annotators + QA
Rate$40–120
[08]Production

Live agent eval

Catch regressions before your users do.

DeliverableContinuous grading stream
Unitper month
CredentialReviewers on retainer
RateSubscription
[09]Enterprise

Enterprise agents

Turn how your best people work into an agent that works.

DeliverableWorkflow capture + task suite
Unitper workflow
CredentialYour SMEs + our authors
RateProgram
[10]Licensing

Data monetization

Turn your workflow exhaust into a licensable asset.

DeliverableRights-cleared dataset
Unitper license
CredentialProvenance + legal review
RateRev-share
[11]Safety

Red-teaming & safety

Find the failure a real adversary would find first.

DeliverableAttack sets + severity labels
Unitper engagement
CredentialDomain + security cleared
Rate$100–250
[12]Consumer

Everyday-life data

Ground everyday tasks in how people really do them.

DeliverableTask demos + preferences
Unitper task
CredentialVetted everyday experts
Rate$25–60
[03]RL environments

An executable replica of the work — not a description of it.

We rebuild the real tool — a CRM, a spreadsheet, an EHR — as an environment the agent can act inside, with a reward function an expert wrote. The agent doesn't answer questions about the work. It does the work.

14
gated actions
~40
avg steps / episode
0–1
dense reward
1:1
tool fidelity
Failure mode
Agents that ace the chat and fail the click. Fluent plans, wrong buttons.
env: salesforce_clone_v3episode 214
› agent.observe(state)
open_opportunity("Northwind — Q3")
› agent.act(step=1)
set_stage("Negotiation") ok
› agent.act(step=2)
attach_quote(v2.pdf) ok
› agent.act(step=3)
email_customer(no_approval) blocked
reward = 0.67 · policy violation: step 3
authored by
enterprise RevOps lead
reward fn
14 gated actions
[04]Rubrics & verifiers

A number you can audit, not an adjective.

A rubric is what "good" looks like, written down by someone who would know. A verifier turns that rubric into a function that returns a score in [0,1] — so you can grade a million outputs the way one expert would grade one.

4–12
criteria / rubric
Σ=1.0
weighted
0.9+
target IAA
[0,1]
verifier range
Failure mode
"Looks good" is not a grade. Un-scored preference data teaches the model your reviewers' mood.
rubric: contract_redline_v4output #82,104
Cites governing law correctly
weight 0.30
0.9
Flags the fraud carve-out
weight 0.30
0.8
Quantifies cap exposure
weight 0.25
0.5
Reasoning is auditable
weight 0.15
0.6
verifier → [0,1]0.71
[05]Vertical professional judgment

The call a board-certified professional would actually make.

Medicine, law, finance, consulting. Adjudicated judgments with the reasoning attached, gated behind a license we verified — not a self-reported bio. Every judgment carries the credential chain that produced it.

4
core verticals
3
reviewers / call
100%
license-verified
full
reasoning attached
Failure mode
Confident, fluent, and wrong. The answers that read best are the ones a layperson can't catch.
adjudication_9920 · law3 reviewers
Question · M&A indemnification

Does the cap survive a fraud carve-out under Delaware law as drafted?

consensus
No — carve-out controls
agreement
3 / 3
✓ JD · bar-admitted DEM&A · 14 yrsNDA on file
[06]How it works

A path, not a menu.

01

Capture the workflow

We sit with the people who do the work and record how it actually happens — the tools, the edge cases, the judgment calls.

3–5 days
02

Author tasks with credentialed experts

Practitioners write the tasks and demonstrations. The credential that qualifies them is verified and attached.

1–2 weeks
03

Build the verifier / rubric

Senior reviewers turn "good" into a scoring function that returns a number in [0,1].

3–7 days
04

Deliver and measure

You get the artifact plus an inter-annotator agreement report and a delta against your gold set.

On delivery
05

Continuously re-grade

Optional: reviewers stay on retainer to grade the deployed agent and catch drift.

Ongoing
[07]Live inventory · priced roles

Every role priced. Every seat gated.

sample of open roles
RoleClusterCredential gateRate
Clinical reasoning reviewerVertical judgmentMD, board-certified$140–300
M&A contract adjudicatorVertical judgmentJD, bar-admitted$150–260
RevOps environment authorRL environmentsSalesforce admin, 5+ yrs$110–170
SWE task graderAgentic codingStaff engineer$90–180
Systematic trading rubric authorRubrics & verifiersBuy-side quant$130–220
Radiology annotation QAVision & embodiedRadiologist + trained team$120–240
[08]Benchmark · research

We publish the score even when it's low.

PERIT-BENCH grades frontier models on real, expert-authored professional work — the same tasks our specialists are paid to judge. The top score is low on purpose. The headroom is the whole point.

PREVIEW · scores below are illustrative. Real, citable results publish with the benchmark.

See the four suites →
PERIT-BENCH · professional workscore / 100
01Frontier model · lab A41.7
02Frontier model · lab B38.2
03Frontier model · lab C34.9
04Open weights · 70B27.1
05Open weights · 8B15.4
human expert baseline · 100.0
[09]Field note

We asked board-certified oncologists to grade a model's treatment-plan reasoning. It passed the reading. It failed the medicine.

[—]TBDof plans a first-year resident would flag were rated “sound” by the model. Real number publishes with the study.

Fluency is not competence. The gap between an answer that reads well and one that is correct is exactly the gap expert data closes — and exactly the gap a generic labeling crowd cannot see.

[10]Transparent economics

You see the price. And the person.

A flat percentage on top of what the expert is paid. No opaque markup, no mystery quote. You approve the profile before the work starts, and the data stays on your platform.

01A flat percentage on expert pay — the same margin on a $40 task and a $300 one.
02You approve the expert profile, with credential and rate, before any work begins.
03The data is created on your platform, in your cloud. We never hold a copy you didn’t authorize.
How the money flows · illustrative
expert · 80%
Perit · 20%
You pay$100 / hr
Expert receives$80 / hr
Flat platform margin20%, fixed
[11]Security · provenance · governance

Who touched your data, and when.

Data residency
Work happens in your cloud or a dedicated enclave. Your choice of region.
Provenance
Every artifact carries the annotator, reviewer, and credential chain that produced it.
Vetting chain
Identity, credential, and license verified before an expert touches a task.
NDA / IP terms
Per-project NDAs; IP assigns to you on delivery. Standard terms published up front.
Access
Least-access, time-boxed workstations. No downloads, full session capture.
Incident policy
Defined severity tiers, named contacts, and a disclosure clock in the MSA.
audit_log · dataset_2210immutable
14:02:11r.okafor · opened task_4417
14:19:44r.okafor · submitted answer
15:03:02j.ikeda · review → accept (0.94)
15:03:02verifier · score committed [1.00]
15:07:30export → client_cloud/us-east
[12]Pricing · starter SKU

A pilot with a price, not a “get in touch.”

Start with one use case, one rubric, one agreement report. Small enough to run in a sprint, real enough to tell you whether expert data moves your metric.

Series A–C teams that need 300 expert-labeled tasks — not 87,000 — start here. So do the labs, on their first cluster.

The Pilot
from $9kpilot
fixed scope · 2–3 weeks
300 expert-labeled tasks in one use-case cluster
One expert-authored rubric + verifier
Inter-annotator agreement report
Named experts with verified credentials
Data delivered on your platform
Scope a pilot →
[13]Definitions & questions

You can't buy a category you can't name.

Plain definitions first. The vocabulary of expert data, without the pitch.

rubric
A written definition of what a good answer looks like, authored by someone qualified to judge — the standard everything else is scored against.
verifier
A function that applies a rubric automatically and returns a score in [0,1], so one expert’s judgment can grade a million outputs.
evaluation
A held-out set of expert-graded tasks used to measure a model’s real capability — not benchmark memorization.
RL environment
An executable replica of a real tool the agent acts inside, paired with a reward function that scores what it did.

An artifact: a set of tasks, a rubric and verifier, an RL environment, or a benchmark — named on each use-case card, priced by a clear unit, and gated behind a verified credential.

Tell us the use case. We'll send a sample task set back.

No mission statement, no discovery-call maze. Name the artifact you need and the credential it requires — we reply with a real sample and a price.

Talk to our research team