[02]Use case · Grading

Define "good" precisely enough to grade a million outputs.

A written rubric authored by a senior reviewer, compiled into a verifier that returns a score in [0,1] — so one expert's judgment can grade at machine scale without drifting.

Deliverable
Scoring rubric + reference set
Priced by
per rubric
Credential gate
Senior domain reviewers
Typical rate
$2k–8k

What you're buying

A weighted rubric (criteria that sum to 1.0), a reference set of graded examples, and a verifier function that applies the rubric automatically and returns a number in [0,1].

Why it holds up

The rubric is authored against real disagreements, not a clean textbook case. We tune criteria until inter-annotator agreement clears the target, then freeze the reference set the verifier is checked against.

How you use it

Point the verifier at a million model outputs and get a defensible score for each — the same judgment one senior reviewer would apply, minus the bottleneck of their calendar.

By the numbers
4–12
criteria / rubric
Σ=1.0
weighted
0.9+
target IAA
[0,1]
verifier range