[04]Use case · Measurement

Publish a score you can defend in a paper.

A held-out set of expert-graded tasks that measures real capability instead of benchmark memorization — hard on purpose, citable by design.

Deliverable
Held-out benchmark set
Priced by
per benchmark
Credential gate
PhD / research-grade
Typical rate
Program

What you're buying

A benchmark built and graded by research-grade experts, held out of any training corpus, with a verifier and a full provenance chain so the result stands up to review.

Honest by construction

The scores are meant to be low. A benchmark that frontier models ace on day one measures nothing — we grade against the ceiling of professional work, not the floor.

What you can cite

Every task carries the author, the reviewer, and the credential behind them. When you publish the number, you can publish exactly how it was produced.

By the numbers
held-out
never in training
PhD
grader bar
[0,1]
verifier range
citable
with provenance