Position · 19 Jun 2026

The human control is not a ceiling

Publishing a human line at 100 makes the benchmark useless the moment a model gets close.

The person we put on the control is a working practitioner doing the same task blind, at market rate, on the clock. They make mistakes: mishear a digit on a bad line, miss a criterion on a long call. Their score lands where it lands, and it is printed with the same error bar as everything else.

That matters for reading the board. A model at 60 against a human at 75 is a different story from a model at 60 against a perfect score, and only one of those two comparisons is honest about what the job actually requires.