Publishing a human line at 100 makes the benchmark useless the moment a model gets close.
The person we put on the control is a working practitioner doing the same task blind, at market rate, on the clock. They make mistakes: mishear a digit on a bad line, miss a criterion on a long call. Their score lands where it lands, and it is printed with the same error bar as everything else.
That matters for reading the board. A model at 60 against a human at 75 is a different story from a model at 60 against a perfect score, and only one of those two comparisons is honest about what the job actually requires.