Because without it the number claims more precision than it has. A ten thousand goal series carries an interval about nine tenths of a point either side, a little under two points end to end, and some of the changes made to this run are worth less than that. Print the number alone and you cannot tell an improvement from a re-run.
The chart on the front page holds one measurement repeated: about ten thousand goals against Nexto from varied starting situations, every thirty five minutes or so. The harness runs other evaluations against the same opponent to the same goal count, and the most frequent of those plays nothing but kickoffs. Its middle reading is a couple of points above the headline and it is a far rougher number than that sounds, having come back as low as 38% and as high as 99%. Four more are one offs: a longer series, two against a different bot, and one against this run's own earlier checkpoint. All real, none of them on that line, because a trend built from mixed protocols averages different questions together.
Inside the one protocol the harness still moves around more than a single interval suggests. Across the last thirty series the middle reading is 69.2%, the lowest 66.7% and the highest 70.4%, and nothing on that line has ever come back above 71%. A reading near 100% attached to this bot came from a kickoff only run having a good day. The headline is the last series confirmed under the charted protocol, not the best number the harness has ever produced.