Predictorous

Every prediction is written down before kickoff and graded afterwards. The first prediction stands — nothing is revised once a result is known.

Calls graded
250
Landed
50.8%
It claimed
47.4%
Exact score
11.2%

Is it honest?

Accuracy can be gamed by only predicting easy matches. Calibration cannot: to be calibrated the model has to be right about its own uncertainty, so its 35% calls must fail about 65% of the time. Each row groups predictions by how confident the model was, and compares what it claimed against what happened.

ConfidenceCallsIt saidLandedSaid v landed
30-40%7536.8%42.7%
40-50%8544.2%41.2%
50-60%5654.7%57.1%
60-70%2865%78.6%
70-80%674.5%100%

Grey is what the model claimed, blue is what happened. Landing at or slightly above the claim means it is honest and a little underconfident — the safe direction to be wrong. Buckets with few calls will move a lot; treat small samples as noise.

By competition

Open a competition for every prediction it has made, graded match by match.