Every prediction is written down before kickoff and graded afterwards. The first prediction stands — nothing is revised once a result is known.
Is it honest?
Accuracy can be gamed by only predicting easy matches. Calibration cannot: to be calibrated the model has to be right about its own uncertainty, so its 35% calls must fail about 65% of the time. Each row groups predictions by how confident the model was, and compares what it claimed against what happened.
| Confidence | Calls | It said | Landed | Said v landed |
|---|---|---|---|---|
| 30-40% | 75 | 36.8% | 42.7% | |
| 40-50% | 85 | 44.2% | 41.2% | |
| 50-60% | 56 | 54.7% | 57.1% | |
| 60-70% | 28 | 65% | 78.6% | |
| 70-80% | 6 | 74.5% | 100% |
Grey is what the model claimed, blue is what happened. Landing at or slightly above the claim means it is honest and a little underconfident — the safe direction to be wrong. Buckets with few calls will move a lot; treat small samples as noise.
By competition
Open a competition for every prediction it has made, graded match by match.