Every idea below was built, tested against the shipping model on fixtures it had not seen, and lost. They are here because a model you can only see the successes of is a model you have no reason to trust — and because several of these are the ideas anyone would try next.
The test is a paired bootstrap on the fixtures an idea actually touches. P(better) is the share of resamples in which the new version beat the old one — about 0.5 is a coin flip. Nothing here got near enough to 0.95 to ship, so nothing here shipped.
A team-quality layer inside the predictor
rejected+0.0001 log lossThe hope. Twenty-eight opponent-adjusted team metrics ought to know something about Saturday that goal-based ratings do not.
What happened. It moved the log loss by +0.0001. Not a small win — no win. The information was already in the chance rates the model reads.
The manager as a separate effect
rejected−0.0132 log loss when fully spell-scopedThe hope. A new manager changes a side, so a manager term should carry signal the club's own history does not.
What happened. Tested four independent ways and rejected each time. What survives is SELECTION — who he picks — and that is already visible in the squad. Scoping the whole history to the current spell made it actively worse.
Availability corrections to a club's history
not provenP(better) = 0.633The hope. A side missing three regulars is not the side its record describes, so the history should be discounted.
What happened. The right shape — it has the property that a full-strength side gets no correction at all — but the effect could not be separated from noise on 232 fixtures.
Blending the process-Elo arm into the shipped model
not provenP(better) = 0.790The hope. Two independent models usually beat one.
What happened. Better on the sample, nowhere near significantly. It has not shipped and will not until a bigger sample says so.
Seeding the model with a manager's record
not provenP(better) = 0.840The hope. Same idea as above, from the other end.
What happened. Closer, still short. Same answer: wait for the data.
Context features for the draw model
rejected28.9% vs 29.6%, inside noiseThe hope. Points-per-game gap, how far into the season it is, whether both sides are comfortable — a drawish match ought to be predictable from its situation.
What happened. The model assigned them coefficients of ±0.001 to 0.04 and out-of-sample performance did not move. Consistent with the separate finding that drawish weeks are random clumping: overdispersion measured 1.0 across ten league-seasons.
Zones built from individual players
rejectedThe hope. If a side's zone profile is the sum of who plays in it, building it from players should beat building it from team totals.
What happened. It did not beat the shipping model. The team totals already had it.
“Volume beats weighting” — my own claim
reversed3.77 projected vs 2.75 actualThe hope. I argued that how much a side creates matters more than how the chances are weighted, and had a result to show for it.
What happened. The result came from a conversion constant I had fitted by log loss instead of solving. Projections came out at 3.77 goals a game against an actual 2.75. Solving it properly INVERTED the ranking and falsified the claim.
What this leaves
A shipping model that beats picking the home side by 9.6 points of accuracy, is 6.9% better than league base rates on log loss, and is calibrated where the mass is — matches it calls at 25–35% happen 29% of the time. Everything above was an attempt to improve on that and did not.
The honest constraint is sample size. Most of these were tested on a few hundred fixtures, which cannot separate ideas that differ by a few thousandths of log loss. They are not dead — they are unproven, and the difference matters. Four more leagues are being collected for exactly this reason.
The track record is what the model that survived all this has actually done.