Skip to content

Model comparison

Several models, trained on identical data cut at identical moments, asked about the same matches and scored by the same code. The question this page answers: if each of these had existed before these matches were played, how many would it actually have called correctly?

Backtested Everything here is retrodiction. Each match was priced from a fit over matches played strictly before its own kickoff, which is an honest test — but the code was written knowing how these seasons ended, so none of it is a forecast published in advance.

Big five + Greece 2018-2026
108,035 predictions · 15,473 matches · 5 hours ago
Model Matches Correct Accuracy RPS Log loss Overconfidence
Ensemble 15,404 8,050 52.26% ±0.79 0.1999 0.9855 +0.2%
Dixon-Coles 15,404 8,050 52.26% ±0.79 0.2000 0.9854 -0.7%
Elo 15,404 8,022 52.08% ±0.79 0.2017 1.0008 +1.4%
Poisson 15,404 7,963 51.69% ±0.79 0.2022 1.0056 +1.6%
League base rates baseline 15,404 6,641 43.11% ±0.78 0.2304 1.0751 +0.6%
Random baseline 15,404 5,111 33.18% ±0.74 0.2353 1.0986 +0.2%
Always home baseline 15,404 6,641 43.11% ±0.78 0.4271 2.6285 +54.9%

Are these differences real?

McNemar's test on the matches where exactly one of the pair was right. Matches they both got right, or both got wrong, say nothing about which is better.

Random — in detail

Confusion matrix

Rows: what happened. Columns: what it said.

Home Draw Away
Home 2,198 2,258 2,216
Draw 1,344 1,301 1,302
Away 1,617 1,597 1,640

By outcome

Precision: when it said this, how often was it right. Recall: of the times this happened, how often did it say so.

Outcome Said Happened Precision Recall
Home 5,159 6,672 42.6% 32.9%
Draw 5,156 3,947 25.2% 33.0%
Away 5,158 4,854 31.8% 33.8%

Does confidence mean anything?

Grouped by how sure it was. A calibrated model matches its own claim.

Band Matches Said Happened Gap
< 40% 15,473 33.4% 33.2% +0.2%

If you only followed it when sure

Coverage matters as much as accuracy: a threshold that is right 80% of the time but fires four times a season is a curiosity, not a strategy.

At least Matches Accuracy Coverage
all 15,473 33.2% 100%

When the models agree

Only the pure models are counted here — agreement with the bookmaker benchmark would be measuring something else.

Agreeing Matches Correct Accuracy
2 / 4 378 139 36.8%
3 / 4 1,754 697 39.7%
4 / 4 13,272 7,229 54.5%

By competition

Group Matches Accuracy RPS
La Liga 2,958 34.7% 0.2330
Super League 1 1,738 33.0% 0.2331
Serie A 2,933 32.8% 0.2334
Ligue 1 2,616 33.2% 0.2357
Bundesliga 2,317 30.6% 0.2366
Premier League 2,911 34.3% 0.2394

By season

Group Matches Accuracy RPS
Super League 1 2025/26 236 34.7% 0.2305
Super League 1 2024/25 236 28.0% 0.2411
Super League 1 2023/24 240 32.9% 0.2347
Super League 1 2022/23 240 36.3% 0.2278
Super League 1 2021/22 242 33.1% 0.2358
Super League 1 2020/21 242 34.3% 0.2295
Super League 1 2019/20 242 30.6% 0.2296
Super League 1 2018/19 60 36.7% 0.2444
Serie A 2025/26 380 33.4% 0.2344
Serie A 2024/25 380 33.9% 0.2304
Serie A 2023/24 380 33.9% 0.2287
Serie A 2022/23 381 28.9% 0.2341
Serie A 2021/22 380 31.1% 0.2348
Serie A 2020/21 380 29.5% 0.2357
Serie A 2019/20 380 35.8% 0.2405
Serie A 2018/19 272 37.1% 0.2269
Premier League 2025/26 380 31.6% 0.2322
Premier League 2024/25 380 37.1% 0.2370
Premier League 2023/24 380 31.8% 0.2418
Premier League 2022/23 380 35.5% 0.2396
Premier League 2021/22 380 35.0% 0.2392
Premier League 2020/21 380 35.5% 0.2413
Premier League 2019/20 380 32.6% 0.2374
Premier League 2018/19 251 35.9% 0.2499