Calibration
How trustworthy a rating is, as a function of games played. Not every game counts equally: early in a season the model has barely seen a team, so both the mean (is it right about the game?) and the sigma (is it right about how right it is?) are still settling. The x-axis is the minimum games played by either team at kickoff, pooled across 6 seasons (2021–2026). Everything here is a walk-forward replay — each prediction was made from the ratings at the cutoff that preceded it, never from hindsight.
The two statistics, and what each answers
| family | statistic | what it answers |
|---|---|---|
| mean | rmse of the spread (+ bias) |
is the model right about the game? (bias vs noise) |
| sigma | sd(z), z = (pred − actual)/σ; 1.0 = honest |
is the model right about how right it is? |
| probability | brier, log_loss |
depends on both the mean and the sigma |
Mean — the spread error, across season progress
cbb_model.calibration
(a few tenths of a percent) of its long-run value — essentially flat. A mean settles no later than a sigma, so this gate is earlier than or equal
to the σ gate below.
Axis spans 10.4 to 13.7 points.Sigma — is the model honest about its own uncertainty?
z = (predicted − actual) / σ.
1.0 is the honest value. Note the direction carefully, because it is
easy to state backwards: z divides the residual by the emitted σ, so if
sd(z) < 1 the residuals are narrower than σ
claims — the sigma is too wide and the model is under-confident;
sd(z) > 1 means σ is too narrow and the model is over-confident.
The cumulative line is the read that decides the gate — the gate is the
first cumulative row whose bootstrap 95% CI covers 1.0, and
no row in this store clears it.
The interval resamples teams, not games: games share a team and are
therefore correlated, and treating them as independent understates the interval about
threefold — enough to decide the gate on the wrong row.
Axis spans 0.83 to 1.29.Sigma only — does it order uncertainty at all?
The read that sd(z) above cannot make.
sd(z) is joint with the mean — it mixes whether the mean is off with whether
the σ is off — so it cannot attribute a departure from 1.0 to the σ. This table can:
games are banded by the σ the model emitted, and each band's
actual residual scale is measured. A σ worth emitting must rise with that
column. A flat column means the σ is describing almost nothing; a falling column
means its ordering is inverted against reality.
| emitted σ band | n | mean σ | actual sd of residual |
|---|---|---|---|
| 6.12 – 11.41 | 5,272 | 10.82 | 11.87 |
| 11.41 – 11.95 | 5,272 | 11.70 | 11.69 |
| 11.95 – 12.41 | 5,271 | 12.18 | 11.59 |
| 12.41 – 12.92 | 5,272 | 12.65 | 11.57 |
| 12.92 – 15.53 | 5,272 | 13.42 | 11.89 |
Tail honesty, which a mean SD cannot show. The model emits Gaussian cover probabilities, so the residual should be Gaussian in σ units. 70.7% of games fall within 1σ (a Gaussian implies 68.3%), 95.8% within 2σ (95.4%), and 0.46% beyond 3σ (0.27%) — a rate of 1.7x the Gaussian rate. Excess kurtosis 0.44; the residuals are measurably non-Gaussian (Kolmogorov–Smirnov p = 0.007, n = 26,359).
Cumulative rows, with bootstrap 95% intervals
The authoritative read. A row whose interval contains 1.0 is calibrated — the sigma is not measurably off; one whose interval excludes it is flagged. Read the interval, not the point: it resamples teams, because a team's games are correlated. The gate is the first calibrated row, so it is not identified at any threshold in this store — no cumulative row's interval covers 1.0, which is the finding, not a missing value.
| cutoff | n | rmse | bias | sd(z) | 95% CI | status |
|---|---|---|---|---|---|---|
| gp >= 5 | 26,359 | 11.72 | 0.15 | 0.97 | [0.97, 0.98] | σ off |
| gp >= 8 | 24,057 | 11.52 | 0.14 | 0.95 | [0.94, 0.96] | σ off |
| gp >= 10 | 22,030 | 11.46 | 0.12 | 0.94 | [0.93, 0.95] | σ off |
| gp >= 12 | 20,075 | 11.38 | 0.16 | 0.93 | [0.92, 0.94] | σ off |
| gp >= 14 | 18,094 | 11.34 | 0.16 | 0.92 | [0.91, 0.93] | σ off |
| gp >= 16 | 15,949 | 11.28 | 0.18 | 0.92 | [0.91, 0.93] | σ off |
| gp >= 18 | 13,842 | 11.26 | 0.09 | 0.91 | [0.90, 0.92] | σ off |
| gp >= 20 | 11,790 | 11.24 | 0.10 | 0.91 | [0.90, 0.92] | σ off |
| gp >= 22 | 9,749 | 11.21 | 0.01 | 0.91 | [0.89, 0.92] | σ off |
| gp >= 25 | 6,838 | 11.25 | -0.01 | 0.91 | [0.89, 0.93] | σ off |
| gp >= 28 | 4,086 | 11.27 | 0.20 | 0.91 | [0.89, 0.93] | σ off |
Per-bucket rows (the noise the cumulative read absorbs)
Show the per-games-played buckets
| bucket | n | rmse | bias | sd(z) | 95% CI | brier |
|---|---|---|---|---|---|---|
| gp 5 | 435 | 13.68 | 0.78 | 1.29 | [1.20, 1.39] | 0.188 |
| gp 6 | 848 | 13.57 | -0.04 | 1.22 | [1.16, 1.28] | 0.197 |
| gp 7 | 1,019 | 13.71 | 0.24 | 1.19 | [1.14, 1.25] | 0.200 |
| gp 8 | 1,007 | 12.38 | 0.23 | 1.07 | [1.01, 1.13] | 0.188 |
| gp 9 | 1,020 | 11.99 | 0.47 | 1.02 | [0.98, 1.07] | 0.166 |
| gp 10 | 968 | 12.39 | -0.38 | 1.06 | [1.01, 1.10] | 0.175 |
| gp 11 | 987 | 12.13 | -0.01 | 1.03 | [0.97, 1.08] | 0.166 |
| gp 12 | 951 | 12.02 | -0.34 | 1.00 | [0.95, 1.04] | 0.176 |
| gp 13 | 1,030 | 11.37 | 0.58 | 0.94 | [0.90, 0.98] | 0.201 |
| gp 14 | 1,060 | 11.57 | -0.40 | 0.96 | [0.92, 1.00] | 0.199 |
| gp 15 | 1,085 | 11.95 | 0.33 | 0.98 | [0.94, 1.02] | 0.206 |
| gp 16 | 1,060 | 11.36 | 0.49 | 0.94 | [0.89, 0.98] | 0.204 |
| gp 17 | 1,047 | 11.51 | 1.09 | 0.94 | [0.89, 0.98] | 0.193 |
| gp 18 | 1,037 | 11.33 | -0.00 | 0.93 | [0.88, 0.97] | 0.189 |
| gp 19 | 1,015 | 11.42 | 0.12 | 0.94 | [0.89, 0.98] | 0.203 |
| gp 20 | 1,020 | 11.35 | 0.59 | 0.92 | [0.88, 0.96] | 0.194 |
| gp 21 | 1,021 | 11.49 | 0.46 | 0.93 | [0.89, 0.98] | 0.204 |
| gp 22 | 995 | 11.08 | -0.22 | 0.90 | [0.86, 0.94] | 0.193 |
| gp 23 | 966 | 11.38 | -0.20 | 0.92 | [0.87, 0.96] | 0.193 |
| gp 24 | 950 | 10.80 | 0.61 | 0.88 | [0.84, 0.92] | 0.198 |
| gp 25 | 946 | 11.47 | -0.37 | 0.93 | [0.89, 0.97] | 0.197 |
| gp 26 | 906 | 11.27 | -0.48 | 0.92 | [0.87, 0.96] | 0.193 |
| gp 27 | 900 | 10.91 | -0.13 | 0.89 | [0.85, 0.93] | 0.190 |
| gp 28 | 864 | 11.34 | 0.41 | 0.92 | [0.87, 0.96] | 0.191 |
| gp 29 | 814 | 11.21 | 0.78 | 0.91 | [0.86, 0.96] | 0.189 |
| gp 30 | 771 | 11.76 | -0.64 | 0.95 | [0.88, 1.03] | 0.188 |
| gp 31 | 822 | 11.14 | -0.05 | 0.89 | [0.84, 0.94] | 0.197 |
| gp 32 | 330 | 10.41 | 0.33 | 0.83 | [0.77, 0.89] | 0.196 |
| gp 33 | 253 | 11.51 | 0.51 | 0.92 | [0.84, 0.99] | 0.195 |
| gp 34 | 126 | 11.37 | 1.36 | 0.90 | [0.79, 1.00] | 0.194 |
Individually these intervals overlap and clear 1.0 irregularly; that is why the cumulative rows above are the ones a claim rests on.
Tournament — the NCAA slice, per season and pooled
A single bracket is small and easy to over-read — the repo's own retraction #1 is a one-season 67-game slice that claimed the sigma was 13–15% overinflated and did not survive pooling. So spread and total are reported separately (a pooled margin error would confound two different quantities), each with its n, and the pooled figure sits beside the per-season ones.
| slice | market | n | rmse | bias | pred σ | actual sd |
|---|---|---|---|---|---|---|
| 2021 | spread | 66 | 12.64 | 1.66 | 12.68 | 14.02 |
| total | 66 | 18.24 | 4.88 | 16.02 | 19.30 | |
| 2022 | spread | 67 | 12.27 | 0.72 | 12.53 | 13.71 |
| total | 67 | 18.48 | 3.20 | 15.76 | 19.63 | |
| 2023 | spread | 67 | 12.50 | 0.87 | 12.94 | 12.73 |
| total | 67 | 17.26 | 4.69 | 15.87 | 17.47 | |
| 2024 | spread | 67 | 12.59 | -0.19 | 12.57 | 15.65 |
| total | 67 | 19.36 | 3.21 | 16.12 | 21.48 | |
| 2025 | spread | 67 | 10.03 | -0.52 | 12.84 | 13.77 |
| total | 67 | 13.88 | 0.05 | 16.19 | 17.13 | |
| 2026 | spread | 67 | 11.48 | -0.18 | 12.62 | 15.59 |
| total | 67 | 14.87 | 2.19 | 16.19 | 18.03 | |
| Pooled | spread | 401 | 11.95 | 0.39 | 12.70 | 14.29 |
| Pooled | total | 401 | 17.13 | 3.03 | 16.03 | 19.36 |
pred σ is the mean sigma the model emitted; actual sd is the spread the results actually had. They should be close — a predicted sigma far from the observed spread is the miscalibration, not the RMSE.
vs Vegas — gated on real odds
Lines are present in the store (42 rows). This block compares the model's margin and total to the de-vigged closing consensus once the line ingest has run for real and predictions join a line — today no edges exist yet.