cbb·system

Calibration

How trustworthy a rating is, as a function of games played. Not every game counts equally: early in a season the model has barely seen a team, so both the mean (is it right about the game?) and the sigma (is it right about how right it is?) are still settling. The x-axis is the minimum games played by either team at kickoff, pooled across 6 seasons (2021–2026). Everything here is a walk-forward replay — each prediction was made from the ratings at the cutoff that preceded it, never from hindsight.

competition Men’s predictions scored 26,359 seasons pooled 6 mean gate 18 gp sigma gate 18 gp

The two statistics, and what each answers

familystatisticwhat it answers
meanrmse of the spread (+ bias) is the model right about the game? (bias vs noise)
sigmasd(z), z = (pred − actual)/σ; 1.0 = honest is the model right about how right it is?
probabilitybrier, log_loss depends on both the mean and the sigma

Mean — the spread error, across season progress

mens calibration — rmse vs games playedmean gate: 18 games played — the RMSE is within a few tenths of a percent of its long-run value
RMSE of the predicted home margin, by games played. The thin line is the per-games-played bucket (faithful to the measured points); the bold line is cumulative — every game where both teams had at least N games — and it is the read to trust, because per-bucket points are noisy. The dashed rule is this curve's own gate: the mean gate at 18 games, where the RMSE is within the tolerance recorded beside the constant in cbb_model.calibration (a few tenths of a percent) of its long-run value — essentially flat. A mean settles no later than a sigma, so this gate is earlier than or equal to the σ gate below. Axis spans 10.4 to 13.7 points.

Sigma — is the model honest about its own uncertainty?

mens calibration — sd_z vs games playedsigma gate: 18 games played (cumulative sd(z) CI first covers 1.0)
The SD of the standardized residual z = (predicted − actual) / σ. 1.0 is the honest value. Note the direction carefully, because it is easy to state backwards: z divides the residual by the emitted σ, so if sd(z) < 1 the residuals are narrower than σ claims — the sigma is too wide and the model is under-confident; sd(z) > 1 means σ is too narrow and the model is over-confident. The cumulative line is the read that decides the gate — the gate is the first cumulative row whose bootstrap 95% CI covers 1.0, and no row in this store clears it. The interval resamples teams, not games: games share a team and are therefore correlated, and treating them as independent understates the interval about threefold — enough to decide the gate on the wrong row. Axis spans 0.83 to 1.29.

Sigma only — does it order uncertainty at all?

The read that sd(z) above cannot make. sd(z) is joint with the mean — it mixes whether the mean is off with whether the σ is off — so it cannot attribute a departure from 1.0 to the σ. This table can: games are banded by the σ the model emitted, and each band's actual residual scale is measured. A σ worth emitting must rise with that column. A flat column means the σ is describing almost nothing; a falling column means its ordering is inverted against reality.

emitted σ bandnmean σ actual sd of residual
6.12 – 11.41 5,272 10.82 11.87
11.41 – 11.95 5,272 11.70 11.69
11.95 – 12.41 5,271 12.18 11.59
12.41 – 12.92 5,272 12.65 11.57
12.92 – 15.53 5,272 13.42 11.89

Tail honesty, which a mean SD cannot show. The model emits Gaussian cover probabilities, so the residual should be Gaussian in σ units. 70.7% of games fall within 1σ (a Gaussian implies 68.3%), 95.8% within 2σ (95.4%), and 0.46% beyond 3σ (0.27%) — a rate of 1.7x the Gaussian rate. Excess kurtosis 0.44; the residuals are measurably non-Gaussian (Kolmogorov–Smirnov p = 0.007, n = 26,359).

Cumulative rows, with bootstrap 95% intervals

The authoritative read. A row whose interval contains 1.0 is calibrated — the sigma is not measurably off; one whose interval excludes it is flagged. Read the interval, not the point: it resamples teams, because a team's games are correlated. The gate is the first calibrated row, so it is not identified at any threshold in this store — no cumulative row's interval covers 1.0, which is the finding, not a missing value.

cutoffnrmsebias sd(z)95% CIstatus
gp >= 5 26,359 11.72 0.15 0.97 [0.97, 0.98] σ off
gp >= 8 24,057 11.52 0.14 0.95 [0.94, 0.96] σ off
gp >= 10 22,030 11.46 0.12 0.94 [0.93, 0.95] σ off
gp >= 12 20,075 11.38 0.16 0.93 [0.92, 0.94] σ off
gp >= 14 18,094 11.34 0.16 0.92 [0.91, 0.93] σ off
gp >= 16 15,949 11.28 0.18 0.92 [0.91, 0.93] σ off
gp >= 18 13,842 11.26 0.09 0.91 [0.90, 0.92] σ off
gp >= 20 11,790 11.24 0.10 0.91 [0.90, 0.92] σ off
gp >= 22 9,749 11.21 0.01 0.91 [0.89, 0.92] σ off
gp >= 25 6,838 11.25 -0.01 0.91 [0.89, 0.93] σ off
gp >= 28 4,086 11.27 0.20 0.91 [0.89, 0.93] σ off

Per-bucket rows (the noise the cumulative read absorbs)

Show the per-games-played buckets
bucketnrmsebias sd(z)95% CIbrier
gp 5 435 13.68 0.78 1.29 [1.20, 1.39] 0.188
gp 6 848 13.57 -0.04 1.22 [1.16, 1.28] 0.197
gp 7 1,019 13.71 0.24 1.19 [1.14, 1.25] 0.200
gp 8 1,007 12.38 0.23 1.07 [1.01, 1.13] 0.188
gp 9 1,020 11.99 0.47 1.02 [0.98, 1.07] 0.166
gp 10 968 12.39 -0.38 1.06 [1.01, 1.10] 0.175
gp 11 987 12.13 -0.01 1.03 [0.97, 1.08] 0.166
gp 12 951 12.02 -0.34 1.00 [0.95, 1.04] 0.176
gp 13 1,030 11.37 0.58 0.94 [0.90, 0.98] 0.201
gp 14 1,060 11.57 -0.40 0.96 [0.92, 1.00] 0.199
gp 15 1,085 11.95 0.33 0.98 [0.94, 1.02] 0.206
gp 16 1,060 11.36 0.49 0.94 [0.89, 0.98] 0.204
gp 17 1,047 11.51 1.09 0.94 [0.89, 0.98] 0.193
gp 18 1,037 11.33 -0.00 0.93 [0.88, 0.97] 0.189
gp 19 1,015 11.42 0.12 0.94 [0.89, 0.98] 0.203
gp 20 1,020 11.35 0.59 0.92 [0.88, 0.96] 0.194
gp 21 1,021 11.49 0.46 0.93 [0.89, 0.98] 0.204
gp 22 995 11.08 -0.22 0.90 [0.86, 0.94] 0.193
gp 23 966 11.38 -0.20 0.92 [0.87, 0.96] 0.193
gp 24 950 10.80 0.61 0.88 [0.84, 0.92] 0.198
gp 25 946 11.47 -0.37 0.93 [0.89, 0.97] 0.197
gp 26 906 11.27 -0.48 0.92 [0.87, 0.96] 0.193
gp 27 900 10.91 -0.13 0.89 [0.85, 0.93] 0.190
gp 28 864 11.34 0.41 0.92 [0.87, 0.96] 0.191
gp 29 814 11.21 0.78 0.91 [0.86, 0.96] 0.189
gp 30 771 11.76 -0.64 0.95 [0.88, 1.03] 0.188
gp 31 822 11.14 -0.05 0.89 [0.84, 0.94] 0.197
gp 32 330 10.41 0.33 0.83 [0.77, 0.89] 0.196
gp 33 253 11.51 0.51 0.92 [0.84, 0.99] 0.195
gp 34 126 11.37 1.36 0.90 [0.79, 1.00] 0.194

Individually these intervals overlap and clear 1.0 irregularly; that is why the cumulative rows above are the ones a claim rests on.

Tournament — the NCAA slice, per season and pooled

A single bracket is small and easy to over-read — the repo's own retraction #1 is a one-season 67-game slice that claimed the sigma was 13–15% overinflated and did not survive pooling. So spread and total are reported separately (a pooled margin error would confound two different quantities), each with its n, and the pooled figure sits beside the per-season ones.

slicemarketnrmse biaspred σactual sd
2021 spread 66 12.64 1.66 12.68 14.02
total 66 18.24 4.88 16.02 19.30
2022 spread 67 12.27 0.72 12.53 13.71
total 67 18.48 3.20 15.76 19.63
2023 spread 67 12.50 0.87 12.94 12.73
total 67 17.26 4.69 15.87 17.47
2024 spread 67 12.59 -0.19 12.57 15.65
total 67 19.36 3.21 16.12 21.48
2025 spread 67 10.03 -0.52 12.84 13.77
total 67 13.88 0.05 16.19 17.13
2026 spread 67 11.48 -0.18 12.62 15.59
total 67 14.87 2.19 16.19 18.03
Pooled spread 401 11.95 0.39 12.70 14.29
Pooled total 401 17.13 3.03 16.03 19.36

pred σ is the mean sigma the model emitted; actual sd is the spread the results actually had. They should be close — a predicted sigma far from the observed spread is the miscalibration, not the RMSE.

vs Vegas — gated on real odds

Lines are present in the store (42 rows). This block compares the model's margin and total to the de-vigged closing consensus once the line ingest has run for real and predictions join a line — today no edges exist yet.