Calibration
How trustworthy a rating is, as a function of games played. Not every game counts equally: early in a season the model has barely seen a team, so both the mean (is it right about the game?) and the sigma (is it right about how right it is?) are still settling. The x-axis is the minimum games played by either team at kickoff, pooled across 0 seasons (–). Everything here is a walk-forward replay — each prediction was made from the ratings at the cutoff that preceded it, never from hindsight.
The two statistics, and what each answers
| family | statistic | what it answers |
|---|---|---|
| mean | rmse of the spread (+ bias) |
is the model right about the game? (bias vs noise) |
| sigma | sd(z), z = (pred − actual)/σ; 1.0 = honest |
is the model right about how right it is? |
| probability | brier, log_loss |
depends on both the mean and the sigma |
Mean — the spread error, across season progress
Not enough games-played levels carry a bucket to draw a mean curve.
Sigma — is the model honest about its own uncertainty?
Not enough games-played levels carry a bucket to draw a sigma curve.
Sigma only — does it order uncertainty at all?
The read that sd(z) above cannot make.
sd(z) is joint with the mean — it mixes whether the mean is off with whether
the σ is off — so it cannot attribute a departure from 1.0 to the σ. This table can:
games are banded by the σ the model emitted, and each band's
actual residual scale is measured. A σ worth emitting must rise with that
column. A flat column means the σ is describing almost nothing; a falling column
means its ordering is inverted against reality.
| emitted σ band | n | mean σ | actual sd of residual |
|---|
Cumulative rows, with bootstrap 95% intervals
The authoritative read. A row whose interval contains 1.0 is calibrated — the sigma is not measurably off; one whose interval excludes it is flagged. Read the interval, not the point: it resamples teams, because a team's games are correlated. The gate is the first calibrated row, so it is not identified at any threshold in this store — no cumulative row's interval covers 1.0, which is the finding, not a missing value.
| cutoff | n | rmse | bias | sd(z) | 95% CI | status |
|---|
Per-bucket rows (the noise the cumulative read absorbs)
Show the per-games-played buckets
| bucket | n | rmse | bias | sd(z) | 95% CI | brier |
|---|
Individually these intervals overlap and clear 1.0 irregularly; that is why the cumulative rows above are the ones a claim rests on.
Tournament — the NCAA slice, per season and pooled
A single bracket is small and easy to over-read — the repo's own retraction #1 is a one-season 67-game slice that claimed the sigma was 13–15% overinflated and did not survive pooling. So spread and total are reported separately (a pooled margin error would confound two different quantities), each with its n, and the pooled figure sits beside the per-season ones.
| slice | market | n | rmse | bias | pred σ | actual sd |
|---|
pred σ is the mean sigma the model emitted; actual sd is the spread the results actually had. They should be close — a predicted sigma far from the observed spread is the miscalibration, not the RMSE.
vs Vegas — gated on real odds
Lines are present in the store (42 rows). This block compares the model's margin and total to the de-vigged closing consensus once the line ingest has run for real and predictions join a line — today no edges exist yet.