cbb·system

Calibration

How trustworthy a rating is, as a function of games played. Not every game counts equally: early in a season the model has barely seen a team, so both the mean (is it right about the game?) and the sigma (is it right about how right it is?) are still settling. The x-axis is the minimum games played by either team at kickoff, pooled across 0 seasons (–). Everything here is a walk-forward replay — each prediction was made from the ratings at the cutoff that preceded it, never from hindsight.

competition Women’s predictions scored 0 seasons pooled 0 mean gate 18 gp sigma gate 25 gp

The two statistics, and what each answers

familystatisticwhat it answers
meanrmse of the spread (+ bias) is the model right about the game? (bias vs noise)
sigmasd(z), z = (pred − actual)/σ; 1.0 = honest is the model right about how right it is?
probabilitybrier, log_loss depends on both the mean and the sigma

Mean — the spread error, across season progress

Not enough games-played levels carry a bucket to draw a mean curve.

Sigma — is the model honest about its own uncertainty?

Not enough games-played levels carry a bucket to draw a sigma curve.

Sigma only — does it order uncertainty at all?

The read that sd(z) above cannot make. sd(z) is joint with the mean — it mixes whether the mean is off with whether the σ is off — so it cannot attribute a departure from 1.0 to the σ. This table can: games are banded by the σ the model emitted, and each band's actual residual scale is measured. A σ worth emitting must rise with that column. A flat column means the σ is describing almost nothing; a falling column means its ordering is inverted against reality.

emitted σ bandnmean σ actual sd of residual

Cumulative rows, with bootstrap 95% intervals

The authoritative read. A row whose interval contains 1.0 is calibrated — the sigma is not measurably off; one whose interval excludes it is flagged. Read the interval, not the point: it resamples teams, because a team's games are correlated. The gate is the first calibrated row, so it is not identified at any threshold in this store — no cumulative row's interval covers 1.0, which is the finding, not a missing value.

cutoffnrmsebias sd(z)95% CIstatus

Per-bucket rows (the noise the cumulative read absorbs)

Show the per-games-played buckets
bucketnrmsebias sd(z)95% CIbrier

Individually these intervals overlap and clear 1.0 irregularly; that is why the cumulative rows above are the ones a claim rests on.

Tournament — the NCAA slice, per season and pooled

A single bracket is small and easy to over-read — the repo's own retraction #1 is a one-season 67-game slice that claimed the sigma was 13–15% overinflated and did not survive pooling. So spread and total are reported separately (a pooled margin error would confound two different quantities), each with its n, and the pooled figure sits beside the per-season ones.

slicemarketnrmse biaspred σactual sd

pred σ is the mean sigma the model emitted; actual sd is the spread the results actually had. They should be close — a predicted sigma far from the observed spread is the miscalibration, not the RMSE.

vs Vegas — gated on real odds

Lines are present in the store (42 rows). This block compares the model's margin and total to the de-vigged closing consensus once the line ingest has run for real and predictions join a line — today no edges exist yet.