Calibration and Brier scores: a good forecast vs a lucky one
Anyone can be right once. Calibration is how you catch who's right on purpose.
Chapter 3 left you with a problem. If one result can’t grade a probability, how do you ever know whether a forecaster (a model, a friend, a desk like ours) is any good? You can’t check one number against one game. You can check a hundred numbers against a hundred games. That check has a name: calibration.
The procedure fits in a sentence: gather every call someone stamped “70%,” and count how many won. About 70? Calibrated; their words mean things. Ninety? They’re sandbagging. Fifty-five? Their 70 is costume jewelry. Repeat for every bucket, plot it, done:
That chart (illustrative data; the live version of ours sits on the track record, updating with every graded slate) is the entire test. Dots on the diagonal: the forecaster’s numbers are a language. Dots below it: every stated probability was flattery, and you now know by how much to discount them.
One number for the whole record
Charts convince humans; sometimes you want one number for comparing. The Brier scoretakes every forecast, measures (probability minus outcome) squared, and averages. An outcome is 1 if it happened, 0 if it didn’t. Perfect certainty, perfectly placed, scores 0. Saying “50%” about everything forever scores exactly 0.25, the score of a shrug. Beating 0.25 means you know something.
Study the third bar, because it’s the one with your name on it. The loud caller and the honest caller won the same seven games. The loud one scored catastrophically worse, because the two misses they’d called at 95% cost (0.95)² each. The score’s square makes unearned confidence the most expensive habit in forecasting. Our own CFB model’s walk-forward Brier is 0.187across five seasons. Under the shrug line, meaningfully; miles from perfect; published anyway. That’s the posture this book keeps trying to teach.
Confidence is a cost you pay when wrong, not a flavor you add when sure.
Why should a tradercare? Because a calibrated source (any calibrated source, including your own journal once chapter 9 makes you keep one) turns prices into decisions. If you know a forecaster’s 70s are real 70s, then their 70 against a market’s 60 is information. If you don’t know that, it’s two strangers shouting numbers.
Frequently asked questions
- What is calibration in sports forecasting?
- A forecaster is calibrated when their stated probabilities match reality bucket by bucket: their 60% calls win about 60% of the time, their 80% calls about 80%. Calibration doesn't ask whether individual picks won. It asks whether the numbers meant what they said across many picks.
- What is a Brier score and what is a good one?
- The Brier score averages (forecast minus outcome) squared across every pick. Zero is perfect, and always saying 50% scores 0.25, so anything meaningfully under 0.25 beats a shrug. For context, our own college football model graded 0.187 across five walk-forward seasons. The score's teeth: it punishes unearned confidence far harder than earned confidence gains.
- How do I check if a sports forecaster is actually good?
- Ask for their whole record, then bucket it: collect every pick they called 70%, and check whether about 70% of that bucket won. Repeat per bucket. No complete record, no verdict. A highlight reel cannot be calibrated, only curated. That standard applies to us too, which is why our record is public.
Take the whole guide with you
Sign up for the daily desk email and all nine chapters arrive as a designed PDF. One email a day, our graded numbers, unsubscribe anytime.
Already on the list? Download the PDF. And as everywhere on Maiden: these are probabilities and mechanics, graded in public. Nothing here is betting advice.