What does “well calibrated” actually mean?
Take every forecast you have ever made with a stated probability near 70%. Collect the outcomes. If about seven in ten of those events happened, your 70% means 70%, and you are calibrated at that level. If nine in ten happened, you were underconfident: your 70% was really a 90%. If four in ten happened, you were overconfident, and anyone acting on your numbers was systematically overpaying.
The definition contains the answer to the most common objection. You cannot evaluate a single probabilistic forecast, ever. A 70% call that fails is not evidence against the forecaster, and a 70% call that succeeds is not evidence for them. Only the collection can be judged — which is why forecasters who never publish their record are not making a checkable claim, and why a record with the losses removed is not a record.
This also explains a phrase that sounds like a paradox and is not: a well-calibrated forecaster should be wrong 30% of the time on their 70% calls. Being wrong at the right rate is the evidence that the number is meaningful. A source whose 70% calls come in 95% of the time is not better than one whose come in 70% of the time — it is a source that should have said 95%, and whose numbers you cannot size a position from.
Why isn’t accuracy the right measure?
Accuracy is the share of calls that were right, which requires turning a probability into a yes-or-no by picking a threshold. Two things go wrong immediately. The threshold is arbitrary and usually undisclosed. And accuracy is dominated by the base rate — how often the event happens at all.
So an accuracy figure is uninterpretable without three companions: the base rate of the event, the threshold used, and the sample size. “90% accurate” with none of those is not a claim, it is a decoration. Ask what the same method would score by always predicting the more common outcome — that is the baseline, and beating it is the minimum bar.
The concept that accuracy is groping toward is discrimination, sometimes called resolution: the ability to separate the cases that will happen from the cases that will not, by giving them different probabilities. Calibration and discrimination are independent, and you need both.
- A forecaster who says “the base rate” to every question is perfectly calibrated and has zero discrimination. Useless, but honest.
- A forecaster who separates outcomes sharply but always states 95% is well discriminated and badly calibrated. Useful, once you correct the numbers.
- Calibration can be repaired after the fact by re-mapping stated probabilities onto observed frequencies. Discrimination cannot: it requires information the forecaster either has or does not have.
That asymmetry is the practical reason to care about calibration first. It is the cheap half of the problem, and a system that has not fixed the cheap half is telling you something about how carefully it was built.
How do you read a reliability diagram?
A reliability diagram — also called a calibration curve — is the standard picture. Group all forecasts into bins by their stated probability. For each bin, put the mean stated probability on the horizontal axis and the observed frequency of the event on the vertical axis. Plot the points. A perfectly calibrated forecaster’s points lie on the 45° diagonal.
- Points below the diagonal: the event happened less often than stated. Overconfidence. If you were buying at those prices, you were overpaying.
- Points above the diagonal: the event happened more often than stated. Underconfidence — a real edge, and one that is invisible if you only look at accuracy.
- A curve that is flatter than the diagonal is the classic signature of a forecaster who does not commit: too few extreme calls, everything squeezed toward the middle.
- A curve steeper than the diagonal is the opposite, and the more dangerous one: extremes stated more often than they occur.
The part that is routinely omitted is the bin counts, and without them the chart cannot be read at all. A point at 0.95 built from twelve observations will land wherever chance puts it; the same point built from ten thousand is a measurement. Any calibration chart worth trusting publishes the sample size per bin — as a column, a histogram beneath, or error bars. If it does not, assume the impressive-looking bins are the small ones.
Two more things to check. Are the bins fixed in advance, or chosen after seeing the data? And does the chart cover the entire forecast history, or a window that happens to look good? Both choices can manufacture a diagonal out of nothing.
Is there a single number for this?
There is, and the standard one is the Brier score: the mean squared difference between the stated probability and the outcome coded as 1 or 0. Say 0.70 and the event happens, and you contribute (1 − 0.70)² = 0.09. Say 0.70 and it does not, and you contribute 0.49. Lower is better; zero is perfection; 0.25 is what you get by saying 50% to everything.
Its useful property is that it decomposes into three terms: reliability (how far the calibration curve sits from the diagonal — lower is better), resolution (how much your forecasts vary from the base rate in the right direction — higher is better), and uncertainty (the irreducible variance of the event itself, which belongs to the question, not to you). The decomposition is what lets you say whether a mediocre score comes from bad calibration or from a hard question.
The main alternative is the logarithmic score, which punishes confident errors far more aggressively — an outright 0 assigned to something that happens carries an infinite penalty. That severity is a feature if you care mainly about catastrophic overconfidence, and a nuisance if a single bad call can dominate a small sample.
注 A score is meaningless without a baseline. Always compare against the base-rate forecast — the same score computed for a forecaster who simply states the historical frequency every time. The difference is the skill score, and it is the number that survives scrutiny.
Why is a calibrated 55% worth more than a miscalibrated 90%?
Because a probability is not a conclusion, it is an input to a decision, and the decision needs the number to mean what it says. On a market quoted 0–1, the arithmetic is direct: your estimate minus the price is your edge per share, before costs. That subtraction is only valid if your estimate is calibrated.
The second-order damage is worse than the first. Any sensible position-sizing rule scales with the stated edge — the Kelly criterion is the explicit version, and most discretionary rules are a rough approximation of it. If the stated edge is inflated, size is inflated in exactly the same proportion, so an overconfident system bets most heavily precisely where it is most wrong. Miscalibration does not shave a little off returns; it changes the sign of the expectation and then multiplies it.
And there is a cost floor underneath all of it. A 5-point edge does not survive a 5-point round-trip cost. This is why the honest way to present a measured edge is next to the cost of acting on it, not on its own.
How do you audit a published track record?
Six questions, in order. They are not hard to ask, and most published records fail at least two of them.
- Is the sample defined in advance? Every forecast emitted, or only the ones still convenient at reporting time?
- What happened to the unresolved cases? Forecasts whose outcome was never retrievable must be published as a share, not silently dropped or counted as failures.
- Is the number decomposed? A single headline rate hides which subset produces it.
- Are the bad bands shown? A calibration table that stops before the ranges where the forecast underperforms is a marketing chart.
- Is the measurement moment specified? A probability read at one timestamp is a different quantity from one read at another, and the difference can be the entire effect.
- Is the cost of acting subtracted, or at least stated alongside?
We hold our own register to those questions, which is the only reason it is worth pointing at. The public register publishes the hit rate on 136,408 verified five-minute slots, together with the 42.54% of emitted slots whose outcome could never be retrieved — because a rate quoted without its coverage cannot be compared from one month to the next. It also publishes the decomposition that undercuts the headline, separating the slots confirmed by a second read from those never confirmed at all — the second group is set aside rather than scored, and printing it is the only way the first number stays readable.
The calibration table itself is the more interesting object, and it is the direct analogue of a reliability diagram for market prices rather than for a forecaster. Each row takes the market price read partway through a five-minute slot and compares it with how those slots actually resolved.
| T+180 秒时的价格 | 时段数 | 如此结算的比例 | 差值(点) |
|---|---|---|---|
| 0.00 – 0.40 | 1,562 | 32.01% | −3.75 |
| 0.40 – 0.50 | 4,681 | 43.43% | −2.74 |
| 0.50 – 0.60 | 11,728 | 55.58% | +0.50 |
| 0.60 – 0.70 | 15,239 | 66.82% | +1.73 |
| 0.70 – 0.80 | 19,108 | 78.11% | +3.04 |
| 0.80 – 0.90 | 25,399 | 88.47% | +3.23 |
| 0.90 – 0.95 | 20,312 | 95.16% | +2.47 |
| 0.95 – 1.00 | 38,031 | 98.83% | +1.13 |
Read it the way you would read a reliability diagram. The rows below 0.50 have a negative gap: those sides resolved less often than they were priced, which is the market pricing them richly and our signal being the one that is wrong there. The widest positive gap sits in the 0.70–0.90 range — and it is of the same order of magnitude as the roughly 3.5% cost of taking liquidity, which is precisely why it is presented as a measurement and never as a return. The sample sizes are in the table because, as above, a calibration figure without its bin count is not a figure.
What does calibration not give you?
It is a necessary condition, not a sufficient one, and the failure modes are worth naming.
- Calibration without sharpness is empty. Stating the base rate every time is perfectly calibrated and carries no information, which is why the Brier decomposition keeps resolution as a separate term.
- Calibration does not transfer between domains. Being calibrated on five-minute crypto slots says nothing about being calibrated on elections, and a record built on one should never be quoted for the other.
- It is backward-looking, and regimes change. A curve measured in one volatility regime can misstate the next one, which is why coverage and dates belong on the chart.
- Small samples lie in both directions. Bins with few observations will wander off the diagonal by chance alone, and the temptation to explain that wandering is strong.
- It says nothing about costs, capacity or variance. A calibrated edge that is smaller than the spread, or that only exists in size you cannot get filled in, is a true statement you cannot use.
What calibration does give you is the ability to check a claim rather than trust it. That is a lower bar than being right, and a much more useful one: it is the difference between a number you can put into a decision and a number you have to believe.
If you want the price side of the same picture, how prediction markets work explains where a 0–1 price comes from and why it is not quite a probability.
常见问题
- How many forecasts do I need before calibration means anything?
- Enough per bin, not enough in total — a thousand forecasts concentrated in two bins tells you about two bins. Treat a bin with a handful of observations as unmeasured, and expect the extreme bins to need the most data because their frequencies are the hardest to pin down.
- Is a market price calibrated?
- It is an empirical question with a per-market answer, which is why it is worth measuring rather than assuming. Our own measurement of short-dated crypto markets is published in full on the public register, including the price bands where the market prices the outcome better than our signal does.
- Can a forecast be recalibrated after the fact?
- Yes — mapping stated probabilities onto their observed frequencies is standard practice, and it is why calibration is the cheaper of the two problems to fix. The catch is that the mapping is estimated on past data and can itself be overfitted, so it needs a held-out sample like any other fitted curve.
这些指南以教学为目的。EdgeMarket 发布的是测量结果与市场机制说明,不是投资建议,这里也没有任何交易建议。