Keeping Score Properly: How to Measure Whether You Are Actually Good at This
Win rate is a terrible measure of forecasting skill. Here is what to track instead — calibration, Brier score and a prediction journal — and how to read your own numbers without flattering yourself.
"I get about 70% of my calls right."
That sentence sounds like a claim about skill. It is almost content-free. If you only ever back heavy favourites, 70% is dreadful — those markets resolve in your favour far more often than that. If you specialise in genuine coin flips, 70% would be extraordinary.
Win rate on its own cannot tell those apart. Here is what can.
Accuracy vs. calibration
Accuracy is how often you are right. Calibration is whether your confidence matches your hit rate.
To check calibration, group your resolved calls by the confidence you assigned and compare each bucket to reality:
- Of everything you called at around 60%, did about 60% happen?
- Of everything at 80%, did about 80% happen?
- Of everything at 95%, did about 95% happen?
A perfectly calibrated forecaster's buckets match the line. Almost nobody's do at first — the overwhelmingly common pattern is that the high-confidence buckets underperform. Things you called at 90% land around 70% of the time. That is overconfidence, and it is the single most correctable flaw in forecasting, because the fix is arithmetic rather than knowledge: shade your extremes toward the middle.
The reverse pattern exists too, and it is rarer and gentler. If your 60% calls land 75% of the time, you are underconfident and leaving value on the table.
You need a decent sample before any of this means anything — a handful of calls in a bucket tells you nothing. Thirty per bucket is where it starts being real.
Brier score: one number that punishes false confidence
A Brier score measures the squared distance between what you predicted and what happened. Say you predicted 0.8 and it happened (outcome 1): the error is 0.2, squared is 0.04. Predicted 0.8 and it didn't happen: error 0.8, squared is 0.64.
Average that across all your calls. Lower is better. Zero is perfect. 0.25 is what you get by saying 50% to everything.
That last number is the benchmark that matters. If your Brier score is worse than 0.25, you would literally do better by refusing to have opinions. A large fraction of confident forecasters are, on measurement, in exactly that position.
Why squaring matters: it punishes confident wrongness disproportionately. Being 95% sure and wrong costs you enormously more than being 60% sure and wrong. This is exactly the incentive a forecaster should face, and it is why Brier score is the standard measure in forecasting research rather than win rate.
The prediction journal
Everything above requires data you have to create deliberately, because it is data about your reasoning and your platform cannot see that.
For each call, before it resolves, record five things:
- The question, in your own words.
- Your probability, as a number, written before you looked at the market price.
- The market price when you committed.
- Your one-sentence reason — the specific thing you believe that the crowd doesn't.
- What would change your mind.
That last field is the one people skip and the one that pays. It forces you to state in advance what evidence would falsify your view — which means when that evidence appears, you notice it, rather than absorbing it into an unchanged opinion.
When the market resolves, add one line: were you right, and were you right for the reason you gave?
Those are two independent questions with four combinations, and each means something different:
- Right, right reason. Repeatable. Do more of this.
- Right, wrong reason. Luck. The most dangerous outcome, because it teaches a bad habit while paying you for it.
- Wrong, right reason. Perfectly normal. A well-reasoned 70% call fails three times in ten. Change nothing.
- Wrong, wrong reason. The actually useful one. Go find out what you misunderstood.
Profit is a real measure, but a noisy one
Standom's leaderboard ranks by net profit in Stars, and that is the right primary measure for a competition — it captures both being right and sizing correctly, which is the complete skill.
But for self-assessment, profit alone is misleading over short windows. It mixes three things: how good your probabilities were, how well you sized, and how the variance broke. A month of bad luck on well-priced calls looks identical to a month of bad calls.
So track profit and calibration. Calibration tells you whether the process is sound. Profit tells you whether the process is paying yet. When calibration is good and profit isn't, the problem is usually sizing or market selection, not judgement — and those are much easier to fix than judgement.
Accuracy percentages need one more caution: a small sample of correct calls produces a spectacular-looking percentage that means nothing. Three-for-three is 100% and is worth exactly nothing as evidence. This is why a leaderboard should weight accuracy by volume rather than take raw percentages at face value, and why you should apply the same scepticism to your own numbers.
A review routine that takes ten minutes a week
Weekly: open every call that resolved this week. Mark each right/wrong and right-reason/wrong-reason. Read your "what would change my mind" notes and check whether you actually noticed when it did.
Monthly: bucket every resolved call by confidence and compare to reality. Compute your Brier score. Compare it to 0.25.
Quarterly: look for category patterns. Almost everyone has one domain where they are consistently overconfident — usually the one they care about most. Find yours and apply a standing discount to it.
What to do with what you find
Three common diagnoses and their fixes:
Overconfident in a specific category. Apply a blanket haircut: whatever you were going to say, move it ten points toward 50 in that category. Crude, effective, and you can refine it later.
Good calibration, poor profit. You are sizing wrong or playing markets where your edge is too small to survive the price you paid. Concentrate on the calls where your number and the market's differ most.
Bad calibration everywhere. Stop committing large positions entirely and run a paper month — record numbers, commit nothing, measure. It costs nothing but time and it is the fastest way to fix a broken process.
None of this is glamorous. It is also the only difference between a player who improves over a season and one who has the same year repeatedly while feeling like an expert. Write the number down before the result. Everything else follows from that one habit.
Keep reading
How Many Stars Should You Commit? A Sizing Guide
Good calls sized badly still lose. A practical framework for deciding how much of your balance a prediction deserves — and why the answer is almost always less than it feels.
5 min readWhen Prediction Markets Get It Wrong: The Limits of Crowd Wisdom
Prediction markets are good aggregators — but they are not oracles. There are specific, identifiable conditions where market prices are systematically wrong, and knowing them is a genuine edge.
5 min readRecency Bias: Why Your Last Match Always Feels More Important Than It Is
The team that won last week feels unstoppable. The team that lost three in a row feels broken. Recency bias is the most common error in sports prediction — and the easiest to identify once you know what to look for.