StrategyProbability

The Only Honest Way to Know If You're a Good Predictor

Most people think they are better predictors than they are. Here is the one metric that separates genuine skill from lucky streaks — and how to use it to actually improve.

Standom EditorialWritten by the Standom Editorial team··5 min read

Most people who have been predicting for a while believe they are reasonably good at it. They have a few memorable wins — the call no one else made, the team they backed at a long price that came through. The losses are harder to recall. The question is: is that self-assessment accurate, and if not, how would you know?

The honest answer is that you would not know without tracking the right thing in the right way. Prediction has a specific, measurable definition of quality. Most participants never measure it, which means they are operating on self-selected memory rather than evidence.

Why Accuracy Alone Is Not the Measure

"I got 7 out of 10 right" is not a useful accuracy metric. It tells you nothing about whether your performance was skillful or lucky.

Consider: if you predict the favourite in every match, you will be "right" more often than not, because favourites win more often than not. That 70% win rate reflects the structure of the markets you chose, not any analytical contribution you made. You could have picked randomly from among the favoured side in every event and achieved the same rate.

The metric that separates skill from structure is calibration — the alignment between how confident you said you were and how often events at that confidence level actually occurred.

If you made fifty predictions at 70% confidence, approximately 35 of them should have come true. Not 35 exactly — variance is real, and 32 or 38 is consistent with a well-calibrated predictor. But if only 24 came true, you were systematically overconfident at that probability level. If 44 came true, you were systematically underconfident. In either case, you have something to fix.

What Calibration Measures and What It Does Not

Calibration is a measure of honesty about uncertainty. It does not tell you whether you are finding events others have missed. It tells you whether the probabilities you state reflect reality.

A well-calibrated predictor who picks moderately priced markets correctly at the right rate is not the same as a skilled edge-finder who identifies mispriced markets. Both are useful, and the best predictors have elements of both, but they are different skills and they improve in different ways.

Calibration improves by learning to state probabilities that reflect your actual uncertainty rather than your preferred outcome. Most people systematically overstate certainty — they say 80% when the honest answer is 65%, because 80% feels more decisive and confident. The practice of explicitly considering: "what are the three or four things that could make me wrong, and how likely are each of them?" tends to pull overconfident estimates toward reality.

Edge-finding improves by identifying where market prices systematically diverge from actual probabilities — which requires first knowing what actual probabilities look like, which brings you back to calibration.

How to Actually Track Your Performance

The minimum viable tracking system requires three things: a record of what probability you assigned, what position you took, and what happened.

Write the number before the event resolves. Not after — after is when hindsight quietly overwrites your memory of what you thought. The written number is the thing that cannot be revised. If you said 65% and it lost, that is your actual data point, and it will look different six months later from 200 data points than it looks from the one event you just lost.

The useful output is a table of predictions grouped by probability band — all your 60-70% calls in one group, all your 70-80% calls in another — and the actual resolution rate in each band. This is your calibration curve, and it is the most honest mirror the game can show you.

If your 70-80% band resolves at 60%, you know you are systematically overconfident at high-confidence predictions. That is a specific, fixable error. If your 50-60% band resolves at 70%, you know you have genuine information at that confidence level but are understating your certainty — which means you are undersizing your positions and leaving value behind.

The Pattern That Separates Good Predictors From Confident Ones

There is one pattern that appears consistently in people who improve over time, and almost never in people who stagnate: they treat incorrect predictions as information.

The natural response to being wrong is to explain it away — the umpiring call, the weather, the freak result. These explanations are sometimes accurate. But the habit of always finding a reason why the loss does not count is the habit that prevents calibration from ever improving.

The useful question after a wrong call is not "why did I lose?" but "was my probability estimate accurate given what I knew at the time?" If you said 75% and it lost, that is not automatically a failure — 75% events lose 25% of the time. If you said 75% and you honestly assess that the true probability was closer to 50%, that is a failure of a specific kind: overconfidence, probably driven by anchoring to the narrative or the name rather than the underlying variables.

That distinction — between a correctly-estimated probability that resolved against you and an incorrectly-estimated probability — is the thing that makes the difference. Every loss is either variance or error. Knowing which one it was is the only honest way to improve.

The Number to Care About After a Season

At the end of a prediction season, ask three questions:

Did my predicted probabilities match resolution rates? This is calibration.

Did I make calls where the market disagreed with me, and did those calls perform better or worse than expected? This is your edge, or the absence of it.

Did my high-confidence calls perform better than my low-confidence calls? If they did not, your confidence is not correlated with accuracy — which means you are not actually more certain when you think you are. That is a calibration failure of a different kind, and a common one.

These three questions, answered honestly with actual data, tell you more about your prediction quality than any single memorable win or loss. The memorable moments are not the data. The pattern across hundreds of calls is the data.

One year of tracked predictions, reviewed honestly, is worth ten years of untracked ones. The tracking is not administrative overhead. It is the thing that makes experience into skill rather than just into habit.