What a Confidence Score Should Actually Mean (and Why Most Are Wrong)
Most betting models show a confidence number that has no relationship to how often those picks win. We tested ours against 1,634 real NFL closing lines and rebuilt it around measured results.
The short version: a confidence score should be the measured hit rate of picks like the one in front of you. If the number says 57, picks carrying that number should win about 57% of the time in data the model never saw. Most published confidence scores fail that test, including an earlier version of ours. Here is what we found and how we fixed it.
The test that matters
Any model can print a number next to a pick. The question is whether that number survives contact with results.
The honest test has two parts:
We ran that test on 1,634 NFL player props from 2025, graded at real DraftKings and FanDuel closing prices.
What we found
Our model's raw probability carried almost no information against the closing line. The technical measure is a logistic slope, where 1.0 means fully honest and 0 means no information at all. Ours came in at 0.07.
In plain terms: when the model said a prop had a 59% chance, it hit 49%. When it said 69%, it hit 50%. The confidence bands ranked backwards from week to week — picks badged 55 to 57 hit 51%, picks badged 58 to 61 hit 60%, and in the other half of the season the ordering flipped.
That is not a model with a small calibration problem. That is a number with nothing behind it.
This is the norm, not the exception. The same pattern showed up in our baseball and basketball models when we tested them the same way. A probability that emerges from a projection model is not automatically a probability that predicts anything, and almost nobody publishes the test.
Why models fail this test
The projection is usually fine. The probability is the problem.
A prop model typically projects a stat — say 61 receiving yards — then converts the gap between that projection and the line into a probability using a distribution. That conversion assumes your projection error is well-behaved and that you know the true variance. Neither is true in football, where 17 games is a small sample and role changes happen weekly.
More fundamentally, the closing line already contains everything the market knows. To have real information beyond it, your projection must know something the market does not. Usually it does not. It just disagrees, and disagreement is not edge.
We measured that directly. Props where our projection sat far from the book's line performed worse, not better:
| Projection vs. line | Hit rate | Return |
|---|---|---|
| Within 10% | 55.6% | +5.9% |
| 10-20% apart | 50.8% | -3.9% |
| 20-35% apart | 45.9% | -12.0% |
The further we strayed from the market, the worse we did. That is the opposite of what an edge looks like.
How we rebuilt it
We stopped showing the model's opinion and started showing the measured record.
The engine still picks a side — that part works, and it is what selects which props are worth looking at. But the number displayed is now the measured hit rate of picks like this one: same prop market, same side, chosen the same way, graded at real closing prices. Thin categories are shrunk toward the broader rate so a 6-bet sample cannot produce a wild number.
The out-of-sample result, fitting on weeks 1 through 5 and scoring weeks 6 through 9:
| | Displayed confidence | Actually hit |
|---|---|---|
| Rebuilt score | 56.4 | 55.4% |
Within one point, on data the model never saw. That is what a confidence score is supposed to do.
For comparison, the old approach displayed numbers up to 69 on picks that hit 50%.
What honest numbers look like
They are lower than you would like.
At real closing lines, our measured rates land between roughly 49% and 57% depending on the market and side. Break-even at -110 is 52.4%. The sharpest professional prop bettors operate in the 55% to 57% range.
If a service advertises 70% hit rates against closing lines, one of two things is true: they are measuring against something other than the closing line, or they are not measuring at all. We made that mistake ourselves. An earlier version of our own backtest, run against constructed lines rather than real ones, showed 61.8%. Against real closing lines the same model returned -0.3%.
The constructed-line result was not a lie. It was measuring a different and much easier question: can we beat a naive median? Yes. Can we beat the sportsbook? That is the question that pays.
What to ask any betting model
Ours: at 300 live graded NFL picks, if displayed confidence misses the realized hit rate by more than 3 points, the layer failed and we will say so.
See the measured number
Every NFL prop on BetBlum shows the measured hit rate of picks like it, graded against real closing prices, with the sample size behind it visible.
Start your free 48-hour trial and see honest numbers on every prop.
Frequently asked questions
What should a betting model's confidence score mean?
It should mean the measured hit rate of past picks like this one. If a model displays 57, picks carrying that number should win about 57% of the time in results the model did not train on. If the displayed number and the realized rate diverge, the score is decoration rather than information.
Why do most betting model confidence scores not work?
Because they report the model's opinion of itself rather than a measured track record. A model can output 68% on a prop and have those picks win 50% of the time. Without grading past picks against real closing lines, a confidence score is untested by construction.
How do you test whether a confidence score is honest?
Split the data by time. Fit the score on an early period, then score a later period the model never saw. If the displayed confidence and the realized hit rate track within a few points out of sample, the number means something. If they diverge, it does not.
What confidence level is realistic for NFL player props?
At real closing lines, the sharpest professional prop bettors operate around 55% to 57%. Break-even at standard -110 pricing is 52.4%. Any model advertising 70% hit rates against closing lines is either measuring against something other than the closing line or not measuring at all.
See this week's NFL props, measured
Every NFL prop with the measured hit rate of picks like it, the line-shopped number, and the honest expected value. 48-hour free trial, no credit card required.
Start Free Trial