What 16 Rounds of Calibration Taught Us About Sports Prediction

Every assumption we had about what matters in basketball prediction was wrong. Here's what actually works.

6 min read

We Thought We Knew What Mattered

When we started building EdgeGoat's prediction engine, we had a mental model of what should matter in basketball. Rest days. Hot streaks. Rivalry games. The "eye test" stuff that every sports commentator talks about.

We were wrong about almost all of it.

Sixteen rounds of rigorous backtesting — hundreds of configurations tested against thousands of historical games — systematically dismantled our assumptions. What emerged was a prediction engine that looks nothing like what we expected. Here's what we learned.

Lesson 1: Minutes Are Everything

This sounds obvious in hindsight, but it wasn't where we started.

Early rounds focused on predicting per-game stat totals directly. How many points will Jayson Tatum score? What's his average? Adjust for matchup, home/away, done.

The breakthrough came when we shifted to predicting minutes first, then projecting stats as a function of playing time.

Think about it: if Tatum plays 38 minutes, he might score 28 points. If he plays 24 minutes because of foul trouble or a blowout, he might score 16. The per-minute production is relatively stable — it's the minutes that swing wildly.

We spent five sub-rounds (R15 through R15e) just optimizing minutes prediction:

  • R15: Introduced decay weighting and outlier trimming — MAE dropped from 5.17 to 4.80
  • R15c: Added 3-game recency blend and median mixing — down to 4.77
  • R15d: Multi-pass projection accounting for blowout probability — 4.75
  • R15e: Fine-tuned lambda decay and margin thresholds — 4.75

Lesson 2: 96% of Error Is Irreducible

We did a full decomposition of our prediction errors, splitting them into two components:

  • Between-player bias (4%): Systematic tendencies to over/under-predict certain players. This is fixable.
  • Within-player randomness (96%): The same player, same matchup, same situation — wildly different outcomes game to game. This is noise.

We tried per-player bias correction. Walk-forward estimates tracking each player's projection error over time, then adjusting future projections. It sounded perfect in theory.

It made predictions worse.

The estimates were too noisy. With maybe 30-40 games of data per player per season, the bias estimates bounced around so much that "correcting" them just added noise. The model was better off ignoring individual player tendencies and trusting the base rates.

This was humbling. No matter how sophisticated your model, basketball players are humans making split-second decisions. The randomness isn't a flaw in your model — it's a feature of the sport.

Lesson 3: The Signals Everyone Talks About Don't Work

Here's a list of things that every basketball analyst considers important, that we tested extensively, and that produced zero improvement or actively hurt accuracy:

Rest and back-to-back games

  • Tested: RestB2B flag adjusting projections for teams on second night of a back-to-back
  • Result: MAE increased by 0.037
  • Why: Modern NBA load management has largely eliminated the back-to-back disadvantage. Teams rest players, adjust rotations, and the effect is already priced into minutes projections.

Hot/cold streaks (HMM regimes)

  • Tested: Hidden Markov Model detecting "hot" and "cold" player states
  • Result: Zero improvement
  • Why: The hot hand is real but tiny and unpredictable in timing. By the time you detect a regime, it's already shifted. The base rate projection is just as good.

Game importance

  • Tested: Playoff implications, rivalry multipliers, elimination games
  • Result: No measurable effect
  • Why: Both teams are equally motivated. The intensity might increase, but it doesn't consistently favor one side or change scoring patterns predictably.

Schedule density

  • Tested: Games played in last 7/10/14 days as a fatigue signal
  • Result: Noise
  • Why: Similar to rest — modern sports science and load management handle this. The signal is already in the minutes projection.

Changepoint detection (PELT algorithm)

  • Tested: Statistical changepoint detection to identify when a player's baseline production shifted (role change, return from injury, etc.)
  • Result: No improvement over simple exponential decay
  • Why: Our lambda decay (weighting recent games more heavily) already handles recency. The changepoint model adds complexity without adding signal.

Opponent pace adjustment

  • Tested: Adjusting projections based on opponent's pace
  • Result: Already captured by pace blend
  • Why: The pace blend signal (combining both teams' pace tendencies) already accounts for this. Adding a separate opponent pace adjustment was double-counting.

Lesson 4: The Boring Stuff Wins

Here's what actually matters, ranked by impact:

Tier 1: Critical (removing these is catastrophic)

  1. Pace blend (MAE +2.3 if removed) — Combining both teams' historical pace into a game pace estimate. This is the single most important signal in the entire pipeline.
  2. Strength of schedule (Win% -3.1pp if removed) — Not all wins are equal. A team that beat the Celtics is different from one that beat a tanking team.
  3. Chemistry/lineup effects (Win% -1.8pp if removed) — How the specific lineup on the court tonight performs together, not just individually.

Tier 2: Important

  1. Venue weighting — Home/away splits, especially for player props
  2. Team defense — Matchup-specific defensive adjustments
  3. Minutes-weighted rates — Per-minute production rather than per-game averages
  4. Median blend — Mixing mean and median to handle outlier games

Tier 3: Marginal but real

  1. Coach embeddings — Encoding coaching tendencies as vectors
  2. Second-pass blowout adjustment — Reducing minutes for starters in projected blowouts
  3. Bayesian sigma priors — Shrinking variance estimates toward league averages for players with small sample sizes

None of these are sexy. Nobody's writing articles about "pace blend" or "exponential decay lambda." But they're what separates a 55% model from a 65% model.

Lesson 5: Chemistry Is Not Double-Counting

This one deserves its own section because we nearly made a catastrophic mistake.

Our chemistry signal measures how players perform together vs. their individual averages. When Tatum and Brown are on the court together, do they produce more or less than the sum of their individual rates?

A reasonable concern: isn't this just measuring lineup quality? If Tatum plays with the starters, of course the team scores more. You're just double-counting the fact that the other starters are also good.

We tested an "absence-only" model — only adjusting when key players are out, not when they're together. Win accuracy dropped from 66.8% to 64.5%. That's worse than disabling chemistry entirely.

The positive lift for present teammates isn't double-counting — it's a lineup-quality signal. Historical averages span all lineups a player has been in. Chemistry adjusts toward tonight's specific lineup output. Different teammates create different spacing, different passing lanes, different opportunities. The model needs to know who's on the court together, not just who's on the roster.

Lesson 6: Your Model Will Lie To You

Overfitting is the constant enemy. Some war stories:

The prop-optimized mode disaster: We built a special configuration that optimized specifically for player prop accuracy instead of game-level accuracy. It improved PropMAE by a hair. It destroyed win accuracy by 6.2 percentage points. The model was contorting itself to match individual stat lines while losing all understanding of game context.

The team bias correction trap: Round 11 introduced per-team bias corrections — if we systematically over-predict the Lakers' scoring, adjust down. It worked in the backtest. Then we realized it cost too much compute for live projections and disabled it. When we re-tested later with more data, the improvement had vanished. Classic overfitting to the training window.

The nonlinear margin curve mirage: We tested nonlinear blowout adjustments — polynomial curves for how minutes change with game margin. Identical MAE to the simple linear model. The complexity added nothing. Always start simple.

The Meta-Lesson

Sixteen rounds of calibration taught us one overarching lesson: sports prediction is a game of small, boring edges compounded over hundreds of decisions.

There's no magic signal. No secret formula. No single feature that unlocks 80% accuracy. It's pace blend plus SOS plus chemistry plus team defense plus minutes modeling plus calibrated probabilities plus smart bet selection — each contributing 1-3 percentage points, stacking up to something meaningful.

The exciting breakthroughs (ML ensembles! neural networks! embeddings!) add real value. But they add it on top of the boring foundation. A neural network trained on bad pace estimates is worse than a weighted average with good pace estimates.

Get the fundamentals right. Measure everything. Trust the backtest over your intuition. And know when you've hit the floor.


We've run over 200 configurations across 2 leagues. The biggest lesson: respect the base rates, and don't fall in love with complexity.

See the model on today's slate

Every line the model prices, with its probability next to the sportsbook's.