From Spreadsheet Vibes to 65% Win Rate: How We Built a Prediction Engine
The story of turning crude averages into a 16-stage machine learning pipeline — one painful calibration round at a time.
6 min read
The First Prediction Was Embarrassing
The first time EdgeGoat tried to predict a basketball game, it used a method I'm generously going to call "weighted averages." Take a player's last 10 games, average their stats, adjust vaguely for home/away, and call it a projection.
It was basically a spreadsheet with a nicer UI.
And honestly? It wasn't terrible. It got about 55% of games right on the moneyline, which is... roughly what you'd get picking the home team every time. Not exactly edge.
But the infrastructure was in place. I could pull game data, run projections, and compare them to actual results. The question was: how far could I push the accuracy?
The Calibration Grind
What followed was 16 rounds of calibration — each one a cycle of "hypothesize, implement, backtest, measure, keep or discard." The commit history reads like a lab notebook written by someone slowly losing their mind:
Round 1-4: Basic calibration. Getting the fundamentals right — league averages, venue adjustments, seasonal trends. Win accuracy crept from 55% to around 59%.
Round 5: "props plateau hit..." — the moment I realized that player prop predictions weren't improving no matter what I tweaked at the game level. The projections needed to understand individual players, not just teams.
Round 8-10: The engine got serious. Play-by-play data integration. Pace analysis from actual game clock events. Per-player rate calculations based on minutes context instead of raw totals. Win accuracy jumped to ~62%.
Round 10: "my best guess" — sometimes that's the commit message you write at 2 AM when you've tried 40 configurations and picked the one that backtests best, knowing full well you might be overfitting.
Round 14-15: Minutes prediction optimization. This was the breakthrough I didn't expect. Getting minutes right turned out to be the key to everything else — if you know a player is going to play 34 minutes instead of 28, every stat projection changes. We got minutes MAE down to 4.75.
Round 16: The final push. Team defense adjustments, per-stat rebounding caps. NBA win accuracy landed at 64.8%, with a prop MAE of 2.43.
Each round took hours. Some took days. You change one parameter, re-run the backtest across hundreds of games, wait, check the metrics, and decide if that 0.1% improvement is real signal or noise. Then you do it again.
What The Numbers Actually Mean
Let me break down the metrics we obsess over, because they tell the story:
MAE (Mean Absolute Error): 15.1 points — On average, our total score prediction for an NBA game is off by about 15 points. That sounds like a lot until you realize the average NBA game total is around 220. We're within 7% of the actual outcome.
Win Accuracy: 64.8% — When we predict which team will win, we're right almost two-thirds of the time. For context, the home team wins about 57% of the time in the NBA, so we're adding meaningful signal beyond just home court.
Prop MAE: 2.43 — For player props (points, rebounds, assists, etc.), we're off by about 2.4 units on average. When a line is set at 22.5 points and we project 25.1, that 2.6-point gap is meaningful edge.
ATS (Against The Spread): 77.1% — This one surprised us. Against the spread, which is supposed to be the hardest thing to beat, we're hitting at an absurdly high rate. We're cautious about this number — it's on our backtest set, and real-world performance will be lower. But even half that rate would be extraordinary.
The Signals That Actually Matter
Through all those calibration rounds, I tested dozens of signals. Here's the hierarchy of what actually moves the needle:
The essentials (removing these destroys accuracy):
- Pace blend — combining team pace tendencies with opponent pace. Remove it and MAE jumps by 2.3 points.
- Strength of schedule — adjusting for opponent quality. Remove it and win accuracy drops 3.1 percentage points.
- Player chemistry — how lineups perform together vs. their individual averages. Worth 1.8 percentage points of win accuracy.
The important (meaningful improvement):
- Venue weighting — home/away matters more for props than most people think
- Team defense — accounting for matchup-specific defensive impact
- Minutes-weighted rates — using per-minute production instead of per-game
- Median blend — mixing mean and median of recent performances to handle outlier games
The surprisingly useless (tested and discarded):
- Rest/back-to-back — everyone thinks this matters. In our data, it actively hurt accuracy.
- Schedule density — how many games in the last week. No signal.
- Game importance — playoff implications, rivalry games. No measurable effect.
- Opponent pace adjustment — already captured by the pace blend.
- HMM regime detection — fancy Hidden Markov Models to detect hot/cold streaks. Zero improvement.
That last one stung. I spent real time implementing Hidden Markov Models because it sounded so cool and so theoretically sound. The backtest came back flat. Sometimes the unsexy weighted average beats the fancy statistical model.
The Pipeline Today
What started as "weighted averages in a function" is now a 16-stage pipeline:
- Minutes prediction — how many minutes each player will get
- Minutes ensemble (ML) — LightGBM correcting the statistical minutes projection
- Player embeddings — contrastive learning representations of player style
- Temporal convolution — neural network processing game log sequences
- Prop projection — per-stat predictions for every player
- Props ensemble (ML) — stacking model refining prop predictions
- Team scoring — aggregating player projections into team totals
- Team ensemble (ML) — correcting team-level predictions
- Game projection — combining teams into margin/total predictions
- Second pass — adjusting minutes for blowout/overtime probability
- Game ensemble (ML) — final game-level correction
- Selection — picking candidates with edge against the sportsbook
- Quantile regression — probability distributions, not just point estimates
- Calibration — isotonic regression ensuring probabilities are well-calibrated
- Strategy — optimizing bet type, sizing, and combinations
- Final report — aggregate results and confidence metrics
Each stage builds on the last. Each was born from a specific calibration round where I realized the previous approach was leaving accuracy on the table.
The Emotional Rollercoaster
Building a prediction engine is an exercise in managing your own psychology. Here are the stages:
- Naive optimism: "I'll just add this one feature and accuracy will jump 5%"
- Reality check: It improved by 0.3%. Maybe.
- Plateau despair: "I've tried everything, nothing moves the needle" (the "props plateau hit..." commit)
- Breakthrough euphoria: You find the one signal that unlocks the next level (minutes prediction was ours)
- Overfitting paranoia: "Is this real or am I just fitting to noise?" (the "less overfitting" and "overfitting guard" commits)
- Acceptance: You realize that 96% of prediction error is irreducible randomness — players are humans, not robots — and the remaining 4% is where you earn your edge
That last point deserves emphasis. We did a full error decomposition on our minutes predictions and found that 96% of the variance is within-player randomness. The same player, in the same situation, will play anywhere from 28 to 38 minutes on a given night. No model can predict that. The 4% that's between-player bias — systematic over/under-projection — is the only part we can fix.
Sports prediction is fundamentally a game of small edges. The sportsbooks have armies of quants. The lines are efficient. If you're looking for a magic formula that picks winners 80% of the time, it doesn't exist. But 65%? With good bankroll management and smart bet selection? That's a real edge. And that's what we're chasing.
What's Next
The engine is never done. There are always more signals to test, more models to try, more data to incorporate. We're currently exploring:
- Spatial shot chart embeddings for predicting 3-point props
- Live in-game model updates as the game unfolds
- Cross-sport model transfer for newer leagues with less data
But the core lesson from 16 rounds of calibration hasn't changed: the boring stuff matters most. Get minutes right. Get pace right. Respect the base rates. And don't fall in love with a signal just because it's clever — fall in love with it because the backtest says it works.
EdgeGoat's prediction engine was built one calibration round at a time. We're still calibrating.