Teaching a GPU to Watch Basketball: Our ML Pipeline
How we went from statistical projections to LightGBM ensembles, neural networks, and contrastive learning embeddings — and the GPU crashes along the way.
7 min read
The Limits of Statistics
For the first couple weeks of EdgeGoat's life, the prediction engine was pure statistics. Weighted averages, exponential decay, Bayesian priors, pace adjustments. No machine learning. No neural networks. Just math.
It got us to about 62% win accuracy. Respectable. But we kept finding patterns in our errors that the statistical model couldn't capture. Players whose projection was consistently off in the same direction. Matchup effects that simple adjustments couldn't model. Game-level dynamics that emerged from the interaction of all the player projections.
The statistical model was leaving signal on the table. We could see it in the residuals. We just couldn't extract it with more math.
Time to bring in the machines.
Layer 1: LightGBM Ensembles (The Workhorses)
Our first ML models weren't deep learning. They were gradient-boosted trees — specifically LightGBM — trained to correct the statistical pipeline's predictions.
The key insight: don't replace the statistical model, stack on top of it. The stats pipeline produces a projection. The ML model looks at that projection along with dozens of contextual features and outputs a correction. The final prediction is a blend of both.
We built four ensemble models:
Minutes Ensemble
The statistical model projects minutes. The ensemble looks at 16 features — team share of minutes, recent variability, opponent pace, game spread, and more — and adjusts. The most important feature, at 75% of total importance? team_share — what fraction of the team's minutes this player typically gets.
The ensemble uses an alpha blend of 0.95, meaning it heavily trusts its own correction over the base projection. That's how confident the training made it.
Props Ensemble
Per-stat models for points, rebounds, assists. 43 features each — 24 statistical features, 16 from player embeddings, and 3 from the temporal convolutional network. The alpha is capped at 0.70 with a bias penalty of 2.0, because we learned the hard way that letting the ML model fully override the stats pipeline leads to systematic bias.
Team Ensemble
Takes 10 features about the team-level projection (pre/post shrinkage scores, sigma, bench production, coverage ratio, player count, total minutes, home indicator) and adjusts the team total.
Game Ensemble
The top-level model. 18 features about the full game projection (total, margin, pace/defense deltas, matchup strength, SOS, rest, overtime probability) producing the final game-level prediction.
Each model is trained with walk-forward cross-validation on temporal splits — we never let the model see future data during training. Hyperparameters are searched automatically. GPU-accelerated LightGBM on an RTX 5090 makes the search practical.
Layer 2: Player Embeddings (Teaching the Model What "Style" Means)
Here's a problem: how do you tell a model that Nikola Jokic and Joel Embiid are similar players in a way that matters for prediction, without manually encoding every aspect of their games?
Answer: contrastive learning embeddings.
We train a neural network to produce 16-dimensional vectors for each player, where players with similar impact patterns end up close together in the embedding space and dissimilar players end up far apart. The training signal comes from how players affect game outcomes when they share the court.
These 16 numbers — the player's "style fingerprint" — become features in the props ensemble. The model learns things like "when a player with this embedding profile faces a team with that defensive profile, the statistical projection tends to overshoot rebounds by X."
We do the same thing for coaches (coaching style embeddings) and for shot chart patterns (spatial shot embeddings). Each captures a different dimension of context that the statistical model can't easily represent.
Layer 3: Temporal Convolutional Network (Reading Game Logs Like Text)
A player's last 15 game logs are a sequence. And sequences are something neural networks are really good at processing.
Our TCN (Temporal Convolutional Network) reads a player's recent game log sequence and produces a "form residual" — how much the player is currently over/under-performing their baseline, accounting for the pattern of recent games.
Unlike a simple rolling average, the TCN can detect patterns: a player trending up across several games, a player who had one great game surrounded by mediocre ones (noise vs signal), a player whose minutes are gradually increasing (role change).
The TCN output becomes another feature in the props ensemble. It's not the most important feature — the statistical projections and embeddings contribute more — but it adds about 0.3 percentage points of win accuracy. In a game of small edges, that matters.
Layer 4: Quantile Regression (Probabilities, Not Point Estimates)
This was a conceptual shift more than a technical one, but it changed how we think about predictions.
Before quantile regression, our pipeline produced point estimates: "Jayson Tatum will score 26.3 points." But what we actually need for betting is a probability: "There's a 62% chance Tatum goes over 24.5."
The naive approach is to assume a Gaussian (normal) distribution around the point estimate and compute the probability from that. But player stat distributions aren't Gaussian — they're skewed, with fat tails. Tatum can score 45 on a random Tuesday but is very unlikely to score below 10 if he plays normal minutes.
Our quantile regression models (LightGBM trained on 11 quantile levels per stat) produce a full CDF — a complete probability distribution. The model learns that Tatum's distribution of scoring outcomes is different from, say, a role player who has a much tighter, lower range.
This directly impacts bet selection. A Gaussian assumption might say 60% probability on an over. The quantile model might say 55% because it knows the downside tail is fatter than normal. That 5% difference is the difference between a bet with edge and one without.
The CUDA Fork Crash: A War Story
Here's a fun one from the trenches.
Our backtest engine processes hundreds of games. To make it fast, we parallelize — 16 worker processes each projecting games simultaneously. On a 16-core machine, this turns a 30-minute backtest into a 2-minute one.
Python's multiprocessing uses fork() on Linux/Mac, which creates child processes that share the parent's memory via copy-on-write. Very efficient. One problem: CUDA (the GPU framework) cannot be re-initialized in a forked subprocess. If the parent process has touched any PyTorch GPU tensor, every forked child will crash silently.
"Silently" is the key word. The workers don't error out. They don't throw exceptions. They just... produce garbage. We spent hours getting bizarre prediction results before realizing the CUDA state was corrupted.
The fix: an environment variable (HOOPS_FORCE_CPU=1) that forces all PyTorch models to CPU before forking. The TCN and embedding models check this flag in their _get_device() function. During backtesting, everything runs on CPU with 16 parallel workers. During live single-game projections, GPU is fine.
It's the kind of bug that doesn't show up in any tutorial. You only find it by running a real workload at scale and noticing that the numbers don't make sense.
The RTX 5090 In the Closet
Our "GPU server" is an RTX 5090 in a desktop machine accessible via Jupyter notebook at a local IP address. It's not exactly a data center.
But it works. The optimization pipeline — testing hundreds of configurations, training ensemble models, running quantile regression across all stats and leagues — would take days on CPU. On the 5090, the full 14-stage tuning pipeline completes in a few hours.
We built a file sync system that uploads code changes to the Jupyter server, clears module caches (Python caches compiled modules, so you have to manually invalidate them after syncing new code), and kicks off training runs. It's held together with WebSocket connections and hash-based change detection.
It's scrappy. But when you're tuning 212 configurations across 2 leagues for an ablation study, you need that GPU.
What ML Adds (And What It Doesn't)
Let's be honest about the incremental value:
| Component | Win% Contribution |
|---|---|
| Statistical pipeline (base) | ~62% |
| LightGBM ensembles | +1.5-2% |
| Embeddings | +0.5% |
| TCN | +0.3% |
| Quantile regression | Better bet selection (hard to quantify as Win%) |
| Calibration | Better probability accuracy (same) |
The statistical pipeline does 90% of the work. The ML layer adds the last 10%. But in sports betting, that last 10% is the difference between having edge and not having edge. The sportsbook lines are already pretty good — they capture most of the obvious signal. The ML models capture the non-obvious stuff.
Is it worth the complexity? For us, yes. The ensemble models run in milliseconds for a single game. The embeddings and TCN are computed once per player per day. The infrastructure cost is a GPU we already had and some Python code.
But if someone asked us "what's the single most important thing for prediction accuracy?" the answer wouldn't be any ML model. It would be: get the minutes right, get the pace right, and respect the base rates. The machines help. The fundamentals matter more.
Our ML pipeline is four neural networks and four gradient-boosted forests stacked on top of statistics that would work fine on a TI-84. The statistics do the heavy lifting. The ML finds what the statistics miss.