# The Data We Wish Existed (And How We Make Do)

*Building a sports prediction engine means constantly hitting the limits of available data. Here's how we get creative when the data we need doesn't come in a nice API.*

## The Fantasy

In an ideal world, we'd have a real-time feed of:
- Every player's exact position on the court, every second
- Every play's strategic intent (was that a designed play or a broken one?)
- Every player's fatigue level, sleep quality, and mental state
- The sportsbook's actual probability model (not just the line)
- Complete injury details (not "knee soreness" but "Grade 2 MCL sprain, 85% recovered")

We'd plug all that into a model and predict games with 90% accuracy.

Back in reality, we work with ESPN APIs, reverse-engineered sportsbook feeds, and a lot of creative inference.

## What ESPN Actually Gives You

ESPN is our primary data source. Their APIs are powerful but undocumented — there's no official developer program, no API keys, no rate limit documentation. Everything we know comes from reverse-engineering their web and mobile apps.

What we can get:
- **Scoreboard data**: Live scores, game status, period scores, game clock
- **Box scores**: Player stats for in-progress and completed games
- **Game logs**: Historical stat lines for every player, every game
- **Play-by-play**: Event-level data (made shot, missed shot, turnover, substitution)
- **Team standings**: Win-loss records, conference standings, division rankings
- **Injury reports**: Player injury status (out, doubtful, questionable, probable)
- **Roster data**: Player IDs, names, jersey numbers, positions

What we can't get (or can't reliably get):
- Real-time lineup data (who's on the court right now)
- Coaching staff information
- Referee assignments (some games, not all)
- Practice report details
- Detailed shot location data (ESPN has some, but it's inconsistent)

## The NBA CDN: Our Secret Weapon

The NBA's stats system runs on a CDN that serves JSON data for their website. The endpoints are different from ESPN's and contain different data — notably, more detailed shot chart information and advanced metrics.

We built a client that queries the NBA CDN for supplementary data that ESPN doesn't provide. This gets us:
- Advanced player metrics (usage rate, true shooting, plus/minus)
- Some shot location data (distance from basket, shot zone)
- Pace statistics at the team level

The catch: the NBA CDN is even less stable than ESPN's API. Endpoints change. Response formats shift. Rate limits are aggressive. It's the most fragile part of our data pipeline.

## Reverse-Engineering Play-by-Play Into Insights

Raw play-by-play data is a stream of events:

```
12:00 Q1  Jump ball: Team A wins
11:42 Q1  [Player A] makes 2-pt shot (assist: Player B)
11:20 Q1  [Player C] misses 3-pt shot
11:18 Q1  [Player D] offensive rebound
...
```

This is useful for real-time tracking, but the real value is in what you can *derive* from it:

### Pace analysis
By counting possessions (field goal attempts + turnovers + free throw trips - offensive rebounds), we can calculate real-time pace for each game. This is more accurate than using the final score because it accounts for game flow — a game might be slow in the first half and fast in the second, and our prop projections need to know that.

### Minutes context
Play-by-play tells us when players check in and out. This lets us compute not just total minutes but *minutes context* — "this player played 28 minutes, but 8 of those were garbage time in a 25-point game." A player's per-minute production in competitive minutes is more predictive than their raw per-minute rate.

### Lineup combinations
By tracking substitution events, we can reconstruct which five players were on the court at any given time. This feeds our chemistry model — how does each 5-man lineup perform relative to the sum of its parts?

### Play-by-play aggregation
Our `pbpAggregator` service processes the raw event stream into structured data: possessions per quarter, scoring runs, lineup stints, turnover sequences. This aggregated data becomes features in the prediction pipeline.

## Coaching Data: The Hard Way

There's no "coaching tendencies API." Coaches don't have stat lines. But coaching decisions — rotation patterns, timeout usage, late-game strategy — significantly affect game outcomes.

We approached this problem with embeddings. Instead of trying to manually encode coaching traits ("this coach plays small ball" or "this coach shortens the rotation in the playoffs"), we let the model learn them.

The coach embedding trainer looks at every game a coach has coached, treats it as a training example, and learns a 16-dimensional vector that captures coaching "style." Two coaches with similar embedding vectors tend to produce similar game patterns — similar pace, similar rotation depth, similar scoring distributions.

This works because we don't need to *understand* what the embedding dimensions mean. We just need them to be predictive. And they are — coaching embeddings add about 0.5 percentage points of win accuracy, which is meaningful in our system.

## Shot Chart Data: The Spatial Frontier

This is our newest data initiative: using shot location data to build spatial embeddings for shooters.

The idea: a player who shoots 45% from the left corner three is different from one who shoots 45% from the top of the key, even if their overall 3-point percentage is the same. Teams defend different zones differently. A player's shot chart — where they take shots and how often they make them — is a fingerprint of their offensive game.

We're building contrastive learning embeddings from shot chart data. Two players with similar spatial shot distributions end up close in the embedding space. These embeddings then become features for predicting 3-point and scoring props.

The challenge: shot location data is inconsistent across sources. ESPN provides zone-level data (left corner, right wing, etc.) but not precise coordinates for every shot. The NBA CDN has better resolution but only for NBA, not college. We're working with what we have and interpolating the rest.

## What We Synthesize

Some of the most valuable features in our pipeline don't come from any API — they're calculated from combining multiple data sources:

### Strength of schedule
No API provides "this team's strength of schedule." We calculate it from the full schedule of every team — who they played, and how good those opponents were (which requires knowing how good *those* opponents' opponents were, recursively). It's a graph problem that we solve iteratively at the start of each backtest run.

### Rest and schedule context
"Is this team on a back-to-back?" Seems simple, but you need the full schedule to answer it. "Did they travel?" You need to know where the previous game was. "How many games in the last 7 days?" You need a rolling window over the schedule.

None of this comes from a single API call. We build the full schedule for every team, compute rest metrics, and use them as pipeline inputs.

### Player reliability scores
How consistent is a player's production? A player who averages 20 PPG but scores between 15-25 every night is different from one who averages 20 but alternates between 10 and 30.

We compute reliability scores from game log variance — coefficient of variation, percentile stability, consistency across different opponents. These scores feed into the prediction pipeline as uncertainty modifiers: for unreliable players, we widen our confidence intervals.

### Bayesian priors for small samples
Early in the season, a rookie has played 5 games. His stats say he scores 22 PPG, but that's a tiny sample. Is he really a 22 PPG scorer, or is that noise?

We use Bayesian priors — league-average or position-average expectations — that get overwhelmed by data as the season progresses. After 5 games, the prior dominates. After 40 games, the data dominates. The transition is smooth and automatic.

## The Missing Data Problem We Can't Solve

Some data gaps genuinely hurt our predictions:

**Real-time injury severity**: "Questionable" means anything from "he's definitely playing" to "he's definitely not playing." The sportsbook knows more than we do here because they have team contacts and insider information. This is a structural disadvantage we can't close.

**Minute-to-minute lineup decisions**: Will the coach play his starter 36 minutes tonight or rest him at 30? This is decided in real-time based on game flow, and sometimes the coach doesn't know in advance. Our minutes model is good (4.75 MAE), but it can't predict a coach's real-time rotation decisions.

**Sportsbook probability models**: We see the odds (e.g., -115), and we can infer the implied probability (53.5%). But the sportsbook's actual model might say 55% — the extra 1.5% is their margin. We don't know their true probabilities, only the margin-loaded ones.

**Private information**: A player who went out last night. A locker room conflict. A nagging injury that hasn't been disclosed. This information exists but isn't in any API. It's why sportsbooks will always have some information advantage over public models.

## The Philosophy

Our approach to the data problem is pragmatic: extract every drop of signal from public data, synthesize features that no single source provides, and accept the structural information disadvantages without pretending we can overcome them.

We can't know that a player is secretly dealing with a nagging ankle injury. But we can know that his last three games show declining minutes, decreasing burst in play-by-play data, and a shift toward mid-range shots instead of drives. The data tells us *something* is off, even if we don't know *what*.

That's the game. Not perfect information. Not insider knowledge. Just relentless extraction of signal from the data everyone has access to, processed better than anyone else is processing it.

---

*The best data in sports prediction isn't the data you have — it's the features you synthesize from combining everything together.*

---

Source: EdgeGoat model projections (https://edgegoat.com) — CC BY 4.0
Canonical: https://edgegoat.com/learn/the-data-we-wish-existed
License: CC BY 4.0 — https://creativecommons.org/licenses/by/4.0/
Methodology: https://edgegoat.com/methodology
Model probabilities are estimates, not guarantees. 21+. Gambling problem? Call 1-800-GAMBLER.
