# Scraping Odds From 5 Sportsbooks Without Getting Banned

*The cat-and-mouse game of aggregating betting odds — featuring user agent spoofing, headless browsers, and a Caesars anti-bot system that made us question our life choices.*

## Why Odds Comparison Matters

Before we get into the war stories, let's talk about why this matters.

The same bet — say, Jayson Tatum over 24.5 points — can have different odds at different sportsbooks. Hard Rock might have it at -115. DraftKings at -110. FanDuel at -108. BetMGM at -112.

The difference between -115 and -108 might look small. It's not. On a $100 bet:
- At -115: you risk $115 to win $100
- At -108: you risk $108 to win $100

Over a season of, say, 500 bets, consistently taking -108 instead of -115 saves you roughly $3,500 in vig. That's free money left on the table by not shopping around.

Line shopping is the single easiest edge any bettor can get. No models needed. No prediction engine. Just take the best available number.

The problem? Sportsbooks don't make it easy to compare. They each exist in their own app, their own ecosystem. And they definitely don't want you automating the comparison.

## The First Attempt: Just Hit the API

Our first approach was the obvious one: find each sportsbook's API endpoints, make HTTP requests, parse the JSON. Clean and simple.

Hard Rock has a GraphQL API that powers their web app. We reverse-engineered the queries from their network traffic, built a client, and started pulling odds. It worked. For about two days.

Then we started getting empty responses. Not errors — just... no data. We added logging. The requests were going through, returning 200 OK, but the response body was empty where odds should be.

The commit log from March 16th tells the story:

```
12:30 PM  fix: hr missing user agent
12:45 PM  fix: potential blocked by HR?
1:03 PM   more headers
2:00 PM   better logs
2:33 PM   use curl for hardrock
2:37 PM   fix: install curl
```

Hard Rock's API had started fingerprinting requests. No user agent? Blocked. Wrong user agent? Blocked. Missing referrer header? Blocked. We kept adding headers until our HTTP client looked exactly like a browser request. That bought us another week.

## The Headless Browser Era

When HTTP requests with all the right headers still started getting blocked, we escalated to headless browsers. If they want browser requests, we'll give them an actual browser.

A headless Chromium instance running on the server, navigating to the sportsbook's website, waiting for the JavaScript to render the odds, and scraping the DOM. It's slower, uses more resources, and is infinitely more annoying to maintain. But it works.

```
8:44 PM   use browser instance for odds
```

The key challenges with headless browsers:

**Resource management**: Each browser instance eats 200-300MB of RAM. Running five of them simultaneously for five sportsbooks means 1.5GB just for odds scraping. We implemented instance pooling — reusing browser contexts instead of launching fresh ones for each request.

**Timing**: You have to wait for the page to fully render, which means waiting for JavaScript to execute, APIs to respond, and the DOM to update. Too little wait time and you scrape an empty page. Too much and you're wasting time. We settled on a combination of element detection (wait until the odds elements exist in the DOM) and a maximum timeout.

**Detection**: Modern anti-bot systems don't just check headers. They check browser properties — WebGL renderer, canvas fingerprint, plugin list, timezone, screen resolution. Our headless instance needed to look like a normal user's browser.

## The Caesars Incident

Caesars (the sportsbook, not the salad) has the most aggressive anti-bot system we encountered. It's not just header checking. It's a full JavaScript challenge — think Cloudflare Turnstile but custom.

The commit log from March 21st, 1-2 AM:

```
1:14 AM   fix AN
1:31 AM   anti bot workaround
1:55 AM   fix caesars
```

Their system injects JavaScript that:
1. Runs a proof-of-work challenge in the browser
2. Checks for automation indicators (is `navigator.webdriver` set to `true`?)
3. Monitors mouse movement patterns (or lack thereof)
4. Fingerprints the rendering engine

Our headless browser was failing silently. The page would load, the challenge would run, and instead of odds data we'd get a "please verify you're human" page. At 1:31 AM, after multiple failed approaches, we found a combination of browser flags and JavaScript property overrides that passed the checks.

We're not going to detail exactly what we did here (the arms race continues), but let's just say that making a headless browser look sufficiently human is a non-trivial software engineering problem.

## The Name Matching Problem

Here's a problem nobody talks about when discussing odds aggregation: every sportsbook names things differently.

Hard Rock might list a game as "Boston Celtics vs Miami Heat" with player props under "Jayson Tatum." DraftKings might have "BOS Celtics @ MIA Heat" with props under "J. Tatum." FanDuel might use "Celtics at Heat" with "Jayson Tatum - Boston."

Same game. Same player. Different text. Now multiply by every game and every player across NBA, NCAAB, and NHL.

We built a fuzzy matching system that normalizes team names (mapping all variations to canonical names), handles player name abbreviations, and deals with the occasional genuinely different spelling. The "fix team assignment" and "fix: flickering and name matching" commits were about edge cases in this matching — what happens when two players have the same last name on the same team?

## The Timing Dance

Different sportsbooks update their odds at different times. Hard Rock might update NFL lines at 6 AM, while DraftKings updates at 4 AM. Player props might appear on one book 24 hours before another.

This creates a comparison problem: if you show "Hard Rock -115 vs DraftKings -110," but the Hard Rock line is from 8 hours ago and has since moved, the comparison is misleading.

We implemented freshness tracking for every odds snapshot. Each displayed comparison shows when the odds were last confirmed. Lines older than a threshold get flagged. It's not perfect — we can't poll every book every second — but it prevents the worst case of stale data misleading someone's bet placement.

## The Rate Limiting Tightrope

Poll too infrequently and your odds are stale. Poll too frequently and you get blocked.

Each sportsbook has different tolerance levels. Some will happily serve you every 30 seconds. Others start returning errors after more than a few requests per minute. One particularly aggressive book started rate-limiting us after a dozen requests in an hour.

Our polling strategy is adaptive:
- **Off-peak hours** (2 AM - 10 AM): poll less frequently, odds don't move much
- **Pre-game** (2 hours before tip): ramp up polling frequency, lines are moving
- **Live**: different endpoints entirely (live odds change by the second, different technical approach)

We also implemented exponential backoff — if a request fails, wait before retrying, increasing the wait each time. Getting permanently banned from a sportsbook's data feed would be catastrophic.

## What We Learned

### The arms race never ends
Every time you solve one anti-bot challenge, the sportsbook updates their defenses. This isn't a "build once and forget" feature. It requires ongoing maintenance.

### Browser-based scraping is fragile
A single CSS class change on the sportsbook's website can break the scraper. We get alerted when odds stop flowing and usually have to investigate a broken selector.

### The legal gray area
Odds data is published publicly on sportsbook websites. Aggregating it for comparison purposes falls into the same category as flight price comparison sites. It's not illegal, but sportsbooks don't love it. We're accessing public data — just more efficiently than a human with five browser tabs.

### Multi-book comparison is the easiest edge
After all the technical drama, the actual user value is simple: "this bet is -108 here instead of -115 there." No prediction model needed. No machine learning. Just comparison shopping. And it's worth thousands of dollars per year to anyone who bets regularly.

## The Current State

Our odds aggregation system runs continuously on the server, managing multiple headless browser instances, polling at carefully tuned intervals, matching games and players across different naming conventions, and presenting unified odds comparisons in the EdgeGoat UI.

Is it elegant? Not particularly. There's a Chromium instance running 24/7, a curl fallback for books that don't require JavaScript rendering, fuzzy string matching that occasionally gets confused by traded players, and a monitoring system that pages us when a scraper breaks.

But it works. And every time a user sees that their favorite bet is 5 cents cheaper on another book, all those 2 AM debugging sessions feel worth it.

---

*Line shopping is the simplest edge in sports betting. Making it automatic was anything but simple.*

---

Source: EdgeGoat model projections (https://edgegoat.com) — CC BY 4.0
Canonical: https://edgegoat.com/learn/scraping-odds-without-getting-banned
License: CC BY 4.0 — https://creativecommons.org/licenses/by/4.0/
Methodology: https://edgegoat.com/methodology
Model probabilities are estimates, not guarantees. 21+. Gambling problem? Call 1-800-GAMBLER.
