Flat-track bullies: we built the stat, then tested whether it means anything
The thesis
There's a theory you'll hear in every sport, and it goes like this: some teams are flat-track bullies. They hammer bad opponents, bank stats and wins across a long regular season, and look terrific — right up until the postseason, when the schedule turns into nothing but good opponents and their whole approach stops working. The production was real, the theory says, but it was *concentrated in the wrong places*, and concentration is a hidden weakness that standings and season-long averages can't see.
It's a seductive idea for the NFL. Playoff defenses are better than regular-season defenses. If your offense only moves the ball against the league's soft underbelly, January should expose you. And the reverse should hold too: an offense that keeps producing against top defenses — a *matchup-proof* offense — should be built for the postseason even if its season-long numbers look ordinary.
So we built the measurement, put it on the site, and then did the thing the theory's fans usually skip: we tested whether it predicts anything.
The stat
The Bully Index now lives on the play breakdown page. For every team, we split offensive EPA per play by the quality of the defense it faced: the league's top third of defenses, the middle, and the bottom third.
Two details matter, because they're what make the number honest:
Defense quality is graded as of each game's week. A week-3 game is bucketed using only what was knowable before week 3 — season-to-date defensive EPA per play allowed, blended toward the prior season so early weeks aren't guesswork. No end-of-season hindsight leaks backward into September games.
The headline number is the gap: EPA per play against bottom-third defenses minus EPA per play against top-third defenses. Every offense in football does better against bad defenses — the gap measures *how much* better. A big gap is the flat-track-bully profile. A small one is matchup-proof.
The measurement works. Teams genuinely differ: in a typical season the gap runs from roughly zero to a quarter of an EPA per play, which is the difference between an elite offense and a broken one. The question is whether that difference *means* anything.
Test one: does the gap predict playoff performance?
The core claim. If bully production doesn't travel, teams with big gaps should underperform in the playoffs relative to their overall quality.
We took every playoff team from 2015 through 2025 — 144 team-seasons — and regressed playoff offensive EPA per play on two things: the team's overall regular-season offense, and its gap. If the theory is right, the gap's coefficient should be negative: for two teams with identical overall offenses, the bully should do worse in January.
Result: the gap came out negative and significant (−0.24 EPA per play per unit of gap, p=.04). One test in, the theory looks alive.
But 144 playoff observations is a small sample, and playoff EPA is one of the noisiest quantities in football. So before running anything, we pre-registered a second, bigger version of the same question — and committed to believing it over the first.
Test two: the same question with twice the data
If a big gap really means your production can't survive good defenses, that should show up *during the season too*. So: measure each team's gap using only weeks 1–12, then check how that team's offense performed against top-third defenses in weeks 13–18. Same logic as the playoff test, no playoff small-sample problem, 288 observations.
Result: nothing. The gap's coefficient was +0.03 with p=.65 — indistinguishable from zero, and pointing the wrong way. Knowing a team's early-season gap tells you nothing about how it will handle good defenses late, once you know how good its offense is overall.
An effect that appears in 144 noisy playoff games but vanishes in 288 cleaner in-season games is, by the rules we set before looking, unconfirmed. Maybe there's something genuinely playoff-specific — two weeks of defensive game-planning, January weather — that weeks 13–18 can't capture. Or maybe the playoff result is what p=.04 often is: noise that got lucky. We can't distinguish those yet, so the claim doesn't get to stand.
Test three: is bullying at least a regular-season superpower?
The theory has a flip side that's just as testable. If concentrating production against bad opponents is *bad* for January, maybe it's *good* for September through December — beat up on the teams you should beat, bank those wins, let the coin flips against good teams land where they may.
So: regular-season wins, regressed on offensive EPA per play, defensive EPA per play allowed, and the gap. 350 team-seasons. If distribution matters in either direction, the gap picks it up.
Result: the cleanest null we have ever published. The gap's coefficient was +0.001 wins — p=.999 — inside a regression that otherwise explains wins almost perfectly (R²=.76, with offense worth +24.9 wins per unit of EPA per play and defense worth −23.9, as symmetric as theory says they should be). Total efficiency is the entire story. *How* a team distributes its production across opponent quality contributes literally nothing to its record.
What actually happens when good offenses meet great defenses
Since we had eleven seasons of graded matchups — 5,790 team-games — we asked the collision question directly.
| vs good D | vs mid D | vs bad D | |
|---|---|---|---|
| good offense | +0.026 | +0.072 | +0.088 |
| mid offense | −0.034 | +0.011 | +0.042 |
| bad offense | −0.096 | −0.058 | −0.013 |
Three things fall out of that table:
Good offenses stay above water against everyone. Even against top-third defenses, they average +0.026 EPA per play — better than a league-average offense in a neutral matchup. A great defense takes its bite, but it does not neutralize a good offense.
The matchup is additive. Predicting the strength-vs-strength corner cell by just adding the row and column effects gives +0.024; the actual value is +0.026. We also fit an explicit interaction term: p=.50. There is no amplification, no "defense wins when both are elite." Add the two ratings and you're done.
The tug-of-war leans offense. A team's offense carries about 0.68 of its rating into any given matchup; a defense imposes only about 0.51 of its own. Offense is the more stable, more portable identity — which is why the same names keep appearing at the top of the scoring tables no matter who they play.
The quarterback version, and why your memory is lying to you
Everyone has a QB in mind right now — the guy who "always shrinks against good defenses." So we ran the same split for every quarterback with at least 200 dropbacks against each outer third since 2015. Seventy qualified, and over a full career the spread looks dramatic: some quarterbacks show a gap of +0.30 EPA per dropback or more, feasting on bad defenses and cratering against good ones, while others produce nearly identically no matter who's across the line.
Then we checked whether it's a real trait, two ways:
- Year over year: a quarterback's gap in one season correlates with his next season's gap at r = 0.01, across 215 season-pairs. Nothing.
- Within the same career: split every qualified QB's games into even and odd weeks and compare the two gaps — r = −0.09. A quarterback's matchup sensitivity in half his own games doesn't predict it in the *other half of the same games*.
For calibration, plain EPA per dropback — actual quarterback quality — has a split-half correlation of +0.76 in the identical setup. The method finds real traits just fine. Matchup sensitivity isn't one of them.
So when a career table shows a journeyman with a monster gap, that's not a scouting report — it's what sorting 70 noisy numbers always produces at the extremes. The one durable pattern is the boring one: good quarterbacks are good against everyone, bad ones are bad against everyone. "Matchup-proof" isn't a style. It's just being good.
What the Bully Index is for
After three pre-registered tests and a stability study, the honest summary:
- The gap describes how a season happened — who piled it up against the soft part of the schedule, who earned it the hard way. That's genuinely worth knowing when a team's raw numbers are inflated by the defenses it happened to face.
- The gap predicts nothing — not playoff performance (unconfirmed at best), not wins (emphatically not), and it isn't a stable trait at the player level.
- Strength-vs-strength football is additive. There is no hidden interaction where elite defense unlocks a special counter to elite offense.
We built the stat because the thesis was interesting. We're keeping it because the description is useful. And we're publishing the nulls because the alternative — letting a good story imply a prediction the data won't back — is how most bad betting advice gets made.
The full methodology, regression tables, and verdict files are committed alongside the model scorecard. The Bully Index itself is live on the play breakdown page — go see who's been bullying, and now you'll know exactly how much it means.
Run this on your own league.
Connect ESPN or Sleeper to get trade suggestions, league-aware rankings, and manager grades.