College Football · Analytics
A statistical breakdown of how much each factor explains AP Top 25 poll points.
This uses a technique called Shapley-value decomposition to split credit for AP poll points among a handful of factors: win percentage, points scored/allowed per game, yards gained/allowed per game, turnover margin, strength of schedule, conference, and the team's own AP points from the previous week's release. Every FBS team is included each week, not just the 25 that were ranked - an unranked team just counts as zero points.
Together, these factors explain 91.9% of the variation in 2025 AP Top 25 points. The breakdown below shows each factor's share of that explained portion - not of the whole picture. The rest is whatever these factors don't capture: narrative, preseason momentum, week-to-week volatility, and plain human unpredictability.
This is a descriptive analysis, not a statistical significance test, and it doesn't prove bias of any kind. Conference showing up here means a team's conference has some relationship with its poll points beyond what its own performance and schedule strength already explain - it doesn't establish why.
The previous week's AP points are broken out separately below rather than mixed into the bar chart - it consistently dwarfs every other factor, which is itself the finding: AP voting is heavily anchored to where a team was already ranked, not fully re-derived from that week's results.
2025 · 1904 team-weeks analyzed
The single biggest factor, by a wide margin - a sign the poll is heavily anchored to where voters already had a team ranked, rather than fully re-derived from that week's results. Shown separately from the factors below so it doesn't flatten them out of view.
2025 · 1904 team-weeks · 10,000 permutations
The factor breakdown above shows conference carries real weight in the poll overall. This section goes further - not one number, but a sequence of seven tests, each one checking (and sometimes correcting) the one before it, on the 2025 season (1904 team-weeks). Bars extending right are getting more poll credit than performance predicts; bars extending left are getting less. Only bars tagged "significant" have a 95% confidence interval that excludes zero - the rest are directionally suggestive but not statistically distinguishable from noise given the sample size.
A regression predicts AP poll points from performance alone (win percentage, scoring, yards, turnover margin, strength of schedule) - no conference, no team name. Each conference's bar is the average gap between what its teams actually got and what that model predicted, for the full season.
R² = 0.366 · spread p = 0.4611 · ANOVA F(10, 125) = 1.897, p = 0.0519
16 teams · 95% CI [40.9, 374.9]
2 teams · 95% CI [-84.2, 438.3]
13 teams · 95% CI [-38.7, 170.3]
18 teams · 95% CI [-139.9, 210.8]
12 teams · 95% CI [-64.5, 105.8]
14 teams · 95% CI [-108.6, 35.0]
17 teams · 95% CI [-137.4, 69.9]
12 teams · 95% CI [-156.9, 5.3]
16 teams · 95% CI [-174.0, 33.1]
2 teams · 95% CI [-254.9, 55.3]
14 teams · 95% CI [-185.0, -23.7]
Our first significance test treated every team-week as independent evidence, which overstated confidence - the same ~16 SEC teams reappearing across 14 weeks each isn't 224 independent data points, it's 16. Every result in this section is clustered by team (each team's weeks are averaged into one number first) before comparing conferences, and 95% confidence intervals are added so individual conferences can be judged on their own, not just the group as a whole.
With clustering fixed, the broad claim ("conference matters, full stop") sits right at the edge of significance. But two conferences individually clear the bar on their own: the SEC (positive) and the American Athletic (negative). Every other conference's confidence interval includes zero - its point estimate might be directionally right, but the data can't currently tell it apart from chance.
Adding a team's own AP points from the previous week as a predictor measures marginal bias - how much extra credit shows up in a single week's vote, on top of wherever a team already stood. Once that's controlled for, the SEC's edge collapses and stops being statistically significant: this isn't a weekly thumb on the scale.
R² = 0.918 · spread p = 0.8849 · ANOVA F(10, 125) = 0.889, p = 0.5381
16 teams · 95% CI [-10.8, 39.5]
13 teams · 95% CI [1.7, 25.5]
2 teams · 95% CI [-15.0, 41.4]
12 teams · 95% CI [-6.1, 18.0]
18 teams · 95% CI [-19.7, 21.7]
2 teams · 95% CI [-20.5, 19.8]
14 teams · 95% CI [-9.4, 5.1]
16 teams · 95% CI [-24.3, 19.8]
12 teams · 95% CI [-19.2, 7.4]
14 teams · 95% CI [-19.3, 1.1]
17 teams · 95% CI [-25.5, 0.7]
Splitting the season at its midpoint (week 9) and refitting separately on each half tests whether bias compounds gradually or is already present early. It's the latter: the effect is significant in the first half and fades by the second, as the performance-only model's R² rises from 0.31 to 0.45 and results start explaining more of the poll than reputation does.
Early (weeks 3-9)
R² = 0.312 · spread p = 0.3900 · ANOVA F(10, 125) = 2.031, p = 0.0465
16 teams · 95% CI [43.0, 403.4]
2 teams · 95% CI [-72.5, 298.9]
13 teams · 95% CI [-18.1, 169.3]
18 teams · 95% CI [-147.5, 231.9]
12 teams · 95% CI [-69.2, 93.4]
17 teams · 95% CI [-146.4, 130.8]
14 teams · 95% CI [-103.3, 33.4]
12 teams · 95% CI [-134.5, -2.5]
2 teams · 95% CI [-222.6, 37.5]
14 teams · 95% CI [-187.9, -35.9]
16 teams · 95% CI [-211.8, -14.8]
Late (weeks 10-16)
R² = 0.446 · spread p = 0.4742 · ANOVA F(10, 125) = 1.382, p = 0.1966
2 teams · 95% CI [-88.3, 542.6]
16 teams · 95% CI [-7.9, 371.9]
13 teams · 95% CI [-64.5, 190.0]
12 teams · 95% CI [-59.6, 135.8]
18 teams · 95% CI [-147.1, 180.6]
14 teams · 95% CI [-115.1, 52.2]
16 teams · 95% CI [-157.3, 95.6]
17 teams · 95% CI [-151.8, 36.8]
2 teams · 95% CI [-258.7, 93.1]
12 teams · 95% CI [-199.3, 24.3]
14 teams · 95% CI [-196.0, 2.4]
About 80% of team-weeks get zero AP points, which could distort the fit. Restricting to just the 45 teams that received any votes all season changes the story again: the SEC's bias is no longer distinguishable from zero here, but the American Athletic and Big 12 both show up as significantly underrated even among teams already earning support - the single most consistent result across every test on this page.
R² = 0.384 · spread p = 0.0432 · ANOVA F(6, 38) = 2.514, p = 0.0331
1 team · 95% CI [212.3, 212.3]
12 teams · 95% CI [-136.8, 227.1]
8 teams · 95% CI [-184.9, 137.9]
9 teams · 95% CI [-280.1, 169.1]
9 teams · 95% CI [-325.3, -147.5]
5 teams · 95% CI [-425.6, -232.6]
1 team · 95% CI [-648.9, -648.9]
Tests 1-5 treat the eleven real conferences as fixed. This test resamples which conferences appear - drawing 11 with replacement from the real 11, each drawn conference's actual teams intact - 5,000 times, and tracks which conference comes out most-overrated and most-underrated each time.
5,000 resamples · spread 282.6 (95% CI [128.2, 311.7]) · ANOVA F 1.723 (95% CI [0.444, 3.460])
Most overrated, how often
Most underrated, how often
A meta-analytic heterogeneity test: treats each conference's own team-clustered mean and standard error (from Test 2) as one data point, the way a meta-analysis combines separate studies, and asks whether the eleven estimates disagree with each other by more than their own uncertainty would predict.
Q = 17.19 (df = 10) · I² = 41.8% · permutation p = 0.4200
Not significant - which agrees with, rather than contradicts, Test 2's borderline ANOVA. Two independently-built tests landing in the same place: the broad "conferences differ" claim still isn't proven, even though the SEC and AAC individually hold up.
This is one season of correlational data - not a controlled experiment, and not proof of intent behind any single ranking. Independents and the Pac-12 are two teams each after realignment, too small a sample for a conference-level claim on their own. See the full write-up in Articles for the complete methodology.