simulating…
How this works
The short version: the rest of the 2026 season is played out 2.5 million times, and your team's odds are counted in the seasons where one game went one way against the seasons where it went the other. The difference is what that game is worth to you.
Playing out a season
Every game still to come gets a score. The margin is drawn around what ESPN's FPI says the gap between the teams is, plus home field, with the spread of real results around that prediction: about 14 points. Ratings are not exactly right, and their errors last, so each simulated season also draws one error per team and keeps it all year. A team the ratings overrate is overrated in September and in November, which is what makes a whole season plausible rather than a string of independent coin flips.
Rating what happened
A simulated season is a full set of results, so it gets rated from scratch the way the real one would be, with Kenneth Massey's model. Two stages: a maximum likelihood fit that asks which set of ratings makes the scoreboard most likely, then a Bayesian correction that reads each team's wins and losses against the quality of who it played. Margin counts, but less than the scoreboard suggests, because the committee does not reward blowouts the way a pure margin model does.
Ranking, championships and the bracket
A rating is not a ranking. On top of it the model adds the committee's own variability, a bump for beating good teams and for losing only to good ones, a head-to-head rule, and the boost a conference champion gets. Conference races are then settled by each league's own published tiebreakers, the champions are crowned, and the playoff field is picked under this season's rules: the four power conference champions, the best team from the other six conferences, Notre Dame if it is ranked in the top 12, and at-large bids for the rest. The bracket is played out, so "win the national title" means winning four games, not being ranked first.
Turning that into a rooting guide
Because every unplayed game is simulated independently, one run answers every question. Split the seasons by who won a given game and compare your team's odds across the two halves, and you have that game's effect without simulating anything twice. Both halves share every other game's outcome, so the comparison is cleaner than running two separate simulations would be.
Why some games are marked as not significant
A guide compares hundreds of games at once, and at that many comparisons some differences look real by chance alone. Each difference gets a confidence interval and a p-value, and the whole set goes through a false discovery rate procedure, so what is shown as significant is what survives the multiple comparisons rather than what happened to look big. Games that do not survive are still listed, just marked plainly.
How well it does
Tested against every real playoff committee poll from 2023 through 2025, the model's ranking sits about 2.9 places from the committee's on average, and its playoff field contains 23 of the 28 teams the committee actually picked. Where the two disagree, it is usually the same way: the model rates teams with fewer losses from weaker leagues higher than the committee does. Settings were tuned on those seasons and checked by leaving one season out.
What it is not
The odds are the output of a model, not a prediction anyone should bet on. It does not know about injuries, suspensions, weather, or a quarterback changing everything in October. The scores update every hour on game days, and the numbers move with them.
The whole thing is open source: github.com/goodsellkai/cfb-rooting.
Win total
Playoff seed
Your own games
Who to root for
League-wide odds
One simulated season
Simulating a season…