For the 2026 World Cup I built a model that rated every national team and put a win, draw and loss chance next to every match. An AI agent wrote most of the code, and it went from an empty repo to a backtested model in about two days.
The maths isn’t hard and the data is free. The hard part is knowing what to build in what order, and when not to believe what you’re looking at. This is the guide I’d want if I were starting again. (I did the same for rugby league.)
You’re building two things. A rating that says who’s stronger, and a second layer that turns the gap between two ratings into match probabilities. Keep them apart, so when a number looks off you know which half to dig into.
Start with the data
You want every men’s international ever played. Mart Jürisoo keeps a free, public dataset on GitHub called international_results. When I last pulled it, it had 49,945 matches between 337 teams, going back to Scotland against England in 1872. Each row has the date, teams, score, tournament, venue and a neutral flag, which is everything Elo needs.
A few things to sort out first.
- Scores include extra time. A knockout won in extra time shows up as a win, not a 90-minute draw. At the 2026 World Cup, 9 of the 32 knockout ties were level after 90 minutes, so decide which question you’re answering.
- The data gets fixed upstream. My backtest gained 12 matches between two runs without me touching a thing. Save a dated copy of whatever you test on.
The rating update
Every team starts at 1,500. Before a match, take the home rating, subtract the away rating and add 100 points for home advantage. At a neutral venue, add nothing. Then run the gap through the standard Elo curve.
d = home_rating − away_rating + home_advantage
E = 1 / (1 + 10^(−d / 400))
E is the home side’s expected score. After the result, the home rating moves by K × G × (S − E) and the away side moves by exactly the opposite, so the pair’s total never changes. S is 1 for a home win, 0.5 for a draw and 0 for an away win.
Pick a K
K is how much a match is allowed to matter. The standard football weights are 60 for a World Cup match, 50 for a continental championship, 40 for qualifiers and the Nations League, 30 for minor tournaments and 20 for friendlies. G scales for the margin. It’s 1 for a draw or a one-goal win, 1.5 for two goals and (11 + margin) / 8 for three or more.
None of that is mine. It’s the formula behind World Football Elo Ratings, and I copied it on purpose. A rating engine is the wrong place to get creative. You can tune the weights, but the best tuned set I found improved log loss by 0.0001 and won 6 of 11 test periods. That’s a coin toss.
A worked example. Two teams on 1,500, one of them at home. The 100-point bump makes E about 0.640. With K = 40, a one-goal home win adds 14.4 points to the home team and takes 14.4 off the away team. A draw costs the home side about 5.6 points, because it was expected to do better than a draw.
Have a play.
Elo, step by step
How an Elo rating changes after one match
Change the match and see what the maths does.
That’s an expected score out of 1, where a draw is worth half. It’s not a 64% chance of winning. Whatever one team gains, the other loses.
40 × 1 × (1 − 0.640) = +14.4 points
The layer that knows about draws
Adjusted difference:
These numbers are just for show (β = 0.7, cut points ±0.55). They’re not from the fitted model.
The calculation, without the shorthand
Both teams start at 1500. The gap slider pushes them apart by the same amount each way. Home advantage, when it’s on, adds 100 points to the difference. It doesn’t touch either team’s actual rating.
d = R_home − R_away + home advantage
E = 1 / (1 + 10^(−d / 400))
Change = K × G × (S − E)
S is 1 for a home win, 0.5 for a draw and 0 for an away win. G is 1 for a margin of zero or one goal, 1.5 for two, and (11 + margin) / 8 for three or more. A draw always has a margin of zero. The K values are the standard ones.
The probability layer is a separate step. It starts with z = 0.7 × d / 100. Then, with σ(x) = 1 / (1 + exp(−x)), away is σ(−0.55 − z), draw is σ(0.55 − z) − σ(−0.55 − z) and home is 1 − σ(0.55 − z). I made those coefficients up on purpose, because the fitted weights aren’t part of this page.
Here’s the example. Both teams are on 1500, home advantage is 100, it’s a qualifier and the home team wins by a goal. Home goes up 14.4 points to 1514.4, and away drops 14.4 to 1485.6. Expected score 0.640. With the made-up numbers, it’s home win 53.7%, draw 24.0%, away win 22.3%.
Expected score isn’t a win probability
This is the easiest mistake to make. Elo’s E is an expected score, with a win worth 1 and a draw worth half. For a three-way forecast:
expected score = P(win) + 0.5 × P(draw)
A 50% chance of winning with a 28% chance of a draw gives you 0.640. So does a 60% chance of winning with an 8% chance of a draw. Those are very different matches, and the Elo curve can’t tell them apart. It has no opinion about draws at all, and draws are about a fifth of competitive internationals.
Turning ratings into probabilities
The fix is a second layer, an ordered logit. Picture the rating gap as a position on a line with two cut points on it. Left of the first is an away win, right of the second is a home win, and the space in between is the draw. A bigger gap slides the match along the line, and how far it slides per 100 rating points is a number the model learns, beta. With sigmoid(x) = 1 / (1 + exp(−x)):
z = beta × d / 100
P(away win) = sigmoid(c1 − z)
P(draw) = sigmoid(c2 − z) − sigmoid(c1 − z)
P(home win) = 1 − sigmoid(c2 − z)
Write the upper cut point as c2 = c1 + exp(gap) so the two stay in order and the draw can never go negative. Then fit the three numbers by maximum likelihood, which just picks the values that make the actual results look most likely. I fitted mine on competitive matches since 1990.
It only sees one number per match, so it knows nothing about lineups or a side with nothing left to play for.
For a knockout tie, the chance of going through is roughly the 90-minute win chance plus half the draw chance. That treats extra time and penalties as a coin toss.
Test it properly
To test a forecast you need the ratings as they stood before each kickoff, or the forecast has already seen the answer. Store the pre-match ratings on every row as you replay history.
Then walk forward. Tune on an early period, then step through the rest two years at a time, refitting the probability layer on earlier data and scoring the next window. The ratings keep updating match by match, because last Tuesday’s result is fair game for this Saturday. This Saturday’s result is not. I tuned on 1994–2005 and tested on 9,927 competitive internationals from 2006 onwards.
Before you compare it with anything clever, beat two dumb models. One just repeats how often home wins, draws and away wins happen. The other gives every match the same draw chance. Losing to either is embarrassing, which is exactly why they’re useful.
| Model | Log loss ↓ | Ranked probability score ↓ | Accuracy ↑ |
|---|---|---|---|
| Ordered logit | 0.844 | 0.164 | 62.2% |
| Fixed-draw baseline | 0.867 | 0.168 | 62.3% |
| Base-rate baseline | 1.046 | 0.230 | 48.2% |
Log loss measures how surprised you were by what happened, and it punishes a confident miss hard. Ranked probability score knows the outcomes are in order, so calling a home win when it was a draw costs less than calling one when the away side won. Lower is better for both.
Don’t lean on accuracy. It only asks whether your most likely outcome happened, so it can’t tell 40% from 90%. In the table, the fixed-draw model edges it on accuracy and loses on the two scores that matter.
Check the calibration
Then ask the question people actually ask. When the model says 70%, does it happen 70% of the time? Bucket the forecasts and compare.
Across 10,195 out-of-sample matches, mine gave the favourite 63.8% on average, and the favourite won 62.3% of the time. Close. In the 85–90% bucket the favourite won 550 of 663, or 83%. A touch overconfident at the top end.
Use the bookies as a benchmark
Beating the floors tells you that you have a model. The bookmakers tell you how good it is.
Their prices carry a margin, so the implied probabilities add up to more than 100%. Taking it out is called de-vigging, and the simple way divides each probability by the total. Shin’s method assumes some of the margin is protection against punters who know more than the bookie, so it takes more off the longshots. For me the choice barely mattered.
Compare against the closing price, the last one before kickoff. It has all the late team news baked in, so it’s the toughest benchmark you’ll find. Expect to lose to it. If you want the market inside your model, feed in the opening price and keep the closing price as the scoreboard. Blend the closing price in and you’re marking your own homework.
Working with an AI agent
The agent will write code faster than you can read it. Your job is to be the sceptic.
Write the plan first. Before any code, get the agent to write the plan into a file in the repo. The data, the rating, the probability layer, the test and the floors. Then build in that order.
Ask for the backtest before you believe anything. Every time the agent says something improved, ask to see it against the floors. Most ideas don’t survive. I tried squad values, confederation adjustments, travel, rest and time decay, and none of them cleared the bar.
Check nothing from the future leaks in. The neatest check is to cut the data off the day before a match, predict it, and confirm you get the same number as the full run. Ask the agent to write that as a test.
Read rows by hand. Averages can look healthy while something underneath is badly wrong. Sort by the biggest surprises and read the top 20 one at a time. My biggest was Luxembourg winning 2–1 in Switzerland in 2008, at about one chance in 400. That one’s real. Not every extreme row will be.
Keep notes. Keep a lessons file the agent reads at the start of each session, so you’re not teaching it the same thing twice.
Freeze your forecasts. Commit them to git before kickoff. It’s the only score nobody can argue with, including you.
Replay all of history
Once the engine works, replay all 49,945 matches and you get a world #1 for every day since 1883. You need a rule for who counts. Mine was at least 20 matches played and one in the last four years.
Nineteen nations have held the top spot. Brazil held it longest, 36.8 years all up, then England on 27.7. Scotland is third on 15.6, more than half of it before 1900, when it was mostly the four home nations playing each other. The 2026 World Cup final was #1 Argentina against #2 Spain, and Spain won 1–0 after extra time to take it.
Ratings creep up as more teams join the pool, so compare teams within an era, not across centuries.
Ballon p’Oor
World number one in men’s football, 1872 to 2026
I put every men’s international since 1872 through the same Elo engine that ran my World Cup site. Nineteen countries have had a turn at number one, and the 2026 final was #2 knocking off #1 for the top spot.
Since 1883, when a team first had the 20 matches you need to qualify, nineteen different countries have been number one. Brazil has spent the longest on top, 36.8 years all up, then England (27.7), Scotland (15.6) and Argentina (13.0). The 2026 World Cup final was #1 Argentina (2,245) against #2 Spain (2,233). Spain won 1–0 after extra time and is still number one.
| Nation | Years at #1 | Separate spells | Longest spell |
|---|---|---|---|
| Brazil | 36.8 | 59 | Apr 1995 to Jul 2000 |
| England | 27.7 | 29 | Apr 1891 to May 1902 |
| Scotland | 15.6 | 12 | Mar 1883 to Apr 1891 |
| Argentina | 13.0 | 43 | Jan 1941 to Jul 1943 |
| Italy | 11.2 | 16 | Oct 1934 to Feb 1940 |
| Hungary | 8.5 | 13 | Jun 1918 to Apr 1921 |
| Germany (incl. West Germany) | 8.0 | 22 | Mar 1993 to Jul 1994 |
| Netherlands | 4.8 | 10 | May 1915 to Oct 1917 |
| USSR / Russia | 4.6 | 16 | Apr 1963 to Oct 1964 |
| Spain | 4.1 | 11 | Jun 2012 to Jun 2013 |
| Austria | 3.2 | 5 | Sep 1931 to Oct 1933 |
| France | 3.1 | 10 | Jul 2000 to May 2002 |
| Czechoslovakia | 1.8 | 7 | Sep 1927 to Jun 1928 |
| Uruguay | 1.1 | 2 | Jun 1924 to Apr 1925 |
| Mexico | 54 days | 1 | Mar to May 1990 |
| Belgium | 36 days | 2 | Sep to Oct 2020 |
| Denmark | 14 days | 1 | Jun 1918 |
| Sweden | 6 days | 1 | Jun 1988 |
| Czech Republic | 4 days | 1 | Jun 2005 |
If you want to try this
- Download the international results and save a dated copy.
- Get the agent to write the plan before it writes any code.
- Build the Elo update with the World Football Elo weights and 100 points of home advantage. Don’t tune yet.
- Add an ordered logit on the rating gap and walk it forward against the two floors.
- Check the calibration, then benchmark against de-vigged closing prices.
- Freeze your forecasts before kickoff.
Credit to Arpad Elo for the ratings, World Football Elo Ratings for the football version, Peter McCullagh’s Regression Models for Ordinal Data (1980) for the ordered logit and Hyun Song Shin for the de-vig. I just plugged them together with an agent at the keyboard.