Research / FPL optimization

Simple models, optimized teams

Why a better forecast does not always produce a better FPL score

Abstract

FPL optimization often begins with an impressive forecast and ends with a simple instruction: choose the players with the largest numbers. This study tests the full chain. It estimates player value, constructs a legal starting eleven, chooses a captain, and compares several ways of handling uncertainty.

The surprise is not that optimization helps. It is that several elaborate forecasting methods fail to separate themselves from much simpler ones.

01

Prediction enters as a price

An optimizer needs one expected return for every player. The study calls these estimates cost vectors and builds them in several ways: recent averages, exponential smoothing, ARIMA, linear regression, hybrid scores, the FPL ICT index, and simulation.

Once the estimates exist, integer programming selects one goalkeeper, legal numbers of defenders, midfielders and forwards, and one captain. The model uses an £83.5 million starting budget, reserving the remaining money for four inexpensive substitutes.

  1. 01Estimate every player's return
  2. 02Apply budget and position rules
  3. 03Select eleven players
  4. 04Double one captain's score
A forecast becomes useful only after it survives the game rules.
02

Simple methods stayed competitive

Across the out-of-sample gameweeks, weighted averages finished among the leading methods most often. Monte Carlo simulation also performed well. The hybrid method produced the single largest weekly score, 83 points in gameweek 27, one point above the Monte Carlo team.

ARIMA and linear regression fluctuated without a clear winner. Robust optimization did not reliably improve the average-based models. More complexity changed the selections, but it did not guarantee a better score.

5 GWsWeighted average led most often
82Best Monte Carlo week
The paper describes weighted averages as the safer choice, closely followed by simulation.

A model can be mathematically advanced and still be weak at the only test that matters: future points.

nil nil interpretation
03

The scoring target matters

The study also compares different ideas of player quality. The official ICT index generally beat objectives based only on expected goals scored and conceded. A hybrid that weighted actual total points more heavily than predicted points also did better than versions that trusted the prediction more.

This does not prove that ICT is the best universal FPL signal. It shows that a broad measure aligned with the game's scoring rules can beat a narrow football statistic when the objective is fantasy points.

04

Optimization cannot rescue a weak forecast

Integer programming finds the best team for the numbers it receives. It cannot tell whether those numbers describe the future well. When forecasts are noisy, a precisely optimized team can still be precisely wrong.

The practical hierarchy is clear. Test the forecast first. Then optimize the squad. Finally, compare the complete decision system against a simple baseline that a manager could actually use.

Limits

A short test window

This preprint evaluates methods on part of the 2023/24 season and focuses on a starting eleven rather than the full fifteen-player squad and season-long transfer path. Weekly FPL scores are volatile, so a handful of wins does not establish lasting superiority. The reported results are best read as a comparison of candidate methods, not a finished winning system.

Read the original research ↗