Research / FPL prediction

Can the crowd predict FPL?

Statistics, betting markets, public opinion, and one striking season

Abstract

Most FPL forecasts begin with the player: minutes, goals, assists, form, and fixtures. This paper asks whether useful information also sits outside the official statistics. It combines performance data with betting markets, fixture difficulty, social discussion, and published expert opinion.

In its 2018/19 test, the combined system beat conventional statistical predictors by more than 300 points. The result suggests that the crowd can carry information, but it also raises a harder question: how do we separate real signal from noise and hindsight?

01

Four views of the same player

A player's historical score describes what happened. Betting prices summarize a market's estimate of what may happen. Injury reports and team news change expected minutes. Public and expert discussion reacts to information that may not yet appear in a numerical record.

The proposed system collects these streams, turns the text into sentiment signals, and joins them with player history and fixture context. The final forecast is therefore not purely statistical or purely human. It is an ensemble of both.

  1. 01Historical FPL performance
  2. 02Fixtures and betting markets
  3. 03Public and expert opinion
  4. 04Combined player forecast
Independent information streams are brought into one weekly estimate.

The crowd is useful when it knows something the table does not.

nil nil interpretation
02

A large reported advantage

Tested on the 2018/19 Premier League season, the system finished with more than 300 points above the comparison statistical predictors. The authors describe the gain as roughly eleven points per week and a final rank near 30,000 from more than 6.5 million managers.

That is large enough to demand attention. It is also large enough to demand careful replication. A single season can reward choices that fail later, and public opinion may encode information that was only available after a clean historical timestamp.

+11Average weekly advantage
Top 0.5%Reported final rank band
Performance reported by the authors for the 2018/19 test season.
03

Markets and crowds do different work

Betting markets aggregate money and professional pricing. Social discussion aggregates attention, rumor, observation, and repetition. They should not be treated as interchangeable signals.

A good model needs to know whether a signal adds information after the existing forecast is considered. Popularity alone can merely repeat recent points. A market price may add a forward view of team strength. Team news may mainly change the chance of starting. Each stream earns its place only if it improves a future test.

04

The model points toward a product

The commercial lesson is direct. An FPL decision tool should not force a reader to choose between data and context. It should show the statistical expectation, then state which live facts moved it.

That makes the forecast easier to challenge. It also prevents a score from hiding the reason behind a recommendation. The useful output is not only who to buy, but what new information changed the estimate.

Limits

One season and a moving information set

This preprint reports a single-season evaluation. Text, news, odds, and social signals are difficult to reconstruct without leakage because their timing matters. The abstract reports the headline comparison, but the result needs a transparent prospective test before it can support a general claim that crowd-informed systems outperform statistical models.

Read the original research ↗