The Moneyball Timeline of Football Analytics: From Notebooks to AI

Trace the evolution of football analytics from handwritten match records and early computer models to xG, tracking data, data-driven recruitment and artificial intelligence.

The history of football analytics began long before the word “Moneyball” entered the sporting vocabulary. Analysts were recording match events by hand in the 1950s, coaches were experimenting with computers in the 1970s, and betting organisations were building statistical models before most clubs employed full-time analysts.

Moneyball nevertheless gave this wider movement a powerful name. It described a way of competing in which organisations use evidence to identify what conventional opinion undervalues. In football, that idea has developed through several stages: manual notation, video analysis, event data, expected goals, tracking technology, data-led recruitment and increasingly sophisticated artificial intelligence.

The important story is not that data replaced football expertise. It is that clubs learned to combine data, scouting, tactics, technology and organisational decision-making.

What does Moneyball mean in football?

Moneyball in football means using evidence to find an advantage that richer or more conventional competitors have overlooked.

The term originates from Michael Lewis’s 2003 book about the Oakland Athletics baseball team. Faced with a much smaller budget than its rivals, Oakland attempted to identify players whose contribution was more valuable than the market believed.

The central principle was not simply “use statistics”. Professional sports had used statistics for decades. The more important idea was to identify a difference between:

  • how the market values a player;
  • how much that player is likely to contribute;
  • and what the organisation must pay to acquire that contribution.

Football’s version is more complicated than baseball’s. Football is continuous, low-scoring and highly interactive. The value of an action frequently depends on the positions and movements of 21 other players. Roles also change across tactical systems, leagues and game states.

Football Moneyball therefore became broader than player statistics. It now covers recruitment, match analysis, tactical preparation, injury prevention, player development, betting-market modelling and the design of decision-making systems.

Before Moneyball: football’s first statistical questions

Football has always produced numbers. Goals, appearances and league points have been recorded since the sport became formally organised.

Those figures described outcomes, but they offered limited explanations of how those outcomes were produced. A final score could show which team won without revealing the balance of chances, territory or control.

Early football analysts began asking a more ambitious question: could the events within a match be systematically recorded and used to understand how football works?

That question created the foundation for modern football analytics.

1950s: Charles Reep and handwritten match analysis

One of the earliest widely documented attempts to analyse football systematically came from Charles Reep, an accountant and former Royal Air Force officer.

In March 1950, Reep began recording events during a match between Swindon Town and Bristol Rovers. Using pencil, paper and a system of symbols, he documented passing sequences, field position and attacking outcomes.

Reep eventually analysed thousands of matches. This was an extraordinary undertaking before digital video and automated data collection. A single match could require many hours of additional work after the final whistle.

His most influential conclusion was that many goals followed relatively short passing sequences. This evidence became associated with direct football and helped influence thinking around territory, possession and attacking efficiency.

Reep’s work was historically important, but it also illustrates a lasting analytical warning.

If most goals follow short sequences, it does not automatically mean that deliberately playing fewer passes causes more goals. Many possessions are short because defenders win the ball close to goal, because a set piece creates an immediate chance or because the final sequence begins after earlier possession has moved the opposition.

The distinction between correlation and causation was not always handled carefully. Reep showed that football could be measured, but he also demonstrated how the interpretation of data can matter as much as its collection.

1960s and 1970s: football meets systems thinking

Football analysis developed along different paths across countries and coaching cultures.

In the Soviet Union, Valeriy Lobanovskyi became one of the most important pioneers of scientific football management. Working with statistician Anatoliy Zelentsov at Dynamo Kyiv, Lobanovskyi approached the team as a complex system rather than simply a collection of individuals.

Training, physical preparation and match performance could be measured. Players had individual responsibilities, but their value also depended on coordinated actions within the team’s structure.

This was a significant conceptual advance. Traditional evaluation often focused on visible individual ability. Systems thinking asked how reliably a player performed the actions required by the collective model.

The available computers were primitive by modern standards, and the data could not capture football in today’s level of detail. However, the underlying philosophy remains highly relevant: analyse repeatable processes, define tactical requirements and evaluate players in relation to the system.

Modern recruitment models still face this challenge. A player’s statistics cannot be separated completely from the role, teammates and tactical environment that produced them.

1980s: video changes match preparation

The growing availability of video transformed football analysis.

Coaches no longer had to rely entirely on live observation and memory. Matches could be replayed, paused and reviewed. Analysts could examine defensive organisation, repeated attacking patterns and individual decisions away from the ball.

Video did not initially produce the large structured datasets associated with modern analytics. Its contribution was different: it made football behaviour observable after the event.

This improved several areas of preparation:

  • opposition analysis;
  • individual player feedback;
  • set-piece preparation;
  • tactical review;
  • and post-match evaluation.

Video remains central today because structured data can identify that something is happening while footage helps explain how and why it happens.

1990s: event data and performance-analysis technology arrive

The 1990s created much of the infrastructure required for the modern analytics industry.

Opta began supplying detailed football statistics in 1996, initially serving media clients including Sky Sports. Instead of recording only goals and results, data companies began documenting actions such as passes, shots, tackles, crosses and fouls.

At approximately the same time, performance-analysis systems such as Prozone began giving clubs a more systematic view of player movement and match activity.

These developments changed the scale of football analysis. Information that previously required an analyst to record events manually could increasingly be collected across entire competitions.

Three important changes followed:

  • players and teams could be compared using consistent definitions;
  • analysts could study large samples rather than isolated matches;
  • clubs could create dedicated performance-analysis departments.

The data was still relatively basic. Totals such as passes completed or distance covered could easily be misinterpreted. Nevertheless, the foundations of a searchable, structured football database had been established.

Early 2000s: Moneyball gives the movement a language

Michael Lewis published Moneyball in 2003. Although the book concerned baseball, its influence spread across professional sport.

Its appeal was particularly strong for smaller organisations. If the wealthiest teams could buy established stars, less powerful clubs needed a different type of advantage. Better valuation appeared to offer one.

Football clubs began to ask whether conventional recruitment contained systematic biases. Scouts and executives might overvalue:

  • players from prestigious clubs;
  • recent performances;
  • visible athleticism;
  • international appearances;
  • goals without considering shot volume or quality;
  • and players who fitted a familiar physical profile.

Data offered another source of evidence. It could expand the number of players considered, challenge subjective impressions and reveal production in less heavily watched leagues.

However, early attempts to import Moneyball into football sometimes misunderstood its lesson. The objective was not to replace scouts with spreadsheets. It was to improve the organisation’s estimate of value.

2000s: betting models develop away from public view

While clubs gradually expanded their analytical departments, professional betting organisations had strong financial incentives to develop football models.

A bettor does not only need to predict what might happen. The bettor must estimate probabilities accurately enough to identify when the available odds differ from fair value.

This encouraged modelling of:

  • team attacking and defensive strength;
  • expected scoring rates;
  • home advantage;
  • opposition quality;
  • team selection;
  • schedule effects;
  • and changes in underlying performance.

The details of successful private systems are not publicly available. Claims about particular variables or algorithms should therefore be treated cautiously.

What can reasonably be inferred is that betting organisations had an important advantage over many early club analytics teams: clear and rapid feedback. A probability estimate could be compared with the market price, the closing line and the eventual result.

That did not remove randomness, but it encouraged disciplined forecasting, model testing and calibration. These principles later became influential in data-driven football ownership, as explained in our history of Tony Bloom, Matthew Benham and the evolution of football modelling.

Late 2000s and early 2010s: expected goals changes the conversation

One of the most important developments in public football analytics was the rise of expected goals.

Traditional analysis often treated all shots as broadly comparable. Expected goals models recognised that a tap-in and a speculative 30-metre attempt do not represent the same scoring opportunity.

Each shot could instead be assigned a probability based on characteristics such as:

  • location;
  • angle;
  • body part;
  • type of assist;
  • phase of play;
  • and defensive pressure, where available.

The resulting xG values made it possible to evaluate chance quality more systematically. Analysts could begin separating finishing outcomes from the process that generated the shots.

A team that lost 1–0 after creating several high-quality chances might have played better than the result suggested. Conversely, a team repeatedly scoring from difficult attempts might be unlikely to sustain the same conversion rate.

Our complete guide to expected goals in football explains how the metric works and why different providers can produce different xG values.

xG did not solve football analytics. It did something arguably more important: it gave clubs, bettors, journalists and supporters a better language for discussing performance beneath the scoreline.

2010s: the public football analytics community expands

Football analytics accelerated during the 2010s as data became more accessible and public analysis became easier to distribute.

Blogs, social media, conferences and specialist companies allowed analysts outside professional clubs to publish models and visualisations. Shot maps, passing networks and expected-goals charts helped make abstract statistical ideas understandable.

This public community contributed in several ways:

  • new metrics could be debated and improved;
  • clubs gained access to a wider pool of analytical talent;
  • supporters learned to question conventional statistics;
  • and football analysis became less dependent on goals, assists and possession percentages.

Providers also collected richer event data. Analysts could study carries, pressures, sequences, possession chains and the location of defensive actions.

Metrics such as PPDA attempted to describe pressing intensity, while field tilt offered a more informative view of territory than overall possession alone.

The central analytical unit was gradually moving from the final outcome towards the processes that produced it.

2010s: data-driven ownership becomes visible

Analytics became more influential when club owners began building entire organisational strategies around it.

Tony Bloom at Brighton & Hove Albion and Matthew Benham at Brentford became the most prominent English examples. Both had backgrounds connected to professional football betting and both developed clubs associated with data-led decision-making.

Their precise private models remain confidential, and the two organisations should not be treated as identical. Their public success nevertheless demonstrated how probabilistic thinking could extend beyond match prediction.

Data could support:

  • player recruitment;
  • manager selection;
  • contract decisions;
  • squad planning;
  • performance evaluation;
  • and the identification of undervalued markets.

Benham became Brentford’s majority shareholder in 2012 and acquired a controlling interest in FC Midtjylland in 2014. Midtjylland used analytics in areas including recruitment and set-piece preparation before winning the Danish championship for the first time in 2015.

Brighton and Brentford later reached the Premier League and established themselves despite operating with fewer financial resources than the competition’s wealthiest clubs.

Data was not the only reason for that progress. Leadership, coaching, recruitment, player development and execution all mattered. The importance of the model was that analysis became part of the organisation rather than an isolated service provided to decision-makers.

Our comparison of Tony Bloom and Matthew Benham examines the similarities, differences and limits of what can be known about their private systems.

2010s: recruitment analytics moves beyond performance totals

Early recruitment analysis often focused on identifying players with attractive statistical output. More mature systems recognised that raw output is heavily shaped by context.

A striker playing for the dominant team in a weaker league may receive more chances than a comparable player elsewhere. A midfielder’s passing numbers may reflect the team’s possession structure rather than exceptional individual creativity.

Recruitment models therefore began attempting to adjust for:

  • league strength;
  • team quality;
  • possession share;
  • player age;
  • position and tactical role;
  • opposition strength;
  • physical development;
  • and the likely transferability of performance.

The analytical objective also became more specific. Clubs were not simply searching for the “best” player. They were searching for a player who could perform a defined role, within a particular tactical system, at an acceptable cost and level of risk.

This illustrates an important difference between recruitment models and betting models. Both may estimate future performance, but they use different targets, time horizons and feedback mechanisms.

Late 2010s: expected threat and possession-value models emerge

Expected goals improved the measurement of shots, but most football actions do not end with an immediate attempt at goal.

A midfielder might break two defensive lines with a pass that creates no shot because the receiver takes a poor touch. A full-back might repeatedly carry the ball into dangerous positions without receiving an assist. Traditional event totals could miss the value of these actions.

Possession-value models attempted to measure how each action changed the probability of a team scoring or conceding later in the possession.

Expected threat, commonly shortened to xT, is one accessible example. It assigns value to moving the ball between areas of the pitch according to how threatening those areas tend to be.

More advanced models can account for the type of action, pressure, player positions, possession history and possible future states.

This represented another important shift in football analytics: from measuring isolated events to estimating how actions change the state of the game.

2020s: tracking data adds the players without the ball

Event data normally records what happens around the ball. Tracking data records where players and the ball are positioned over time.

This allows analysts to investigate questions that event data struggles to answer:

  • Which passing options were available?
  • How much pressure was applied to the ball carrier?
  • Which defender closed an important space without making a tackle?
  • How quickly did a team reorganise after losing possession?
  • Which off-ball run created room for a teammate?
  • How compact was the defensive structure?

Tracking data makes football’s interactive nature more visible. It offers a route towards valuing players who shape an attack or defence without being credited with the final recorded action.

The challenge is scale. Tracking datasets are large, expensive and technically difficult to process. A greater volume of information also creates more opportunities to find patterns that do not generalise.

Better data expands what can be studied. It does not remove the need for careful questions, validation and football understanding.

2020s: analytics becomes an organisational system

The most advanced clubs no longer treat analytics as a single department producing occasional reports.

Information can flow between recruitment, coaching, sports science, medical teams, academy development, finance and executive leadership.

A recruitment decision, for example, may combine:

  • statistical performance;
  • video scouting;
  • physical and medical information;
  • personality and adaptation evidence;
  • tactical suitability;
  • contract details;
  • transfer-market alternatives;
  • and projected resale value.

The advantage comes from connecting these inputs rather than allowing each department to reach an isolated conclusion.

This is the modern football intelligence stack: a system through which raw information becomes contextualised evidence, a forecast, a decision and eventually a feedback signal.

It also explains why hiring data scientists does not automatically create a data-driven club. The organisation must decide how evidence is communicated, challenged and incorporated into real decisions.

2020s and beyond: artificial intelligence enters football analytics

Artificial intelligence is the latest stage in the evolution, but it is not a complete break from the past.

Machine-learning models have already been used to identify complex relationships within football data. Computer vision can help extract player and ball locations from video. Language models can assist with searching reports, summarising evidence and connecting information stored across an organisation.

Potential applications include:

  • automated video analysis;
  • player similarity and role classification;
  • tactical pattern recognition;
  • injury-risk estimation;
  • personalised coaching feedback;
  • scenario simulation;
  • and faster synthesis of scouting and performance information.

AI also creates familiar risks in a more powerful form. A system can reproduce biased historical decisions, mistake correlation for causation or generate confident explanations unsupported by the underlying evidence.

Data quality, version control, model validation and human accountability become more important as automation increases.

The next phase of football analytics will therefore be shaped not only by which organisations build the most advanced models, but by which organisations use them most responsibly and effectively.

Why xG was a beginning rather than an endpoint

Expected goals is sometimes treated as the defining product of modern football analytics. Its real historical importance may be that it showed a wider audience how probabilities could describe football more effectively than outcomes alone.

But xG only evaluates a particular part of the attacking process. It does not fully capture:

  • how the opportunity was created;
  • which players moved defenders;
  • what alternative actions were available;
  • the defensive structure before the shot;
  • how game state changed behaviour;
  • or how repeatable the attacking pattern may be.

Modern systems therefore combine xG with richer event data, tracking information, tactical analysis and contextual evidence. Our guide explaining why xG is not enough explores these limitations in more detail.

The development from goals to shots, from shots to shot quality, and from shot quality to entire possession states illustrates the direction of football analytics: increasingly detailed attempts to understand the process behind the result.

What football can learn from the real Moneyball story

The history of football analytics is sometimes presented as a battle between numbers and human expertise. That framing misses the central lesson.

Every analytical method contains judgement. Someone must decide:

  • which question to ask;
  • which events to record;
  • how a metric should be defined;
  • which variables belong in the model;
  • how uncertainty should be communicated;
  • and when the evidence is strong enough to justify action.

Scouting and data are not natural opponents. A model can search thousands of players consistently, while an experienced scout can investigate role, technique, behaviour and adaptability in greater depth.

The strongest process uses each source for what it does well.

The other lasting lesson is that an advantage is temporary. Once a useful metric or recruitment method becomes widely adopted, the market begins to price it more efficiently. Organisations must keep learning, testing and finding new questions.

That is why modern football intelligence increasingly extends beyond public metrics and standard event data.

A simplified timeline of football analytics

  • 1950s: Charles Reep begins systematically recording football matches by hand.
  • 1960s–1970s: scientific coaching and systems thinking develop, most notably through Valeriy Lobanovskyi and Anatoliy Zelentsov.
  • 1980s: video becomes more important for tactical review and opposition analysis.
  • 1990s: Opta, Prozone and other providers help establish structured event and performance data.
  • 2003: Michael Lewis publishes Moneyball, giving evidence-based player valuation a global sporting identity.
  • 2000s: professional betting models and club analysis develop largely away from public view.
  • Late 2000s–early 2010s: expected goals becomes an influential method for measuring chance quality.
  • 2010s: public analytics communities, specialist providers and club departments expand rapidly.
  • 2010s: Brighton, Brentford, FC Midtjylland and other clubs make data-led organisational strategies more visible.
  • Late 2010s: possession-value metrics expand analysis beyond shots and final actions.
  • 2020s: tracking data, computer vision and integrated intelligence platforms deepen analysis of space and off-ball behaviour.
  • Present and future: AI accelerates data processing, pattern recognition and organisational learning while creating new governance risks.

Key Takeaways

  • Football analytics began decades before the publication of Moneyball.
  • Charles Reep demonstrated that match events could be recorded systematically, while also showing the danger of drawing causal conclusions from descriptive patterns.
  • Valeriy Lobanovskyi helped establish the idea of a football team as a measurable, interconnected system.
  • Video, event-data providers and performance-analysis technology created the infrastructure for modern club analytics.
  • Moneyball popularised the idea of using evidence to identify players and characteristics undervalued by the market.
  • Expected goals shifted attention from final scores towards the quality of the process behind them.
  • Data-driven clubs extended analytical thinking into recruitment, manager selection, squad planning and organisational design.
  • Tracking data and possession-value models allow analysts to study space, movement and actions before the final shot.
  • Artificial intelligence can increase analytical scale, but it cannot remove problems involving data quality, bias, uncertainty and human accountability.
  • The strongest football organisations combine models, scouting, tactical knowledge and disciplined decision-making.

Understand football through evidence, not hindsight

Subscribe to the GoalIQAI newsletter for educational guides to football analytics, probability and betting-market intelligence. Learn how to assess the process behind the result and make more disciplined decisions under uncertainty.