Why Football Statistics Differ Between Data Providers
A practical framework for understanding why football statistics vary between providers and how to reconcile conflicting numbers without assuming one is wrong.
Football statistics differ between data providers because the numbers are produced through a chain of human definitions, data collection, coordinate systems, model choices and later corrections. Two sources can analyse the same match and report different totals for shots, expected goals (xG), expected assists (xA), PPDA or possession sequences without either source necessarily being wrong.
The key question is not simply which number is correct. It is whether the providers measured the same event, used the same inputs and applied the same calculation. Analysts should compare definitions and methodology before treating a difference as evidence of an error.
Why Can Football Data Providers Report Different Statistics?
A football statistic is not a direct copy of reality. It is a structured representation of what happened. A provider must decide where one event begins, how it should be labelled, where it occurred and whether later actions belong to the same sequence.
Simple match facts can therefore contain judgement. Was a deflected attempt a shot, cross or own goal? Did a defender block an on-target effort as the last line, or was the goalkeeper still able to intervene? Did a receiver retain enough control for a subsequent shot to remain connected to the original pass?
Derived metrics add another layer. An expected-goals model converts recorded shot information into a probability. If providers record or model that information differently, their xG values can differ even when they agree that the shot occurred.
The Seven Main Causes of Provider Variation
1. Event definitions
Providers need operational rules for shots, assists, tackles, interceptions, crosses, pressures and many other actions. Small definitional differences can change totals.
Stats Perform's published Opta event definitions, for example, classify a deliberate on-target attempt stopped by a last-line defender as a shot on target. An attempt blocked by an outfield player with other defenders or the goalkeeper behind the blocker is classified differently. A viewer may describe both incidents as blocked shots, but the data rules separate them.
This distinction is especially important when interpreting player-shot statistics and settlement labels. The displayed total is tied to the nominated provider's definition, not every reasonable interpretation of the television pictures.
2. Human collection and review
Event data can be collected live by trained operators, supported by video and automated systems. Marginal incidents require judgement, and the first live classification may later be reviewed.
Two operators may initially disagree about whether a touch was a pass, interception or loose-ball recovery. Quality-control teams can then amend the event after examining another angle. This means two websites may disagree temporarily even when one ultimately receives corrected data from the same underlying supplier.
3. Pitch coordinates and spatial precision
Location-based analysis depends on where events are recorded. Providers may use different pitch dimensions, coordinate scales or conventions for the point attached to an action. One system might represent a pitch from 0 to 100, while another uses 120 by 80 or physical metres.
Coordinates can be transformed, but conversion does not remove collection uncertainty. A shot recorded one or two metres closer to goal can receive a different angle, distance and xG value. The effect is also visible in shot maps, where apparently different dots may describe the same attempt on differently scaled pitches.
4. Model inputs and assumptions
xG, xA, post-shot xG and possession-value figures are model outputs rather than universally fixed event counts. Providers choose which variables to include, how to train the model and how to treat unusual situations.
A basic xG model may use location, angle, body part and assist type. A richer model may also include goalkeeper position, defenders between the shooter and goal, pressure or shot impact height. Hudl StatsBomb explains that its xG data can incorporate goalkeeper and defender locations captured in shot freeze frames. Different information sets can legitimately produce different estimates for the same chance.
The same applies to expected assists. One approach assigns the later shot's xG to the creator; another estimates the probability that a completed pass becomes an assist. Both may be called xA, but they answer different questions.
5. Possession and sequence rules
Metrics based on chains of play require rules for when a possession starts, ends or changes hands. A deflection, duel, unsuccessful touch or momentary interception may split one provider's sequence while another treats play as continuous.
Those decisions affect possession counts, sequence length, direct attacks, build-up speed and possession-value models. They can also change whether an earlier pass receives credit for a later shot. A difference in possession-chain data may therefore originate in the segmentation rule rather than the visible actions themselves.
6. Metric formula and eligible actions
Even apparently transparent calculated metrics can vary. PPDA is broadly opposition passes divided by qualifying defensive actions in defined areas of the pitch. Providers or analysts can use different pitch zones and include different defensive events in the denominator.
Before comparing two PPDA figures, check the permitted passes, defensive-action list and field boundary. Matching labels do not guarantee matching formulas.
7. Model versions and correction timing
Data products change. Providers improve collection rules, correct events and retrain models. Historical figures may be recalculated when a new xG version is introduced, while another platform retains the value first published after the match.
Hudl StatsBomb's published discussion of an xG upgrade describes changes to blocker and goalkeeper positioning, long-shot behaviour and post-shot inputs. A model update can therefore alter historical outputs without the underlying match changing.
Timing matters as well. A live feed, a post-match report and a database checked several days later may represent three different revision states. Always record the provider, extraction time and model version where available.
A Same-Match Comparison Framework
Suppose Provider A reports 14 shots and 1.42 xG, while Provider B reports 13 shots and 1.18 xG for the same team. The figures alone do not identify the cause. Use this framework before drawing a conclusion.
| Check | Question to ask | Possible explanation |
|---|---|---|
| Event inventory | Do both sources include the same 13 or 14 attempts? | One incident may be classified as a cross, blocked attempt or own goal rather than a shot. |
| Coordinates | Are the shot locations aligned after transforming the pitch scales? | A location difference can change distance, angle and model value. |
| Context | Which body part, assist type, pressure and player-position inputs are recorded? | One model may see information that the other does not use. |
| Model scope | Are penalties, rebounds and blocked shots treated the same way? | The providers may calculate different eligible shot populations. |
| Version and time | When were the figures collected and which model version produced them? | One feed may contain a correction or recalculated historical value. |
First reconcile the event count. If Provider B excluded one attempt worth 0.04 xG, that explains only part of the 0.24 difference. Match the remaining shots by time and location, then compare their individual values. Larger gaps on cut-backs, headers or crowded shots may indicate different contextual inputs rather than an arithmetic error.
The same method works beyond xG. For possession chains, align the underlying events before comparing sequence metrics. For xA, establish whether the metric is shot-based or pass-based. For post-shot xG or xGOT, check which attempts qualify and whether shot placement, velocity, goalkeeper location or trajectory is available.
How Analysts Should Reconcile Conflicting Football Data
- Name the provider. Never store an unexplained column simply called xG, xA or PPDA.
- Preserve one consistent source for longitudinal analysis. Switching providers mid-season can create artificial changes in performance.
- Save the definition and extraction date. Record the methodology or data dictionary version where possible.
- Compare raw events before derived totals. Establish whether the disagreement begins with collection, classification or modelling.
- Normalise carefully. Convert coordinate systems, time conventions and units without assuming that differently defined events are equivalent.
- Use video for material discrepancies. Review the relevant sequence, but apply the provider's operational definition rather than intuition alone.
- Run sensitivity checks. Ask whether the analytical conclusion changes under either provider's figure.
If a team's match xG is 1.42 in one dataset and 1.18 in another, the decision may remain unchanged: both describe a modest attacking performance. If a recruitment ranking or probability forecast changes substantially, the source difference becomes a material model risk requiring investigation.
Common Mistakes When Comparing Providers
- Assuming one source must be wrong: the metrics may have different definitions or information sets.
- Mixing providers in one trend: an apparent improvement may be a break in methodology.
- Comparing rounded totals: displayed figures can conceal meaningful differences at event level.
- Treating model estimates as observed facts: xG and xA are estimates conditional on a particular model.
- Ignoring corrections: live data can change after review.
- Choosing the preferred number: selecting whichever figure supports an existing opinion creates confirmation bias.
- Overreacting to small gaps: numerical disagreement does not always change the football conclusion.
What Provider Disagreement Can Reveal
Differences are not merely an inconvenience. They can expose which assumptions drive an analysis. If a shot receives much higher xG only when defender and goalkeeper positions are included, the disagreement highlights the importance of pressure and goalmouth obstruction. If xA totals diverge because one model values every completed pass, the comparison reveals different ideas of creative contribution.
Provider variation should be treated as one source of measurement uncertainty. It sits alongside sampling variation, tactical context and model error. The practical discipline is similar to separating signal from noise in football data: identify the mechanism, test whether it persists and avoid claiming more precision than the inputs justify.
GoalIQAI Interpretation
No football statistic is independent of its definition and production process. Raw event counts are constructed through classification rules; advanced metrics add modelling assumptions; live figures may later be corrected.
For descriptive reporting, naming the provider and acknowledging modest variation may be enough. For modelling, recruitment or market analysis, provenance should be part of the dataset: source, definition, coordinate convention, collection time and version. Consistency usually matters more than selecting the provider with the most appealing total.
When two credible sources disagree, treat the gap as a diagnostic prompt. Reconcile the underlying events, identify the methodological cause and test whether the decision survives both measurements. That produces a stronger conclusion than declaring one figure correct without examining how either was made.
Key Takeaways
- Football statistics can differ because providers use different event definitions, collection judgements and correction processes.
- Coordinate systems and recorded locations can change spatial metrics and model outputs.
- xG, xA, post-shot xG and possession-value figures depend on provider-specific inputs, training data and assumptions.
- Possession chains and PPDA can differ because sequence rules, pitch zones and eligible actions are not always identical.
- Model upgrades and post-match corrections can change figures after their first publication.
- Analysts should preserve provider provenance, compare underlying events and test whether discrepancies alter the decision.
- Different numbers do not automatically mean that one source is wrong.
Related Guides
Stay Ahead of the Market
Subscribe to GoalIQAI for evidence-based football analysis, practical guides to football data and clear explanations of how analytical models should—and should not—inform decisions.