Why Do xG Models Disagree on the Same Match?

Andy
August 28, 2026
5 Views
Why Do xG Models Disagree on the Same Match?
One match, two answers

The final whistle goes, one app reports 1.8–0.9 xG, and another shows 1.2–1.1. Both watched the same shots, so the natural suspicion is that one has made a mistake. Usually, neither total is plainly “wrong.” Expected goals are estimates, not measurements such as possession time or the final score.

Our Top Sportsbook Bonuses for August 2026

Use Code: BTCSWB750

Bovada Welcome Offer

5/5
Up To $750 Bonus
18+ only. Full terms apply.
Use Code ACES250

Vegas Aces Casino Welcome Offer

5/5
250% Up To $1,000 Deposit Bonus
18+ only. Full terms apply.
Use Code WELCOME1

Slots of Paradise Welcome Offer

5/5
250% Up To $2,500
18+ only. Full terms apply.
Load More - Link

Disagreement can begin with the event feed itself: providers may record slightly different shot locations, label chances differently, or apply separate rules to blocked efforts, rebounds and penalties. The models then add another layer. Each may weigh angle, distance, body part, defensive pressure or assist type differently, depending on its training data and design. Even rounding and later data corrections can shift a displayed total. The useful question is therefore not simply which number is correct?, but what assumptions produced each number?

What an xG number actually represents

Expected goals is not recorded directly like shots, fouls or goals. It is an estimate produced by a statistical model. For each attempt, the model compares features such as location, angle, body part and assist type with similar historical shots, then assigns a probability of scoring. The fundamentals of expected goals in football begin with this distinction: xG describes shot quality, not what actually happened.

A shot valued at 0.20 xG is one the model expects to be scored about 20% of the time across comparable situations. It does not mean one-fifth of a goal was observed, nor does it guarantee that every provider will assign the same value.

A team’s match xG is usually the sum of its individual shot values. Three attempts worth 0.10, 0.25 and 0.40 xG produce a total of 0.75 xG. That total is not necessarily the probability of scoring at least once; it is the combined expected goal contribution of the shots.

Because providers choose their own data, definitions and modelling methods, football has no single official xG figure. Each published total is the output of a particular model.

Event collection

The data can differ before the model runs

Providers do not always build the same list of shots from the same match.

An xG model cannot assess a chance that its data provider never recorded as a shot. Before any probability is calculated, analysts or automated systems must decide what happened, where it happened and which event category applies.

Common differences include:

  • Blocked efforts: One feed may count an attempt blocked near the striker, while another treats it as a failed touch or pass.
  • Shot location: Coordinates may mark the first contact, the final contact after a deflection, or an estimated position derived from video.
  • Deflections and rebounds: A redirected effort might remain the original player’s shot, become a separate attempt, or be classified as an own goal. A quick rebound may also be missed or merged into the preceding event.
  • Special cases: Own goals are often excluded from xG because no attacking shot is credited. Penalty-shootout kicks are usually excluded from match totals, but some displays handle them differently.
  • VAR decisions: A shot following an offside, foul or handball may appear in a live feed and later be deleted. Providers can update at different speeds, leaving temporary—or occasionally lasting—disagreement.

These choices matter most in low-scoring matches. If one provider records a close-range rebound worth roughly 0.55 xG and another omits it, a team’s total might differ by more than half a goal. That gap can look like a modelling dispute even though the models were given different matches on paper.

Model inputs

What the model sees

Extra detail can sharpen an estimate—or introduce unreliable assumptions.

Two models can receive the same shot record yet describe the chance differently. A basic model may rely mostly on distance and shooting angle. Richer versions can add context that materially changes the probability:

  • Body part: headers are generally converted differently from shots with the foot.
  • Assist type: a cutback, cross, through-ball, or rebound creates a different chance profile.
  • Phase of play: open play, corners, free kicks, and counterattacks produce distinct situations.
  • Goalkeeper position: an exposed goal is not equivalent to a keeper set on the line.
  • Defender proximity: a nearby challenge can reduce shot quality.
  • Obstructions: players between the ball and goal may block part of the target.

Location-only models may therefore rate two shots identically even when one is a clean attempt and the other is taken through several bodies. A model with reliable tracking data can separate them more convincingly.

More inputs are not automatically better, however. Defensive pressure is especially difficult to capture consistently: providers may infer it from event labels, measure the nearest defender at a chosen instant, or estimate whether a challenge affected the shooter. Those methods can disagree when defenders are moving quickly or the exact strike frame is unclear.

A noisy pressure variable may make a model look sophisticated while weakening its consistency across leagues, cameras, or data collectors. Understanding how models account for defensive pressure is therefore more useful than simply counting their features. A smaller model built from stable inputs can be more dependable than a detailed one fed uncertain observations.

Training data leaves a fingerprint

A model’s past shapes the probabilities it assigns today.

An xG model learns from the football in its training set. A system trained mostly on recent Premier League matches may find different patterns from one spanning several countries, divisions and older seasons. Playing styles, data quality and law changes can all shift those relationships.

Coverage also controls sample size. Central shots provide thousands of examples; overhead kicks, extreme angles, unusual rebounds and goalkeeper errors may provide very few. Estimates for these rare chances are less stable, so providers may group or smooth them differently.

Different learning choices

Recency is another choice. Giving newer matches extra weight follows modern tactics but increases sensitivity to short-term patterns. A larger historical sample is steadier, yet may react slowly.

Architecture determines how examples are combined. A simple model may learn broad averages for distance and angle; trees or neural networks can capture interactions, such as angle behaving differently for pressured headers. Regularization and feature selection decide whether a small pattern is signal or noise.

Finally, calibration sets the probability scale. One provider may ensure that shots rated 0.15 score roughly 15% of the time across all validation data, while another calibrates separately by league or shot type. Given different evidence and design choices, 0.12 and 0.18 can both be defensible estimates for the same shot—not proof that either model is broken.

Measure matters

Similar labels can hide different questions

Different measures
Every xG figure estimates the chance that a shot becomes a goal.
Pre-shot and post-shot values should not be compared as interchangeable estimates.
Not guaranteed
A team total and a player total should always add up identically.
Matching totals depend on both views including and assigning the same events.
Different question
Goalkeeper xG is simply the opponent’s attacking xG copied across.
Goalkeeper metrics usually evaluate shot-stopping rather than the full quality of chances conceded.
Worked example

How small differences add up

A worked match total shows why a few hundredths matter.

Consider a match containing six ordinary attempts, one clear central chance and a penalty. Two providers might assign the following values:

AttemptModel AModel B
Six ordinary shots0.500.64
Major chance0.360.45
Penalty0.760.79
Displayed match total1.611.89

Across the six routine shots, Model B is only about 0.02 higher per attempt. That barely looks noteworthy on a shot map, yet it creates roughly 0.14 xG in aggregate. The major chance adds another 0.09 to the disagreement, while the penalty contributes 0.03.

The displayed rows need not add perfectly. Providers commonly calculate totals from unrounded internal probabilities, then round each visible shot and the final sum separately. Values such as 0.495 and 0.635 could appear as 0.50 and 0.64, while producing slightly different totals behind the scenes.

Penalties are another small complication. One model may assign every standard penalty a fixed historical conversion rate; another may use a learned estimate that changes with its training sample or treatment of penalty events. A few hundredths there can widen an already noticeable gap.

Check the shots before the scoreline

Before treating a 0.28 xG gap as evidence that one model is wrong, inspect the shot map. Look for many small differences, one heavily valued chance, the penalty rate, and any missing attempt.

Comparison checklist

How to compare xG providers fairly

  • Confirm the metric

    Check whether both figures are pre-shot xG, post-shot xG, open-play xG, or another variant. Similar-looking labels may measure different questions.

  • Align the shot lists

    Compare events one by one, including penalties, blocked attempts, rebounds, own goals, and shots later removed by review. A FotMob–Understat xG comparison is meaningful only when both totals cover the same attempts.

  • Check locations and classifications

    Look for coordinate differences and disputed labels such as headers, through balls, set pieces, or big chances. Video can help resolve obvious event-recording disagreements.

  • Find the largest shot-level gaps

    Sort corresponding shots by the difference in assigned probability. One major chance often explains more than a dozen small disagreements.

  • Rebuild the totals

    Add the matched shot values and separate any unmatched events. Only then judge whether the remaining gap reflects model design rather than different underlying data.

  • Keep one benchmark over time

    For season-long tracking, use the same provider consistently. Switching sources because one total better fits a preferred match story creates a misleading comparison.

Provider totals can both be reasonable without being interchangeable.

Conclusion

Use xG as a Range, Not a Ruling

When providers still disagree after like-for-like checks, the most useful answer is usually a range. A match reported as 1.4–1.7 xG tells a sturdier story than treating 1.62 as a definitive measurement. The shot list and broad chance-quality pattern matter more than the second decimal place.

For casual analysis, the best provider is often the one with clear metric definitions, visible event rules and reliable coverage across matches. Richer feeds become more valuable when analysing large datasets, reproducing published work or studying details such as goalkeeper position and defensive pressure. In those cases, whether paid xG data justifies the cost depends less on a single model’s apparent precision than on access, documentation, consistency and exportable underlying data.

Author Andy

Hi I'm Andy and I love to report on the latest football scores and Tables. I also like to have a bet on the football and occasionaly on the horses. On this website I have new bookmaker offers listed that will give you free bets and bonuses to help you beat the bookies. Enjoy your stay.

0 0 votes
Article Rating
Subscribe
Notify of
0 Comments
Oldest
Newest Most Voted