What Are Expected Goals in Football? A Practical Introduction to xG
Expected goals (xG) estimates the probability that a shot will become a goal. A model…

A precise decimal can hide a surprising amount of judgment.
A striker shoots from eight yards and the match graphic flashes 0.42 xG. To one viewer, it looked almost impossible to miss; to another, a defender’s pressure made it a difficult chance. The decimal appears to settle the argument—but it does not come directly from the laws of football.
An xG value is a model’s estimate, based on how often similar shots became goals in its historical data. Providers may define “similar” differently, use different datasets, or include different details: goalkeeper position, defensive pressure, shot height, body part, and the type of preceding pass. The same attempt might therefore receive 0.35 from one model and 0.47 from another. That variation is not necessarily an error. xG is objective in calculation, but conditional on subjective modelling choices.
A shot’s expected goals (xG) value is an estimate of the probability that the attempt will become a goal, based on similar shots in the model’s data. A value of 0.20 xG means a 20% scoring chance—or, in everyday terms, roughly one goal from every five comparable attempts over many repetitions. The broader idea behind expected goals uses this probability-based view rather than judging chances only by whether they went in.
That does not mean every fifth 0.20 xG shot must score. Any individual attempt has a binary result: goal or no goal. A goalkeeper’s save, a deflection, or a slight finishing error can decide the outcome without changing how promising the chance was beforehand.
The probabilities become more informative when combined. If a team takes shots worth 0.05, 0.20, and 0.35 xG, its match total is 0.60 xG. Adding shot values across matches produces season-level xG, which can reveal how consistently a team creates or concedes quality chances—even when short-term results are noisy.
An xG model begins with a large dataset of previous shots. Each attempt is labeled with its actual outcome—goal or no goal—and described by variables such as location, angle, body part, assist type, defensive pressure, and whether it was a header or one-on-one.
The model then looks for repeated relationships between those variables and scoring outcomes. If shots from a particular area, angle, and situation have historically produced goals 12% of the time, a similar new attempt may receive a value near 0.12 xG. The exact calculation depends on the provider and modelling method; logistic regression and machine-learning algorithms can identify different combinations of patterns.
This means an analyst does not watch every shot and assign a value by personal judgment. Once trained, the model applies its learned relationships to the recorded details of each new attempt. Human decisions still matter when selecting variables, defining events, and cleaning the data, but the shot value itself is generated from observed scoring patterns across many comparable attempts rather than a visual verdict on one moment.
The strongest input is usually shot location. Coordinates establish the distance from goal and the angle available between the posts: central, close-range attempts generally offer a larger target than shots from tight positions. The relationship is not linear, however, as explained in more detail by how distance and shooting angle shape xG.
A model may also consider:
These features interact rather than adding fixed bonuses or penalties. A header is not automatically worth a set amount less than a footed shot: a header from two metres after a precise cross may remain an excellent chance. Likewise, a central shot can be poor if it comes from long range or under heavy pressure.
Standard event data records what happened and where: the shot coordinates, body part, assist, play type and preceding action. Richer tracking data adds where players were at that moment. It can reveal nearby defenders, blocked sightlines, defensive pressure, goalkeeper position and how much of the goal was exposed.
Consequently, two attempts from the same coordinates may receive different estimates in a tracking-based model, while an event-only model may treat them as nearly identical.
Location, body part, assist type, game state, pressure, and other recorded details are converted into numerical inputs. Categories such as “header” or “through ball” are represented in a form the model can process.
In logistic regression, each input has a weight learned from historical shots. Factors associated with more goals push the internal score upward; factors associated with fewer goals pull it downward.
The model starts from a baseline and combines all relevant adjustments. A central footed attempt might gain value, while a sharp angle or heavy pressure reduces it.
A logistic curve converts the internal score into a probability. The result might be 0.08, 0.35, or 0.76 rather than an unrestricted number.
Models are tested and often calibrated on separate data. Among well-calibrated shots rated near 0.20, roughly one in five should result in a goal over a large sample.
Decision trees, boosted trees, neural networks, and related methods can learn more complicated relationships. They still produce the same kind of final output: an estimated scoring probability between 0 and 1.
An xG value cannot usually be reconstructed from distance and angle alone. The exact result depends on the provider’s variables, training data, model design, and calibration choices—even when two shots look identical on screen.
Consider a right-footed shot taken from the centre of the penalty area, about 12 metres from goal. The ball arrives from the byline after an open-play cutback, and the attacker strikes it first time.
The model converts those details into features: coordinates, distance, shooting angle, body part, assist type, phase of play and any other context the provider records. It then compares the attempt with patterns learned from similar historical shots. Suppose the resulting estimate is 0.25 xG.
That means comparable attempts are expected to produce a goal about 25% of the time—roughly one goal per four such shots over a large sample. It does not mean this particular attempt is one-quarter of a goal, nor that four repetitions would guarantee one.
The estimate can move when the description changes:
There is no dependable rule such as “subtract a set amount for a header.” Each feature interacts with the others, and whether defensive pressure is included in the xG model can further change the result.
One provider might report 0.25, another 0.24, and another 0.18 without any calculation being obviously wrong. They may train on different competitions and seasons, define cutbacks or rebounds differently, record coordinates on slightly different pitch scales, or include different contextual variables. Their modelling methods, treatment of blocked shots and data-cleaning choices can also produce reasonable disagreement.
No. It is a model estimate based on past attempts judged comparable. Unrecorded details can still matter.
Not necessarily. Distance, body part, pass type, pressure and goalkeeper position may all qualify the advantage of a central location.
Either is possible depending on the full situation and model. A rebound may expose the goal, but it can also leave a difficult angle or awkward body position.
Consistency matters more than choosing the largest value. Comparisons are clearest when every shot comes from the same provider and model version.
Penalties are usually separated from open-play shots because their starting conditions are unusually consistent: the ball is on the spot, the goal is unobstructed, and only the goalkeeper may defend. Providers can therefore estimate scoring probability from a large historical sample of penalties rather than rebuilding the chance from location and angle.
That produces a fairly stable value, often around 0.75 to 0.80 xG, though the exact figure depends on the dataset and rules used. Some providers apply one fixed number; others adjust the xG value assigned to penalties for selected circumstances.
Ordinary models may miss or only partly represent factors such as:
Several of these details are difficult to measure reliably. Others, especially placement and power, describe how the penalty was executed and may only become known after the shot, making them unsuitable for a pre-shot estimate.
A penalty worth 0.78 xG still misses roughly 22 times in 100 comparable attempts. The figure describes the average outcome of similar chances—not the quality of the strike, the tactical importance of winning the penalty, or certainty that this attempt will score.
Shots closer to goal usually receive more xG, especially from central areas.
A tight angle leaves less visible goal and typically lowers the estimate.
Headers and weak-foot attempts may convert differently from comparable strong-foot shots.
One-on-ones, rebounds, set pieces and open-play attempts belong to meaningfully different comparison groups.
A through ball, cross, cutback or defensive error can change how much time and space the shooter had.
Some models know only event data; others include goalkeeper position, defensive pressure or player tracking. Extra context can explain apparently surprising values.
Small differences between providers are normal because their data and model choices differ.
xG is best read as an estimated probability, not a verdict on finishing. A trained model compares the shot’s encoded circumstances with historical attempts, producing a value that is most useful for comparing chances and patterns over time.