Officiating by Model
Decision-support systems have moved from measuring positions to classifying actions. That crosses a line between fact-finding and judgement that the laws have never addressed.
Photograph: Hullian111 · CC BY-SA 4.0 · Wikimedia Commons
The first generation of officiating technology answered questions of fact. Did the ball cross the line? Was the attacker ahead of the defender at the moment of contact? These are geometric questions with determinate answers, and a well-calibrated measurement system answers them more reliably than a human can.
The systems now entering trial answer a different kind of question. Given this contact, was it a foul? Given this handball, was the arm in an unnatural position? Given this challenge, was it reckless, careless, or excessive force? These are classification problems, and a classifier does not return a fact. It returns a probability.
A confidence score is not a decision
This is the conceptual difficulty and it has not been addressed by any governing body that has deployed these systems. When a model reports eighty-one per cent confidence that an action constituted excessive force, what has it told the official?
It has not identified a fact about the match. It has reported a property of its training distribution — that actions with these characteristics were labelled as excessive force by human annotators roughly eighty-one per cent of the time. That is genuinely useful information. It is not the same thing as the action having been a foul, and the laws of the game contain no provision that assigns meaning to a probability.
The model does not tell you what happened. It tells you what similar events were previously called.
The threshold is a policy decision in disguise
Deploying such a system requires choosing a threshold above which the recommendation is acted upon. That choice is presented as technical calibration. It is a policy decision about the trade-off between false positives and false negatives, and the two are not symmetric in their consequences.
A threshold set to minimise missed serious fouls will produce additional incorrect dismissals. One set to minimise incorrect dismissals will miss serious fouls. Where to sit between those errors is a question about what the sport values, and it is currently being answered by whoever specifies the system, without the question being publicly posed.
The training data reproduces past officiating
A further problem follows from how these models are built. They are trained on historical decisions labelled by human officials, which means they learn to reproduce the distribution of past officiating, including its inconsistencies.
If a particular category of challenge has historically been penalised more readily in one competition than another, a model trained across both will encode a blend that matches neither. If historical officiating applied different thresholds to different competitions, player profiles or match situations, the model will reproduce those patterns and present them with the authority of a computed output.
This is well understood in machine learning generally and it has received almost no attention in sporting deployment, partly because the historical decision archives are not published in a form that would permit independent analysis.
Accountability dissolves
The most serious consequence is institutional. Officiating accountability currently rests on a chain: an official made a decision, the decision can be reviewed, and the official can be assessed, retrained or removed.
Insert a model into that chain and the responsibility becomes genuinely unclear. If an official follows a recommendation that proves wrong, the error belongs partly to the model, partly to the threshold-setter, partly to the labellers of the training data, and partly to the official who chose to defer. If an official overrides a correct recommendation, they will be criticised for ignoring the system — which creates pressure to defer even where their own judgement is better informed.
That pressure is the outcome to watch. A decision-support system that officials feel unable to override is not decision support. It is decision-making, and it is being introduced without the governance framework that a change of that magnitude would ordinarily require.