All data is wrong, some of it is useful

Data companies want to talk to you about their 'line-breaking pass' definitions.

In the last month or so, at least two of them have posted about how they define their particular statistic, with another being roped in by a neutral party as a (negative) comparison. Analytics can seem frightening to outsiders, but little do they know the truth of the difficulty in nailing down concepts like 'line', 'breaking', and, indeed, 'pass'.

Truth be told, it's not as if football is the only place where you can experience the inexactitude of data offerings. There are a rash of YouTube videos comparing the step counts of various wearables; whose view counts themselves have recently been re-engineered for ~reasons. And, like, have you ever seen political polling data?

Unsurprisingly, there is an existing literature on how to think about data quality. While it's not something Get Goalside can claim as a particular area of expertise, there are two heavily-cited 1996 works on the matter -- yes, we're going back to the dawn of time -- co-authored by the same guy, Richard Y. Wang. There is 'Beyond accuracy: What data quality means to data consumers' with Diane Strong, based on survey responses; and 'Anchoring data quality dimensions in ontological foundations' with Yair Wand, which takes a more theoretical starting point.

The simplest framework is the latter's, which draws out four elements of data quality: it should be complete [as in 'a complete representation of the real world'], unambiguous, meaningful, and correct. This may put football, a fuzzy grey sport from a fuzzy grey isle, at an immediate disadvantage. A provider which collects passes and nothing else may be unambiguous (and probably pretty correct in its collection), but doesn't represent a complete picture of real world football. But a provider who tries to collect every action, aiming for completeness, may fall short on being unambiguous, meaningful, and correct in its collection.

(There are, of course, data sources that are neither complete, unambiguous, nor reliably correct, a panda-esque collection of qualities that raises the question of how long they can stave off extinction if left to their own devices).

However, there's a part of Wang's other '96 paper, 'Beyond accuracy', which flags a particular reason why some of these criteria may not be met: accessibility. You need to have access to the data! And so again we come back to the complexity of football: don't let perfect stand in the way of good; gold-standard data quality might not be available due to basic, boring practicalities. Even the silver-standard may be out of reach, depending on your budget.

So what do you do when faced with these decisions?

This question strikes me as similar to the transfer market, where clubs outside the very richest of the rich are usually faced with trade-offs about who to bring in. Players will either have weak spots or question marks over their performance, adaptability, or consistency. It's up to somebody at the club to make the call about which areas are the most important, which pros or cons take priority over which.

Nowadays, it should be a little easier for clubs to make informed choices about decisions around data providers. With large language models, you could -- at least in theory -- create a document about how your coaching staff thinks about things like duels, ball progression, pressing, and then compare that to a provider's stat and event data documentation. How much alignment is there? How meaningful do the provider's definitions seem? How much ambiguity is there between the provider's conceptual boundaries? Is there anything that the provider simply doesn't cover?

Correctness is a little less easy to gauge, but you should at least be able to limit the task by only assessing the stuff that makes it through the above first set of filters. For most providers (hopefully all), shots would be something that'd make it through to this second round. At this point, it's likely that you'd want to look at the things that most affect expected goals models: are the shot locations right, are the player designations right, if the data has any major situational markers (phase of play, pressure, sight at goal) are they correct?

There's a slow, ebbing push towards a generalised framework for the structure of football data (see Common Data Format), but perhaps there should also be one for the semantics as well. I've generally been resistant to the idea, as there can be such a large amount of difference between data providers in certain areas. (Related Get Goalside: 'What is the worldview of your data?'). Maybe that way of seeing the task is getting it backwards though: maybe you should start with the football framework, and if the provider's event data doesn't fit into it then provider's event data be damned. If it doesn't fit, it's either not meaningful or is ambiguous, and if their data leaves gaps in your football framework, then that just means the data's not 'complete'.

This sounds dangerously like it might involve data engineers. Recent trends in football data job postings have reflected what a bunch of people, Get Goalside included, have said for a while about the importance of that domain. The coding abilities of LLMs have only amplified this: software is easier to create, but it's easiest when the building blocks are already in place. If you squint at this a bit and think a bit more abstractly, you could say that this is a case of LLMs being a powerful tool when the information environment is right.

You see this in non-coding cases too: your Notion or your Google Workspace or (god forbid) your Meta accounts all open up 'second brain' opportunities only if all of your relevant info is already being stored in them (or places accessible to them). The way that you understand football and record that understanding, as an individual or a footballing organisation, is an information environment. How suitable is that information environment for the tasks and the tools at hand?

Data providers want to fight over the definition of line-breaking passes. In the marketplace of data quality ideas, it's actually your understanding of the game that they have to compete with the most. 'What is good data quality' is a fairly well-understood question, but what 'good data quality' is depends on your understanding of the world that the data is reflecting.

MORE ABOUT ME