References and citations – a reminder
In a research database, citations are incoming and references are outgoing. The definition is not a private opinion: it is the official position. An article has on average 29 references and on average 11 citations. If it were a closed system these numbers would have to match, but of course there are many references that point to targets outside the corpus.
There is not really any use case of database users clicking on references to jump to a target. The high-value use cases are about citations.
Measuring the capability, not the data product
So what is the right data product to measure the quality of citation data? Consider two options:
Option A: The set of all pairs (reference-occurrence ID, target work-entity ID) for the 2.5 billion reference occurrences in the database, constantly updated. This measures the quality of the matching capability.
Option B: For each work entity, the set of all reference occurrences in the database that point to it. This measures the quality of the data product a user actually cares about: does this article have all the citations it should?
These are very different. Option B's quality has only a limited relationship to the accuracy of reference matching. If a reference is missing because it was not captured – either due to a capturing error or a policy decision – then the set of citations is incomplete regardless of how accurate the matching is. Even worse, if a whole article is missing that should have cited the target, the citation set is incomplete. By looking only at the precision and recall of the matching capability, we are fooling ourselves.
The missing affiliations
From a data tool one of my colleagues built, I learned something that had escaped me for years: fourteen million article records in a major database do not have affiliations. That is a huge fraction of 88 million total articles. Six-digit counts for the number of articles attributed to specific universities radiate absolute certainty – yet there is no error bar, as any responsible scientific article would have. What would that count be if we had affiliations for those 14 million articles?
When I mentioned this to colleagues, the main reaction was that adding those affiliations would be either impossible or too expensive, and would probably affect only old articles anyway.
But the point is not whether we go back and add them. The point is that we know. Before we even begin to think about cost, it is our job to understand the nature of those 14 million articles – so that we can take an informed decision. It can still be a valid business decision not to go back and add anything, but then we need to take some other action. We cannot knowingly let customers use incorrect data.
We also know that 14 million articles without affiliations is visible only if you follow the ground rule respect zero – the rule that says there is a meaningful difference between "no affiliation recorded" and "no affiliation exists". If we had followed that rule consistently, this problem would have been visible immediately.
What data products make obvious
All this becomes obvious the moment we embrace the concept of data products and their quality – and remains obscured when we think in terms of capabilities like normalisation, matching, and linking, each with their own individual quality scores.
In the domain-based operating model, the whole team – product, operations, and technology together – is responsible for the data product, not just the capability. Therefore the whole team is accountable for any gap in that data product, whatever its cause. They cannot afford not to worry about missing citations, regardless of why they are missing.
Data products is not just words. It symbolises precisely the paradigm shift required – in the operating model, in the architecture, and above all in the mental model of what we are producing and why.