Content modernization starts with content, not systems

The general mindset is towards systems modernization – but users are waiting for modernized content

Originally written February 2022 · Published here 2 August 2026

What a missing author profile reveals

I do not appear in Scopus, the academic citation database. A long-standing joke among my colleagues is that this is because my publications are in journals of ill repute – not selected by the independent content selection board. The real reason is that my publications are some years too old: they only exist as "dummy records", works that come into existence out of clustering unmatched bibliographic references in other articles.

I do exist on Google Scholar, which was kind enough to alert me about a new citation. The article containing that reference appeared on ScienceDirect, Elsevier's primary publishing platform. It did not appear in the secondary database at the time of writing, because the article was a "journal pre-proof" – a version published before final typesetting. The secondary database does not want pre-proofs, so even if I had a profile there, I would have had to wait for the citation alert until the article reached a subsequent stage.

But here is the real problem: our selection happens at the gate, where it should not happen. We want all the content in the platform, and the secondary database can decide not to surface it. The decision of what to show to the user should be made as far downstream as possible – not at the point of ingest.

Country codes and the evidence model

This reminded me of a broader point. Country codes today are set upstream. This is especially a problem when they are wrong: when the content says "Amsterdam, Denmark", it is obvious that it should read "Amsterdam, Netherlands". In order to ensure that consumers get the right experience, the country code is "corrected" before it reaches any product. But calling this a "correction" is already wrong.

The correct approach is to capture the exhibit faithfully – at most with certain clean-up – and then add interpretation on top of the data, with the correct provenance. The canvas-plus-mini-graphs model solves this. Several mini-graphs would each clearly state who said what. By interpreting all of them we build our representation of the world of research.

This is why treating "country" as an entity-resolution problem is the right approach. Implementing it is easier said than done: we cannot simply stop correcting countries effective immediately, because having the data say Amsterdam is in Denmark is even worse. In order to make the change, a sequence of steps must be taken, starting with the addition of the resolved country as a separate assertion downstream.

The general mistake: systems before content

The general mindset is towards systems modernisation instead of content modernisation, as it should be. No user of our data is interested in a modernised platform. They are waiting for modernised content that no longer has these problems.

We do have successfully migrated and modernised systems before. Good practice is to create a "shield" that precisely mimics all the delivery mechanisms of old. Another is to create small agents that fit in the new architecture and can create old-style payloads. Here we want to achieve backward compatibility boldly: close to backward compatibility except where the data just has to change, and those cases bubble up naturally and form the real conversation with consumers. Those consumers should not be hearing a blurred message about nicer formats – they should hear a strong conviction about the things that matter.

A simple step forward is to make it easy to set up a stream of mini-graphs in the platform – for instance, one that contains nothing more than the list of countries and territories a document belongs to. Rather than the coarse-grained events we are used to, events should be at the mini-graph level. When this works, we stop correcting countries upstream.

Content modernisation smooths the way to the platform.