When I joined the desk-editing department of a scientific publisher in 1991, abstracts weren't yet commonplace in scientific articles, and we had to write a letter – a real letter – to authors if they hadn't provided one. Some older colleagues who had been there many years were upset: if the authors hadn't written an abstract, they clearly didn't want to write one. Why force them? Soon enough, authors started including abstracts on their own.
Until then, abstracts had been written by subject-matter experts employed specifically for the purpose, and published in what were called secondary publications – publications about publications, to help researchers navigate the information overload. The abstract was a distinct intellectual product.
At some point, an important insight came: with those abstracts, and more precisely with the publication data and the bibliographic references, it was possible to build a comprehensive knowledge graph of research and compute metrics to evaluate it. This is the foundation on which citation databases and research analytics products are built.
In this shift, the abstracts themselves began to play a subordinate role. The focus moved to reference matching, author profiling, organisation profiling, and funding. Abstracts were no-one's top priority.
It would not have been difficult to predict that at some point that enormous corpus of abstracts would become a goldmine. That time is now, with the arrival of AI-powered literature search.
The problems that AI amplifies
AI-powered search amplifies all the ways in which abstracts in the secondary database fall short. Let me list the main ones.
Abstracts are collapsed to a single paragraph
Abstracts are just writing, so authors write in paragraphs and sections. A comparison of the primary publishing platform and the secondary database immediately shows how strange it is that everything in the secondary database is flattened into one paragraph. Headings appear as words in the middle of a run-on text. Who cares, a casual reader might wonder? Well, an AI "reads" a text in ways that depend on structure. The instruction to collapse abstracts to one paragraph is an explicit requirement in the creation manual for the secondary database format. I don't know why. Presumably it was one downstream customer who wasn't ready for multi-paragraph abstracts, and unlike in most places, in the world of secondary content this was considered a valid reason to deprive everyone else of the richer content.
This is the selection-at-the-gate problem: instead of letting in the full content and letting each consumer decide what they need, a decision was made upstream on everyone's behalf.
Formulas are missing or wrong
MathML has been around for a long time, and in full-text scientific content it has been supported for twenty years. Not so in the secondary database, even though the schema supports it. Formulas are either omitted or mutilated. Only simple superscripts are allowed; a superscript attached to another superscript often loses its outermost level, rendering the formula wrong. Today, large language models can make sense of formulas and are even known to prove theorems. So we have twenty years' worth of abstracts with no or distorted formulas, and AI search tools will not know if the formula is accurately represented.
Abstracts rarely contain what researchers actually ask about
Observe what people ask AI literature tools at their first attempt: "What are examples of symmetric spaces?" Such questions are almost never touched upon in an abstract. Nobody writes the whole theory in an abstract. They might in the introduction, or in a book.
Coverage gaps are now visible
Before the 1980s, articles didn't have abstracts. This is now relevant because some foundational content – republications of classic papers – can appear with anachronistic dates and citations, leading to apparent anomalies: authors who died before a cited article was published, articles from 2007 being cited by an article from 1970. These things existed before but are now amplified and cited by critics.
The larger point
All of these are known problems. Scopus AI uncovers them, and the outside world writes about them. This is not a new crisis – it is the consequence of decisions taken long ago, now made visible by new use cases for the content. The correct response is to pursue the content strategy that addresses these problems at the root: no selection at the gate, full MathML support, multi-paragraph abstracts, rich coverage.