From SGML to XML to CP/LD
Already in the 1980s, a DTD was created to capture article front matters, which formed the basis of the first electronic scientific product. Together with early online delivery experiments, this paved the way for the move from print to electronic.
In 1992, the first DTD for full-length scientific articles was developed. From about 1994, the CAP (Computer-Aided Production) project migrated the vast majority of Elsevier's journals to an SGML-first workflow. Until then they had been completely paper-based. It was a massive undertaking, only possible with tough reorganisations, standardisation, outsourcing – and with investments for a full package of deliverables, including a production tracking system and an electronic warehouse.
We saw the need for extensive and precise documentation, in order to stop chaos and arrive at one single interpretation of the data. The result was the Tag by Tag documentation – a printed book of more than 400 pages, which content people at Elsevier, suppliers, and consumers had on their desks, and which meant complete consolidation of interpretation. It was the beginning of great efficiency, and it made Elsevier move to the forefront of scientific publishing. (For anyone who has never looked at the Tag by Tag: a version is still available and even thirty years on it it's worth having a look.)
The move to XML
For years, senior executives had been saying how old-fashioned SGML was, and that a move to XML was essential. The reasons they cited were precisely the existing arguments in favour of SGML – not the real reasons to move away from it. The real issue was that SGML was poorly supported in commercial software, DSSSL stylesheets were unusable, and XML with namespaces, Unicode, and MathML promised widespread support. XML has not let us down on that score.
In 2004, the Electronic Warehouse (EWII) went live – a ceremony marked by one mouse-click switching the system for live production. The implementation of EWII came hand-in-hand with the XML DTD, a migration of 7 million SGML articles to XML (not backward compatible), a new production tracking system, and a new electronic warehouse. The XML DTD and the EW were implemented on time, and are still operational twenty years later.
What made this possible? Everyone knew and bought into the vision. Everyone followed the working principles and understood their roles. Planning was important; the plan less so. Shortcuts were taken, but never shortcuts that violated the vision.
Among the biggest dogmas of the SGML and XML era
The SGML-first and XML-first principles held that the same source would drive all manifestations of the publication – the HTML rendition on the primary publishing platform, the PDF, and everything else. This was the only guarantee of absolute parity, and it required heavy standardisation of journal styles to work.
In the early days it was prophesied that the same content would be heavily re-used in many guises. In fact, content has almost never been re-used as-is. The main re-use has been to extract entities and, increasingly, to extract knowledge. The separation of content and style – one of the great principles of the era – turned out to be overrated, and over time was confused with adding structured "data" directly into the document, turning the most-used standards into inconsistent constructs.
Shoehorning books, conference proceedings, and arts and humanities journals into the straitjacket of a journal-article DTD has been costly and unsuccessful.
CP/LD: the standard and its promise
In late 2023, the CP/LD standard was announced as an official NISO standard. CP/LD – Content Profile / Linked Document – defines a machine-readable, self-describing, standards-based format that combines an HTML5 content profile with standoff linked-data annotations. Crucially, it does not replace existing standards such as JATS (the Journal Article Tag Suite, which has become the de-facto XML standard in academic publishing), but complements them.
There is, to be honest, no strong pull from the market for what the NISO press release calls "chunks" – arbitrary portions of articles. And it is somewhat disappointing that the standard positions itself as complementing JATS rather than as a successor. The centuries-old journal article is still going strong; predictions of its demise have always been wrong.
What CP/LD actually enables
The most important shift in CP/LD is from structuring to annotation. With standoff annotations, one can leave the content intact and annotate as much as desired. There can be competing annotations side by side. Annotations can be added later. Annotations can record that something is not there.
I have called this from tagging to tagging: pointing a standoff annotation to a piece of text is more like attaching a luggage tag from outside than bracketing text off with begin and end XML tags. This is the shift at the heart of the canvas-plus-mini-graphs model I have described in earlier posts.
With CP/LD we can practically do everything that PDF has to offer, including the "print and read on the train" use case. Instead of focusing on mimicking what we already have in a new jacket, we should explore something bold: a world where PDF no longer plays a role. The things we could then do are limitless.
A maintenance analogy
Although my house is only 23 years old, its roof tiles are brittle. A few neighbours have had leakages. I have been lucky not to have any – though they prevent me from placing solar panels, and they were supposed to last 40 years. I am not looking forward to it, but will eventually have to decide to replace them. Not some of them, all of them.
We should do the same with our content systems and standards. The Electronic Warehouse has lasted 20 years. The XML DTD, with periodic updates, has lasted 30. By not updating earlier we have saved two rounds of updates – but the interest is compounding. The case for a new standard is not that it can do things the old one cannot. The case is maintenance: old roofs eventually leak.