Exactly twenty years ago, we published the Tag by Tag – the Elsevier XML documentation for scientific article content – in print. There had been earlier digital versions, but we started with a real printed book of more than 400 pages, which content people at Elsevier, at content suppliers around the world, and at consuming systems had on their desks.
The book came in British Racing Green. It was, in my opinion, a benchmark of technical documentation quality – and anyone who has not looked at it should take a moment to have a quick look (open the PDF in Acrobat). The fact that it was a physical book, shared across every organisation in the chain, meant complete consolidation of interpretation. One interpretation. No ambiguity.
Typeset model standardisation
At the same time, twenty years ago, a new wave of journal typeset-model standardisation was in full swing. New size-reducing fonts and new page layouts were being introduced. A massive poster showing all the typeset models in actual size was used to convince publishers to adopt the standard. They came to see what the styles looked like before agreeing to a transition. It was normal practice to make such investments to inform vendors and seek buy-in from stakeholders.
Who would have thought that twenty years later, the content standards and the typeset models would still be in use?
Where we are now
We do have a successor content standard lined up: CP/LD, the Content Profile / Linked Data standard, now accepted as a NISO standard. It separates the HTML5 presentation layer from the linked-data annotation layer. However, many of the benefits of CP/LD can also be obtained today, especially with all the innovations that came since the XML DTDs were written. There is not much that we want to do with primary content standards that we genuinely cannot do today.
The deficiencies are real, but they are mostly in the secondary content – the citation database schemas – which never underwent a comparable generational upgrade.
The boldest next step
The most interesting possibility that HTML5 at the basis of CP/LD opens up: a world where PDF no longer plays a role.
PDF was designed to guarantee a precise print rendering. HTML5, with all its modern capabilities, can do this too – and more. It can reflow for any screen size, support accessibility features, be searched and indexed with full fidelity, and carry interactive content without any of the constraints of a print format.
If we were to fully commit to this direction, the constraints imposed by PDF – the need for precise page layout, the fixed column widths, the page-break logic – would fall away. Content could be structured for reading, for processing, for AI, without the contortions that print compatibility currently requires.
This is not a proposal for next quarter. It is a question worth sitting with: given that HTML5 can now practically do everything PDF can do, what exactly are we still optimising for?