Books are always on the minds of researchers in the humanities and social sciences – the disciplines where the book, not the journal article, is the primary publication vehicle. They are proud of their books, and they are concerned that coverage in major citation databases remains incomplete.
There have been gaps in book coverage and work is underway to fill them. But filling gaps does not automatically produce a useful data product. The question that should be asked is: will this help researchers?
Why book coverage disappoints
The answer, for many book authors and editors, is that their books appear to attract fewer citations than they deserve, and that their h-indexes and publication counts do not improve as much as expected. The effect of adding their works to the database is smaller than anticipated.
There are structural reasons for this. Even if the capturing rules are correctly followed, we are shoehorning book data into a schema designed for journal articles. This creates persistent problems:
- Edited books and monographs are not carefully distinguished. Editors of collected volumes do not get appropriate credit in author profiles. Monograph authors are treated differently from article authors in ways that disadvantage them.
- Book series volumes are not cleanly distinguished. The relationship between a chapter, the volume it appears in, and the series the volume belongs to is handled inconsistently.
- Citation accumulation is fragmented. Because of the way series, editions, and related works are modelled, citations do not aggregate in the way users expect.
- Chapters rarely have abstracts. In a world moving towards AI-powered answers, chapters without abstracts are increasingly invisible. There is no reason not to generate useful summaries.
What a richer source model would look like
If the Sources entity were more like the Organisation entity – hierarchical, with complex parent-child relationships, a rich curated database to match against, and with the option to associate articles and chapters with multiple source entities – then suppliers could capture "what's on the tin" without having to force the content into a restrictive structure.
Consider this: in some databases, searching for other chapters in the same book returns only a subset of them; or in the case of book series, returns all chapters in the entire series across all volumes. This is because the "source" of a chapter is modelled as the series, not the volume. A hierarchical source model would resolve this.
The deeper question
Even if all of these modelling and capture issues were solved, a deeper question remains: are we measuring books in a way that reflects their true impact?
Citation counts and h-indexes were developed for journal articles. The citation practices of humanists and social scientists are different: they cite books, and they cite older material. A researcher with a strong books-based record will appear disadvantaged when compared to a colleague in an experimental science who publishes many short journal articles.
Without a comprehensive vision for how books should be represented and measured, filling the gaps will not do justice to books – or to the researchers and editors whose work is primarily in book form. Before investing in coverage, the right question is: what kind of data product are we building? Until that question is answered, gaps filled are gaps merely transferred from one column to another.