Welcome to another Q&A, on this last Tuesday of the month.
Today’s question comes from Winona, and references a previous reflection I wrote. Winona asks:
In your post on the subjective side of indexing, you write that two indexers can produce different indexes for the same book, but that there should still be significant overlap because the text’s aboutness remains the same. Is there any practical way to quantify that overlap? For instance, after accounting for defensible differences in wording, formatting, and structure, should a certain percentage — say, 50% or more — of top-level entries overlap?
Thank you, Winona. That is an excellent question. If a certain amount of overlap is expected, it is reasonable to ask if that overlap can be quantified.
My short answer is no, I do not know if there is a practical way to quantify that overlap. I am also not sure if it is possible to ever fully quantify it. I imagine a lot would depend on the subject matter and type of book, as well as, of course, the subjective approach each indexer brings.
To attempt a quantitative answer, I think you would need to analyze multiple indexes from multiple indexers across several books. I have not done that research. If someone else would like to do so, I think the results would be fascinating.
I will, however, still attempt a response. While I do not have hard numbers, this is a question worth considering and I have some thoughts about what to look for in terms of overlap. I am mostly drawing upon my experience working with subcontractors. While subcontracting is not an exact comparison, since there is usually one indexer providing direction, subcontracting does involve two or more indexers working on the same text.
I also highly recommend this article by Jolanta Komornicka, recently published in The Indexer. Jolanta convened two panels, at two different conferences, for which participating indexers indexed the same book. Jolanta attempts to answer this very question, of how similar or different the resulting indexes are. Definitely worth checking out.
When considering overlap between indexes for the same book, I want to distinguish between main headings and everything else.
At the main heading level, I would expect a high level of overlap. Most terms are not up for interpretation. People, places, organizations, and objects should generally be indexed as written in the text. Concepts which are clearly identified in the text should also usually be indexed as presented.
There is still some room for variation and the indexers’ subjective choices. For example, whether or not to include glosses, and how to phrase the glosses; how to handle synonymous terms; what counts as a passing mention, and, conversely, which are important themes and discussions; and any discussions which are implicit and which require more interpretation. But by and large, I would expect main headings to mostly be similar across all indexes.
Beyond main headings, I think there is far more room for different decisions to be made. These include:
- Metatopic: How is the metatopic handled?
- Structure: How is information divided or gathered throughout the index?
- Subheadings: How many subheadings are used? Are subheadings mixed with undifferentiated locators? How are subheadings phrased?
- Cross-references: Are cross-references used sparingly or extensively? In which directions are cross-references pointing? Are cross-references used alongside or instead of double-posts?
At the level of structure and subheadings, there are so many different decisions that can be made. There should be recognizable overlap in terms of themes and correspondence to the text, but I don’t know if it is possible to quantify how much overlap or to expect that certain subheadings and arrays will be interchangeable.
To sum up, I would expect a high level of overlap between main headings, since many terms should be indexed as is. But indexing is a lot more than simply identifying and picking up terms. An excellent index involves organizing all of those terms in a way that is accessible and to highlight the key themes of the book. How all that organization happens is very much subjective. While the indexes should all still clearly point back to the same text, I don’t know if it is possible to fully quantify overlap.
Shifting directions slightly: I think what the creators of AI indexing tools are attempting to do is to quantify or codify indexing decisions. Or at least they believe that a good index is simply the result of following a certain set of rules. But that is not entirely true. Yes, an index is deeply governed by rules and conventions, and an AI tool can probably reproduce those rules and mimic various techniques. But an excellent index is ultimately the result of the indexer’s judgment. It is deciding which rules to apply, and when, and how to combine techniques, and how to make adjustments to fit the material rather than trying to force the material to fit the format. The subjective nature of writing an index, while difficult to quantify and difficult to compare, is also a superpower, enabling indexes and indexers to respond to the text and to readers.
PS. As I mentioned, I highly recommend Jolanta Komornicka’s recent article in which Jolanta attempts to quantify how similar and different the indexes by her panelists are. You can find it here. If you do not have a subscription to The Indexer, try accessing through your local library.