The reviewer error catalogue

The framework went through four external review rounds before its release candidate, and the clearest pattern in the record is an uncomfortable one. Reviewer confidence tracked accuracy inversely, the claims delivered with the most certainty were wrong the most often, and the corrections that mattered arrived hedged, sourced and quiet.

Out of that record came a standing rule, no reviewer claim adopted without independent primary-source verification, with every accept and reject logged with its reasoning, and the logs produced a catalogue.

One worked rejection

The rule is best shown on a specific, so here is one from the log. A reviewer reported that CycloneDX 1.7 had deprecated securedBy, the property family the framework’s data contract relies on. The claim arrived confident and detailed, and acting on it would have meant redesigning the input contract days before publication.

We opened bom-1.7.schema.json and checked the property directly. It is present and not deprecated, the claim was rejected, and the rejection is logged with the schema file as its source. Nothing about the reviewer’s certainty was evidence, and nothing about checking took longer than an hour. The cost asymmetry is the argument, because a wrong adoption ships an error into a normative document, and a verification costs an hour and a schema file.

The pattern has a research literature behind it. Philip Tetlock’s long study of expert political forecasting, published in 2005, found the relationship most people still find surprising. The experts most confident in their judgements were, on average, the least accurate, and the cautious, self-correcting ones outperformed them. A validation process that weights claims by the certainty of their delivery imports exactly that inversion into an assessment, and four rounds of our own logs reproduced the finding without our looking for it.

The catalogue

The ten errors live in the Universal’s Part D, in rough order of frequency, each with how it presents in a finished assessment and what it does to the result. The top of the list will be familiar to anyone who has read this blog. Resolution to product names at layer 4, which shows an institution more custody dependencies than it has firmware families and understates concentration at the most expensive layer. Tool-scoped coverage accepted as estate coverage, where everything a single discovery method cannot see is silently counted as absent.

Marketing language accepted as lineage disclosure, presenting as a suspiciously clean layer 2 with no disclosure programme behind it. And absence recorded as diversified, the inversion that turns a missing entropy record into good news.

The remaining six are quicker to state than to unlearn. Ambiguity resolved to the favourable reading, averaging across layers, averaging coverage across paths, the largest share omitted from a banded result, criticality flags applied by judgement rather than by the rule-based tests, and an institution-level figure reported without its mandatory companions.

Every one of the ten moves the result in the comfortable direction, which is why a catalogue beats intuition. A validator working from memory checks for the errors that annoy them personally. A validator working from the table checks for the ones that flatter the institution, and those are the ones that reach a board pack. Between them, the top four account for most of the material corrections across all four rounds.

After the fourth round

The ratified position is that there will be no fifth model-review round. The next external input into this framework is a pilot or a supervisor, because after four rounds the marginal reviewer opinion is worth less than one primary source or one reproduced computation. Multi-model review stays in place as a quality gate for new documents, and its accept-reject log stays part of the record, but validation of the method has moved to instruments that do not have confidence levels: the published vectors and reference implementation, the retrodictions against the public record, and a catalogue that tells any validator, in advance and in writing, where assessments go wrong.

Confidence is a feeling. The schema file is a fact, and the framework is built to prefer the second. That is the whole discipline, and it fits on one page of Part D.