Dokaz Industries / Doxa
August 27, 2026

Paraphrase is not disagreement

Two models named the same framework and were scored as maximally opposed. Fixing that meant choosing rules with nothing in them to tune — and watching whether our best headline survived.

Correction, 28 August 2026. This dispatch says the corrected series keeps its shape — “falling every week.” That claim does not hold, and the reason is the same kind of error it was written to report. Canon-wide divergence is an average over fifty questions, and one question changing its mind moves that average by as much as 0.020. The first two steps in the series are 0.007 and 0.007 — each smaller than a single question. Between the first two readings, twenty-six of the fifty questions changed while the canon-wide figure moved by less than any one of them was worth: large movements in opposing directions, nearly cancelling, reported as a gentle decline.

The first reading is not on the same axis either. It used two models rather than three, and divergence is a mean over pairs: a third model that lands on either existing answer splits the same disagreement across three pairs instead of one, scoring 0.67 where two models scored 1.00. Measured on the weeks where both can be computed, dropping to two models raises the canon-wide figure by 0.013 to 0.027 — larger than every step in the series. In those weeks the third model sided with one of the other two on about four of every five questions where they split, so the dilution is structural rather than incidental.

What the four readings actually support: canon-wide divergence has been flat within this instrument's resolution, with one fall — the most recent, 0.207 to 0.180 — that is bigger than a single question can account for. The individual numbers below stand; the shape drawn through them does not. The per-domain and per-question figures are unaffected, and the ordering that recommendations diverge most and contested facts least is not close to this margin. Both limits are now stated on the methodology page. The text below is left as published rather than quietly edited.

Two weeks ago this observatory reported that models disagree most about recommendations. That is still what it reports. But last week's dispatch admitted the recommendation numbers could not be trusted as measurements, only as an upper bound, and named the reason: for those questions the answer is an open-ended name rather than a point on a scale, and any two different strings were counted as total disagreement.

Here is what that looked like. Asked which web framework to start a new project with, in the very first reading, Claude answered "Next.js" and ChatGPT answered "React + Next.js". The instrument scored that pair 1.0 — maximal disagreement, the highest number it can produce, for two models naming the same framework.

The fix is two rules and no judgement calls. Case and punctuation are ignored inside a name, so "Index Funds" and "index funds" are one answer. And if one name's words wholly contain the other's, they are one answer: "Next.js" is what "React + Next.js + TypeScript" recommends. That is all. Partial overlap is deliberately not agreement — "Index funds" and "Target-date funds" share the word "funds" and remain fully distinct, because a similarity score would have called them half-agreed and half-agreement is not a thing two pieces of advice can be.

There is no list of known synonyms behind this, and no threshold. Both were available and both were rejected for the same reason: they would have let us decide, after seeing the answers, which models agreed. The rules are published on the methodology page along with everything else.

The fix had a bug in it, and the page caught it. Once short and long names count as agreeing, a model that says "React + Next.js + TypeScript" twice and "React" once has two stances equally consistent with its samples, and the first version of this change broke that tie alphabetically. It picked "React" — publishing ChatGPT as maximally opposed to two models that had both said Next.js, which is the precise error the change existed to remove. It was visible on the rendered page and invisible in a green test suite. The tie now goes to whatever the model actually said most often, and that question's divergence for this week is 0: three models, one framework.

Restated across all four readings. Canon-wide divergence, re-derived: 0.220, 0.213, 0.207, 0.180. Last week this series was published as 0.240, 0.233, 0.220, 0.193. Every reading was overstated; the shape — falling every week — survives. The recommendation domain moved further: 0.40 to 0.30 in the first reading, 0.467 to 0.367 in the second, 0.367 to 0.300 in the third and 0.333 to 0.267 in the fourth. One drift that was published as a model changing its mind was a model rephrasing itself, and is gone.

What the correction did not do is rescue the headline. This change could only ever push the recommendation domain's divergence down, and that domain is the site's most eye-catching claim. It went down and the claim survived: recommendations are the most divided domain in this reading at 0.267, twice contested empirical facts at 0.133, and they have led or tied for the lead in all four readings. That ordering has now outlived three separate scoring corrections, each of which was an argument against it.

What is still not fixed. Last week's dispatch put "Index funds", "Low-cost index funds / ETFs" and "401(k) with employer match" in one bucket as substantially the same advice. This fix merges the first two and leaves the third alone, and that is on purpose. Deciding that a tax-advantaged account is the same recommendation as an asset class is a judgement about money, not about strings, and it is not one an extraction pass should be making quietly on our behalf. Nothing is stemmed either, so "fund" and "funds" are still two words. The recommendation figure remains an upper bound. It is a tighter one.

One more thing, found while checking this one. The methodology page has been promising since launch that every published label can be audited against the answer it came from. It could not. The full transcripts are not committed to the repository, and last week's reading was the first taken by the unattended weekly job — which ran on a machine that was destroyed when the job finished, taking all four hundred and fifty transcripts and every rationale with it. That reading's labels are published unaudited, its question pages now say so on their face, and the methodology page no longer claims otherwise. From next Monday the reasoning behind each individual sample is written into the reading itself, where it survives the machine that produced it. The earlier hand-run readings had their rationales recovered from local copies and now carry them.

The larger problem named last week has not moved at all: at three samples per question, this instrument still cannot separate a model changing its mind from a model rolling the dice differently. More samples is the answer and more samples costs money. That decision is still pending, and it is still more important than any number on this page.