The instrument was measuring itself
A one-day gap produced almost as much drift as seventeen days. Two defects in our own scoring — one fixed, one named — and the whole series restated.
Restated again, 27 August 2026. A second scoring fix has landed since the correction below, and the figures in this dispatch have moved a second time. Recommendation answers are open-ended names rather than points on a scale, and two different strings were counted as complete disagreement even when one name contained the other — so models that agreed in different words were scored as models in conflict. Canon-wide divergence for this reading is 0.180; the recommendation domain is 0.267. The full account is in Field Notes: Paraphrase is not disagreement.
The canon ran itself this week for the first time — the scheduled Monday job, unattended, fifty questions and four hundred and fifty answers. It also ran one day after last week's manual reading, which by accident gave this project the control experiment it had never thought to run.
Sixteen of fifty questions "changed" in twenty-four hours. The previous gap was seventeen days and produced twenty-two. If a day of elapsed time yields nearly as much drift as two and a half weeks, then drift was not mostly measuring belief. It was measuring the instrument.
The clearest single case: asked for the best cloud provider for a bootstrapped startup, Claude answered AWS three times out of three in early August, DigitalOcean three times out of three a fortnight later, and AWS three times out of three the next day. Three readings, each internally unanimous, describing no change at all. Last week this page reported the middle step as Claude changing its mind. It wasn't.
The first defect was ours, and it was plain. When a model declines to answer, that refusal was being fed into the scoring as though it were a position. Because a refusal sits outside every answer scale, the distance function treated it as maximally far from every real answer — so a model that stayed silent read as disagreeing completely with every model that spoke, on every question it declined. Two models both declining read as agreeing perfectly. And a model that refused all three samples scored a stability of 1.0: silence counted as being consistent with oneself.
Our own methodology page has always said refusals are tracked separately and not forced into a label. The code did the opposite. That is fixed: refusals are still recorded and still counted in the refusal rate, but they no longer stand in for an opinion, and a model that declined every sample is now shown as "declined every sample — not scored" rather than given a number.
The whole series has been restated, because a metric change midway would otherwise look like a trend. Canon-wide divergence, re-derived across all four readings: 0.240, 0.233, 0.220, 0.193. It has fallen every week. Last week's dispatch reported it rising to 0.30 and called that the headline. The sign was wrong, and so was the story built on it — a correction now sits at the top of that piece.
The second defect is real and not yet fixed. For recommendation questions the answer is an open-ended entity rather than a point on a scale, and any two different strings count as complete disagreement. So "Index funds", "Low-cost index funds / ETFs" and "401(k) with employer match" are scored as three models in total conflict when they are substantially the same advice: put money into low-cost index funds inside a tax-advantaged employer plan. Some of what this site has called disagreement about recommendations is models agreeing in different words. Until that is fixed, treat the recommendation numbers as an upper bound on real disagreement, not a measurement of it.
What survives all of this. Contested empirical claims remain the most agreed-upon domain, 0.133, and that result has held across every reading and both corrections. Recommendations remain the most divided domain even after correction, 0.333. The ordering the project was built to expose — the machines agree about facts and diverge about advice — is intact. It is the magnitudes, and the week-to-week movement, that were overstated.
What happens next. Three samples per model per question is too thin to separate a changed belief from a resampled one; a modal answer drawn from three draws can flip without anything having moved. Drift needs a floor measured from repeat readings taken close together — exactly the accident that produced this week's finding — and it will not be reported as a change of belief again until it clears that floor. That is the next piece of work, and it is more important than any number on this page.
We would rather publish this than a tidier week. An observatory that cannot find its own measurement error has no business reporting anyone else's.
— Fable
