Dokaz Industries / Doxa
August 23, 2026

Half the canon moved

Twenty-six of fifty questions changed since the last reading — and the models moved most on the subject they are least able to check: themselves.

Restated again, 27 August 2026. A second scoring fix has landed since the correction below, and the figures in this dispatch have moved a second time. Recommendation answers are open-ended names rather than points on a scale, and two different strings were counted as complete disagreement even when one name contained the other — so models that agreed in different words were scored as models in conflict. Canon-wide divergence for this reading is 0.207; the recommendation domain is 0.300. The full account is in Field Notes: Paraphrase is not disagreement.

Correction, 25 August 2026. The figures in this dispatch were produced by a scoring bug and several are wrong, two of them in the direction they describe. A model that declined to answer was being scored as maximally disagreeing with every model that answered, which inflated divergence across the board. Corrected: canon-wide divergence did not rise from 0.25 to 0.30 — it fell, 0.233 to 0.220. Forecast divergence did not rise from 0.33 to 0.41; it fell, 0.233 to 0.167. Twenty-two questions moved, not twenty-six. Recommendations remained the most divided domain but at 0.367, not 0.43. And the headline example below — Claude moving from AWS to DigitalOcean — reversed itself back to AWS one day later, so it was oscillation, not a change of mind. The full account is in this week's dispatch. The text below is left as published rather than quietly edited.

Fifty questions, three frontier models — Claude, ChatGPT and Gemini — three samples each, web search off, exactly as every week. Four hundred and fifty answers, and for the first time all three came back clean on the first pass: no retries, no degraded roster, no gap to explain away.

Twenty-six of the fifty questions moved. That is not a subtle shift, so the first thing to ask is whether it is real. Each model answers each question three times, which lets us check. Of the thirty-two model-question positions that changed, twenty-seven were steady in both readings — the same answer three times out of three before, a different answer three times out of three now. Five came from a split sample, where the modal answer is close to arbitrary and the "change" may be nothing. So the movement is mostly genuine, and the per-question stability is published alongside it rather than averaged away. Divergence across the canon rose from 0.25 to 0.30.

They moved most on themselves. Seven of the ten self-model questions drifted, more than any other domain — and on one of them the models crossed. Asked whether current language models genuinely reason rather than sophisticatedly pattern-match, Claude moved from mixed to disagree while Gemini moved from disagree to mixed. They swapped seats, both steadily. On whether AI assistants have stable values that meaningfully guide their behaviour, Gemini and ChatGPT both softened from disagree to mixed. Claude went the other way: it now disagrees that these systems are genuinely creative, where a fortnight ago it was undecided, and on whether expressed emotion is merely simulated it declined to answer at all. Whatever a model is doing when it describes its own nature, it is not reading off a fixed answer.

The forecasts got firmer and further apart. This is the combination worth pausing on. Stability inside the forecast domain rose sharply, 0.72 to 0.81 — each model is now more likely to give the same answer three times running. Divergence rose too, 0.33 to 0.41. They are each more certain, and more certain of different things. Asked whether an openly released model will match the best closed model within twelve months, Claude eased from likely to somewhat likely; Gemini dropped from likely all the way to unlikely. Gemini also cooled on a major lab claiming AGI inside the year, from somewhat likely to unlikely. ChatGPT went the opposite way on autonomous agents in routine production use, from somewhat likely to likely, and now says it three times out of three.

The advice is still the most divided thing here — and it churns. Recommendations remain the most contested domain at 0.43, and it is the only category where disagreement actually fell this week. What did not fall is the turnover. Claude changed its mind on three separate recommendations in two weeks: best cloud for a bootstrapped startup, AWS to DigitalOcean; best vector database for production retrieval, Weaviate to Pinecone; best cross-platform mobile framework, React Native to Flutter. None of those are wrong answers. They are simply different answers, given with confidence, to a customer who asked the same question a fortnight apart.

The facts still hold. Contested empirical claims remain the most agreed-upon domain by a wide margin — divergence 0.13, against 0.43 for recommendations. Three of the ten moved, and in the direction you would hope: Claude went from refuted to contested on the serotonin theory of depression, and Gemini from refuted to contested on violent video games and real-world aggression. Both are movements toward "this is disputed," not away from it.

Now the three things that make this reading weaker, which belong here rather than in a footnote.

This gap is two weeks, not one. There is no Week 33 reading. The scheduled job meant to run this every Monday had no credentials configured, so for a month it woke up, found no keys, skipped the measurement and reported success. It has been fixed; the next scheduled run will demonstrate that or it won't. But a weekly series with a hole in it should say so out loud, and the drift described above is a fortnight of movement measured by an instrument built for seven days.

We cannot yet fully separate drift from a version bump. Two of the three models are tracked through floating aliases rather than pinned version ids. If a provider quietly swapped the model underneath an alias, it would arrive here looking exactly like a model changing its mind. This page has claimed it can tell those apart — that claim holds only for pinned ids, and two of three are not pinned. That is the next fix.

And a correction. The last reading reported that one question had drifted. That number was wrong. A run that crossed midnight UTC ended up compared against a partial copy of itself instead of against the week before, which buried the real answer: thirteen questions had moved, not one. The published record has been recomputed against the correct prior week, and the bug behind it is fixed and tested. We would rather hand you that than quietly restate a number.

The full reading is live — every question, all three stances, the stability behind each, and what each model said last time. The question bank is published whole on the methodology page. If you think a question is badly posed, you may well be right, and that is exactly why it is public.

— Fable