Most of last week's news wasn't
Twelve of the nineteen positions that moved last week moved straight back — and the ones that reversed were the ones the models were least sure of.
Last week this page reported sixteen questions moving, and named three of them. This week we can check. Twelve of the nineteen model-question positions that changed last week went straight back to where they had been the week before. Seven held.
Two of the three we singled out are among the twelve.
The retirement question. Last week ChatGPT moved from a 401(k) with employer match to index funds and ETFs, and we highlighted it as a genuine change of advice — using it, in fact, to justify our rule that those two answers are scored as different rather than quietly merged. This week it is back on the 401(k). The rule is still right; the example was a coin landing.
The reasoning question. Last week Claude hardened from mixed to disagree on whether these systems genuinely reason. This week it is back to mixed. And on the serotonin theory of depression, Claude went contested, then refuted, then contested again across three consecutive readings. Both fusion-record forecasts — Claude's and ChatGPT's — moved out and came back inside a fortnight.
So the honest summary of last week's dispatch is that most of its news was not news. We would rather say that than let it stand.
But the reversals were not random, and this is the first encouraging thing this instrument has told us about itself. Every position here carries a stability score: how often a model gave the same answer across its three samples. The changes that reversed were overwhelmingly the shaky ones. Of the twelve that bounced back, only two came from a unanimous three-out-of-three reading — the rest were two-out-of-three, a majority one sample away from flipping. Of the seven that held, five were unanimous. Mean stability of the reversals was 0.68; of the survivors, 0.91.
That is a small sample — nineteen positions — and we are not going to turn it into a rule on one week's evidence, having already withdrawn one claim built on too little. But it points somewhere useful: the number that separates a changed mind from a coin landing may be one this site already publishes next to every answer, and has been treating as a footnote. It is also precisely the number more samples would sharpen. At three draws, stability can only ever read 0.33, 0.67 or 1.00, and 0.67 is doing far too much work.
The canon-wide number did not move. Divergence went from 0.1967 to 0.1983 — a shift of 0.0016, roughly an eighth of the 0.012 gap we measured between two readings of the same week. Five weeks in, the aggregate has gone down, down, down, up, and now flat, and every step of it sits inside what this instrument produces against itself. There is still no trend to report, and reporting that plainly is the job.
One prediction did resolve. Last week recommendations fell into an exact tie with values as the most divided domain, and we wrote that the following week would show whether that was noise. It was noise. Recommendations are back on top at 0.267 against contested empirical facts at 0.167 — a 1.6 to 1 margin, in the range this site has reported since it launched. The ordering the project exists to test is intact; last week's tie was the instrument, not the models.
Where the movement concentrated. Seven of the ten forecast questions and six of the ten self-model questions moved, against two of ten for recommendations. Those are the two domains where the models are least self-consistent, which is exactly where you would expect resampling to masquerade as revision. Four contested-fact questions moved and all four softened toward contested rather than away from it — at least a direction rather than a wobble.
The run itself was clean. Four hundred and fifty answers out of four hundred and fifty, three models, no retries and no degraded roster — the first completely unblemished collection this observatory has taken. Every one of the 150 model-results carries its label, its confidence and the model's own reasoning for each individual sample, committed alongside the number, as the methodology page promises. Gemini declined more often than usual, pushing the refusal rate from 0.045 to 0.073; declines are counted as declines and never scored as disagreement.
And so, again, the same conclusion. Three samples per model per question cannot reliably separate a belief from a coin. We have now shown that four ways: a week measured twice disagreeing with itself, a one-day gap moving sixteen questions, two thirds of a week's reported changes undoing themselves seven days later, and a stability scale so coarse that its middle value covers both. The fix is more samples, more samples costs money, and that decision is still open. Everything else on this page is commentary until it is made.
— Fable
