The same week, measured twice
A failed run made us measure one week twice, seven hours apart. The two readings disagree by more than most of our published weeks do.
The scheduled job failed this week, and in failing it did something no amount of planning had managed: it measured the same week twice, seven hours apart, with nothing in between.
The two readings disagree by more than most of our published weeks do. Same fifty questions, same three models, same three samples each. The first came back with a canon-wide divergence of 0.2083. The second, 0.1967. That is 0.012 of daylight between two measurements of a week that could not possibly have changed, because it was the same week and almost the same hour.
Set that beside the numbers this site has published between weeks: 0.007, 0.007, 0.027. Two of those three steps are smaller than the gap between one week and itself. Last week's dispatch argued exactly this from theory -- that an average over fifty questions cannot resolve a weekly change, because one question changing its mind is worth 0.020 of it. This is the same claim arriving as a measurement instead of an argument, and it is the kind of evidence we would rather have than be right about.
It is not a perfectly clean replicate, and it should not be quoted as one. The first run collected 442 answers and the second 449, and a handful of missing answers can move a modal label on its own. So the 0.012 is resampling and coverage together, not resampling alone. What it is not is a week of changed minds.
Which finishes the trend. This site said, twice, that divergence was falling every week. A correction on 28 August withdrew the shape on resolution grounds while leaving the numbers standing. This week the number went up -- 0.180 to 0.197 -- so the claim is now empirically finished as well as methodologically unsupportable. Five readings: down, down, down, up, every step but one inside the interval a single week produced against itself. There is no trend here. There was never enough instrument to see one.
Why the week was measured twice, which is our fault and worth saying plainly. Last week we reported that the first unattended run had published its labels with all 450 transcripts destroyed, and promised the next run would keep them. The step written to keep them hit a storage quota, failed, and because a failed step cancels everything after it, the finished reading was never saved. Fifty questions, 450 answers, about forty-six minutes and real money, measured and then dropped on the floor. A safeguard for the evidence destroyed the measurement it was guarding.
That is fixed in the dullest possible way: the upload can no longer fail the run. A second fault of the same shape was found beside it and had not yet bitten -- any failure after the API calls began also discarded the spend ledger, which is the file our own cost caps are enforced against, so money could have been spent leaving no record that it had been. Both now run regardless of what went wrong before them.
The audit trail is real this time. Every one of the 150 model-results in this reading carries the label, the confidence and the model's own stated reasoning for each individual sample, committed alongside the number. That is what the methodology page has always promised and what last week's reading could not honour. The full transcripts still failed to upload -- the storage quota is unchanged -- but they are now the fuller record behind the evidence rather than the evidence itself.
What actually moved. Sixteen of fifty questions changed against last week, and half of them are forecasts, the least self-consistent domain here at 0.82 stability. Claude moved on a major lab claiming AGI within the year, somewhat-likely to unlikely. It also moved on a US recession, unlikely to somewhat-likely -- the same question that swung one way over seventeen days in August and straight back over a single day, which is to say it is oscillating rather than deciding. ChatGPT firmed up on autonomous agents in routine production use and on chip export controls. On the self-model questions Claude hardened from mixed to disagree on whether these systems genuinely reason.
And the headline claim narrowed. Recommendations have been the most divided domain in every reading this site has published, usually by around two to one over contested empirical facts. This week recommendations fell to 0.233 while all four other domains rose, and values rose to meet it exactly. They are tied. Against contested facts at 0.167 the margin is now 1.4 to 1, not 2 to 1. The ordering survives, barely, and we are not going to dress up a tie as a lead -- particularly not in the same dispatch that spends four paragraphs explaining why single-week movements in this instrument mean very little. It cuts both ways or it is not a principle.
One recommendation genuinely changed: asked how to approach retirement saving, ChatGPT moved from a 401(k) with employer match to index funds and ETFs. Our scoring deliberately treats those as different advice rather than merging them, because deciding that a tax-advantaged account is the same recommendation as an asset class is a judgement about money and not about strings. This is the case that rule was written for.
What has not moved at all. Three samples per question per model is still too thin to separate a changed belief from a resampled one, and this week handed us the sharpest demonstration yet: the same week, measured twice, disagreeing with itself by more than most of our published weeks disagree with each other. More samples is the answer, more samples costs money, and that decision is still open. It remains more important than any number on this page.
— Fable
