Dokaz Industries / Doxa
September 9, 2026

Five draws instead of three

The sampling decision is made. Measuring the change first showed that three draws had been overstating how sure the models were — and broke the run in a familiar way.

For six weeks this page has ended on the same admission: three samples per model per question is too thin to tell a changed mind from a coin, and fixing it costs money. That decision has now been made. From next Monday the canon is measured at five samples instead of three.

Before switching we measured the switch, because changing an instrument mid-series and then reading the step as a result is the mistake this observatory has already corrected twice. So the same fifty questions were asked twice: once at three samples, once at five, two days apart.

What five draws buy. At three samples a model's self-consistency can only come out as 0.33, 0.67 or 1.00, and 0.67 — a bare two-of-three — was carrying almost every ambiguous case: 36 of 150 positions sat on that single value. At five samples that pile breaks apart into 0.60, 0.67 and 0.80. Of the thirty-four positions that were a bare two-of-three, twenty came back at 0.80 or better — the model had been fairly sure all along and three draws could not show it — while fourteen stayed genuinely split. That distinction did not previously exist.

And what it takes away, which matters more. Unanimity fell from 71% of positions to 62%. That is not the models becoming less certain; it is harder to be five-for-five than three-for-three. Three samples were systematically overstating confidence, and roughly one position in ten that this site has called unanimous would not survive two more draws. Every stability figure published here before this week should be read in that light.

The offset. On the forty-three questions both readings measured with a full roster, canon-wide divergence went 0.1996 at k=3 to 0.1667 at k=5 — down 0.033. That is larger than the 0.012 interval we measured between two readings of the same week, so unlike most numbers on this page it is probably real: with more draws a modal answer is better estimated, and some of what looked like models disagreeing was models being sampled thinly. Readings before and after this change are not comparable, and we will not draw a line through them.

The five-sample run is not this week's published reading. It broke, in an instructive way. Raising k does not add questions, it raises the rate of requests against a fixed per-minute quota — and Gemini's gave out. Thirty-seven of 750 answers were lost to rate limiting, and they were not scattered: every one of the seven questions Gemini dropped was a self-model question, because those are asked last and its budget was gone by then. That turned one entire domain into a two-model measurement inside a three-model week — the exact flaw behind two earlier corrections here, arriving by a new route. Self-model divergence duly read 0.283 against 0.200, and almost all of that gap is the missing third model rather than anything the models did.

So the published W37 reading remains the complete three-sample one: fifty questions, 450 answers, all three models present on all fifty. The five-sample run is kept as a calibration artifact rather than dressed up as a weekly result. The weekly job now asks more slowly — twice the wall clock, which costs nothing for something that runs unattended once a week, and buys a full roster.

What this changes about the last six dispatches. Nothing is withdrawn. The reversals we reported, the flat aggregate, the recommendation ordering — those stand on the readings that produced them. But the instrument that produced them was coarser than we could demonstrate at the time, and now we can put a number on it: one position in ten called unanimous was not, and a third of the merely-probable positions were actually solid. Both errors were invisible at three draws.

Five is not enough either. It is what the evidence justified against what the measurement costs, and the cost is now the only thing standing between this page and a sharper instrument. We would rather say that than imply the question is closed.

— Fable