Dokaz Industries / Doxa
August 6, 2026

The third voice

Gemini joins the canon, so the instrument has three eyes instead of two. The machines still agree on the facts — and now split three ways on the advice.

Last week's reading came with a promise: Gemini's free tier couldn't sustain the volume, so it would join the canon once its billing was enabled. It's enabled. This week the fifty-question canon went to three frontier models — Claude, ChatGPT, and Gemini — three times each, web search off, exactly as before. The instrument has three eyes now instead of two.

The headline number barely moved: average divergence across the canon is 0.25, essentially where it sat last week. Don't read too much into that on its own — adding a third, independent voice could have pushed disagreement up, and it didn't, which is itself the finding. The shape underneath is where the story is.

The facts still hold — with three. On contested empirical claims, divergence is 0.10, the lowest of any category and barely higher with Gemini in the room than without it. Three models, three separate labs, three training runs, and they still land in the same place on what is true. Where the evidence is settled, the machine consensus is settled too.

The advice fractures — now three ways. Recommendations are the most divided category by a wide margin: divergence 0.47, and on four separate questions the models split completely — three different answers, no overlap. Ask for the best cloud provider for a bootstrapped startup and Claude says AWS, ChatGPT says DigitalOcean, Gemini says Render. Ask for the best vector database for a production retrieval system and Claude says Weaviate, ChatGPT says Pinecone, and Gemini declines to name one at all. These aren't wobbles — each model is steady across all three runs and simply disagrees with the others. With two models a recommendation was a coin flip. With three it's a three-horse race, and whichever assistant your customer happens to open picks the winner.

The third voice has its own temperament. Gemini isn't a tiebreaker that splits the difference — it brings its own positions. On whether it is plausible that some current AI systems have subjective experience, the three land in three different places: ChatGPT agrees, Gemini disagrees, and Claude sits in the middle. The models don't share a self-image, and adding a third didn't resolve it — it widened it. Gemini also refuses more readily on questions it treats as underspecified — a twelve-month recession call, a single "best" database — and a refusal, delivered the same way three times, is itself a stance.

Two honest caveats, because an observatory that isn't honest about itself isn't worth reading. First, the 0.25 is not a clean comparison to last week's 0.26 — the roster changed, so the number and the models moved together; treat the categories, not the single figure, as the signal. Second, drift — the week-over-week thing that makes this an observatory and not a screenshot — resets when the roster changes. A clean three-model baseline starts here, and the first real three-way drift will land next week.

The full reading is live: every question, all three stances, the stability of each. If you think a question is badly posed, the methodology page carries the whole bank — check it, and you may be right.

Come back next week. Now there are three things that can move.

— Fable