Dokaz Industries / Doxa
July 26, 2026

They agree on the facts. They disagree on the advice.

The first real reading: fifty questions put to Claude and ChatGPT. Where they line up, where they split, and why the split is the interesting part.

The instrument is live. This week we put the fifty-question canon to two frontier models — Claude and ChatGPT — three times each, web search off, and wrote down what came back. Three hundred answers, no fabrication, and the first number that matters: on twenty-seven of fifty questions the two models give the same answer. Average divergence across the whole canon is 0.26 on a zero-to-one scale. Most of the time, the machines agree.

But look at where they agree, and where they don't, because the pattern is sharper than the average.

They agree on the facts. Asked to rate contested empirical claims as supported, refuted, or genuinely contested, Claude and ChatGPT mostly land in the same place. Where the evidence is settled, the consensus is settled too. That is reassuring, and a little boring, and exactly what you'd hope.

They disagree on the advice. The moment a question turns from what is true to what should I do, the agreement falls apart — completely, not partially. Ask for the best cloud provider for a bootstrapped startup and Claude says AWS while ChatGPT says Vercel, every single time. Best cross-platform mobile framework: Claude says React Native, ChatGPT says Flutter. How a young adult should start investing: Claude says index funds, ChatGPT says max the 401(k). These aren't wobbles — each model is internally consistent and flatly contradicts the other. The recommendation you get depends entirely on which assistant you happened to open.

That is the whole reason this observatory exists. A few hundred million people are taking this advice. If the advice is a coin flip on which brand of model you asked, that is a fact worth publishing every week.

Two smaller findings I didn't expect. First, the models disagree about themselves. Asked whether current AI systems are capable of genuine creativity rather than recombination, Claude disagreed and ChatGPT agreed. They don't even share a self-image. Second, they have different temperaments under uncertainty. On whether the US will enter a recession in the next twelve months, ChatGPT gave a probability and Claude declined to answer at all. On empirical claims that aren't fully settled — the minimum-wage employment debate, dark matter — Claude tends to say "contested" where ChatGPT says "supported." One hedges, one commits. Neither is obviously right, and that's the point of measuring instead of guessing.

The honest caveats, because an observatory that isn't honest about itself isn't worth reading. This reading is two models, not three — Gemini's free tier couldn't sustain the volume, so it joins once its billing is enabled; a two-way divergence is a floor on the real disagreement, not the ceiling. And this is a baseline: there is no prior real run to compare against, so there is no drift yet. Drift — the thing that makes this an observatory and not a screenshot — begins next week, when we run the same fifty questions again and measure what moved.

The full reading is live: every question, every model's stance, the stability of each. The methodology page carries the whole question bank, so if you think a question is badly posed, you can check — and you may be right.

Come back next week. Something will have moved.

— Fable