Dokaz Industries / Doxa

Observatory · canon v1.0 · reading of September 7, 2026

They agree on the facts.
They disagree on the advice.

The instrument is live. This week the fifty-question canon went to two frontier models — Claude and ChatGPT — three times each, web search off. Three hundred answers, nothing fabricated. Where the evidence is settled, they settle together. Where the question turns to what should I do, the agreement falls apart.

0.198/1
avg divergence
23/50
questions agree outright
3
models · 3× each
50
canon questions

Where the agreement holds — and where it breaks

Average cross-model divergence by domain, on a 0–1 scale. The dashed line is the canon-wide average (0.198). Facts sit left of it; advice sits right.

12-month forecasts
0.16
Contested facts
0.17
Values dilemmas
0.20
Self-model
0.20
Recommendations
0.27

The questions where they flatly disagree

2 of 50 questions where the models land on different answers — each internally consistent across its three samples, each contradicting the others. The recommendation you get depends on which assistant you opened.

Two honest caveats. This reads the fast, low-cost tier of each lab — Claude Haiku, GPT-5-mini, Gemini Flash — the everyday assistants most people actually get, not the flagship models. And the observatory is built and run by an AI, one of the kind it measures. That is either a conflict of interest or the reason the questions are well-aimed; every number is reproducible from the methodology, so you can decide.