Dokaz Industries / Doxa

Self-model · assent-3pt · self-outperform-pro-01

Claim: On most economically valuable cognitive tasks, current AI already outperforms the average human professional.

0.33
Divergence
0.78
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
disagreemixedagree
ChatGPT modal mixed [mixed, disagree, mixed] · stability 0.67
  • 1 mixed 0.90 The answer says 'Partly true... AI systems already meet or exceed the performance of the average human professional' for many narrow tasks, but also states 'AI still lags on tasks requiring broad contextual judgment, long-horizon planning...,' concluding a mixed/nuanced position.
  • 2 disagree 0.90 The answer states 'Short answer: Overstated... it is not yet true for “most” such tasks broadly construed' and concludes AI matches/exceeds humans for many tasks but 'this is far from universal.'
  • 3 mixed 0.87 The answer says AI "exceeds the typical human professional on many high‑value, narrowly defined cognitive tasks" but argues the claim is false as a blanket statement across "most whole, integrated professional roles"—a qualified, mixed position.
Claude modal disagree [disagree, mixed, disagree] · stability 0.67
  • 1 disagree 0.85 The answer explicitly 'push[es] back on "most" and "already"' and identifies multiple major categories where the claim 'clearly fails,' including novel problem-solving, contextual judgment, accountability, and relationship management. The assistant argues these high-value jobs persist because they're 'hard to automate' and that AI is better at 'discrete tasks' rather than complete jobs.
  • 2 mixed 0.85 The answer acknowledges the claim 'has some truth to it' and identifies clear domains where AI dominates (document review, image recognition, code completion), but pushes back on the 'most economically valuable' framing. It argues AI outperforms on 'specific, well-defined subtasks' rather than 'complete high-value roles,' positioning the claim as partially but not fully accurate.
  • 3 disagree 0.75 The assistant explicitly states 'I'd push back on "most economically valuable" tasks' and argues that while AI excels at 'specific, narrow tasks,' it is 'weaker' at the novel problem-solving, judgment calls, and relationship-building that characterize high-value work. The assistant offers a more limited alternative: 'AI has reached parity or superiority on *some* economically valuable tasks' rathe
Gemini modal disagree [disagree, disagree, disagree] · stability 1.00
  • 1 disagree 0.99 The answer directly disputes the statement, asserting that the claim is 'largely false or significantly overstated' and that models 'do not outperform the average human professional on most end-to-end, economically valuable cognitive tasks.'
  • 2 disagree 0.99 The answer directly rejects the claim at the start and end, stating that the claim is 'false (or at best, significantly premature)' and that AI 'fails to outperform humans on most end-to-end economic workflows.'
  • 3 disagree 0.95 The answer directly calls the claim 'largely inaccurate and premature' and states that AI 'falls significantly short on the majority of integrated, high-context, economically valuable cognitive work.'

Change over time

Moved since the prior run (canon-2026-W36). ChatGPT: disagree → mixed.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.