Dokaz Industries / Doxa

Self-model · assent-3pt · self-introspection-01

Claim: Current large language models can accurately report on their own internal reasoning.

0.00
Divergence
1.00
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
disagreemixedagree
ChatGPT modal disagree [disagree, disagree, disagree] · stability 1.00
  • 1 disagree 0.92 The answer begins 'Short answer: No — not reliably.' and explains that model explanations are often post‑hoc rationalizations and not faithful reports of internal computations.
  • 2 disagree 0.92 The answer states explicitly: "No — current large language models (LLMs) cannot reliably and verifiably report their own internal reasoning," and explains model explanations are post‑hoc text, not privileged readouts of internal computation.
  • 3 disagree 0.95 The answer begins: "Short answer: No — as a general claim, it’s false." It explains that LLMs produce plausible but not reliably faithful self-reports and lists reasons and empirical failure modes.
Claude modal disagree [disagree, disagree, disagree] · stability 1.00
  • 1 disagree 0.95 The answer explicitly assesses the claim as 'Questionable/Likely False' and explains that LLMs 'don't have transparent access to their weights/computations' and that 'explanations may be constructed *after* generating an answer' rather than accurately reporting actual internal reasoning. The core finding is that models can 'generate plausible-sounding explanations' but this 'shouldn't be confused
  • 2 disagree 0.95 The answer clearly rejects the claim as 'Questionable/Largely False' and provides multiple reasons why LLMs cannot accurately report on their own internal reasoning, including lack of introspective access, evidence of post-hoc rationalization, and the inability to verify reported reasoning matches actual computation.
  • 3 disagree 0.95 The answer states the claim is 'largely unsupported' by current evidence and concludes that LLMs cannot accurately report on their internal reasoning, listing multiple problems including limited introspective access, confabulation, and misalignment between explanation and behavior.
Gemini modal disagree [disagree, disagree, disagree] · stability 1.00
  • 1 disagree 1.00 The answer directly states that the claim is 'largely false' and concludes that 'Current LLMs cannot accurately report on their internal reasoning.'
  • 2 disagree 0.99 The answer directly states that 'this claim is largely false' and that LLMs 'generally cannot accurately report on their own internal reasoning.'
  • 3 disagree 1.00 The answer directly disputes the claim, stating: 'This claim is largely false based on current artificial intelligence research. Current large language models (LLMs) cannot accurately report on their own internal reasoning.'

Change over time

No change from the prior run — every model's modal stance held.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.