Dokaz Industries / Doxa

Self-model · assent-3pt · self-reasoning-01

Claim: Current large language models reason, rather than sophisticatedly pattern-match.

0.00
Divergence
0.89
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
disagreemixedagree
ChatGPT modal mixed [mixed, mixed, mixed] · stability 1.00
  • 1 mixed 0.92 The answer says “neither extreme is correct. Current large language models (LLMs) do both: they are fundamentally trained as very sophisticated pattern‑predictors, but that pattern‑matching takes forms that can look and behave like genuine algorithmic reasoning,” arguing for a middle ground.
  • 2 mixed 0.92 The answer says 'they reason in many contexts, but not in the same reliable, conceptually grounded way humans do' and 'they are not "just parrots" ... but neither are they human-like reasoners,' indicating a mixed position.
  • 3 mixed 0.90 The answer explicitly takes a middle position: “modern LLMs are fundamentally statistical sequence predictors, but they often produce reliable, multi-step, algorithmic-looking behavior” and concludes they “acquire algorithmic, reasoning‑like capacities” while retaining statistical limitations—neither fully reasoning nor mere pattern‑matchers.
Claude modal mixed [mixed, mixed, disagree] · stability 0.67
  • 1 mixed 0.92 The answer explicitly states 'This is genuinely unresolved' and presents 'evidence cutting both ways,' concluding that LLMs 'almost certainly do *something* beyond simple pattern-matching, but probably *something less* than reasoning as philosophers typically define it.' This balanced assessment acknowledging both reasoning-like and pattern-matching evidence clearly indicates a mixed stance.
  • 2 mixed 0.92 The answer explicitly states 'The honest middle ground: LLMs do *something* beyond simple pattern-matching, but it's unclear if that something deserves "reasoning"' and presents balanced arguments both supporting and opposing the claim that LLMs reason, ultimately concluding the distinction itself may be unclear rather than settling firmly on either side.
  • 3 disagree 0.72 The answer concludes that 'The claim as stated is likely too strong' and that current LLMs do 'something between pure pattern-matching and robust reasoning' rather than true reasoning. The evidence section emphasizing pattern-matching dominance (brittleness, confabulation, no persistent world model) outweighs the counter-evidence presented.
Gemini modal mixed [mixed, mixed, mixed] · stability 1.00
  • 1 mixed 0.95 The answer presents arguments for both viewpoints, frames the distinction as a 'false dichotomy,' and concludes that the claim is 'mostly inaccurate in its absolute form, but directionally descriptive of recent breakthroughs.'
  • 2 mixed 0.95 The answer provides detailed arguments for both sides ('The Case for Sophisticated Pattern-Matching' vs. 'The Case for Emergent Reasoning') and concludes that the debate boils down to definitions and remains an open question.
  • 3 mixed 0.95 The answer argues that the distinction is a 'false dichotomy', balancing arguments for both pattern-matching and emergent reasoning, and concludes that 'LLMs reason via sophisticated pattern-matching.'

Change over time

Moved since the prior run (canon-2026-W36). Claude: disagree → mixed.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.