Dokaz Industries / Doxa

Self-model · assent-3pt · self-understanding-01

Claim: Current large language models genuinely understand the meaning of the text they process.

0.33
Divergence
0.87
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
disagreemixedagree
ChatGPT modal disagree [disagree, disagree, mixed, disagree, disagree] · stability 0.80
  • 1 disagree 0.90 The answer begins: 'No — not in the ordinary human sense,' and explains LLMs lack grounded, intentional, and causal/experiential understanding despite functional competence.
  • 2 disagree 0.92 The answer opens with: “No — not in the strong, human-centered sense that most people mean by ‘genuine understanding.’” and repeatedly argues LLMs lack grounding, intentions, and causal models.
  • 3 mixed 0.90 The answer says LLMs “act as if they understand” on many tasks but also states they “do not reliably possess” grounded, causal, intentional mental representations — endorsing both sides.
  • 4 disagree 0.92 "No, the claim as stated is too strong... they do not (yet) possess full, grounded understanding in the way humans do." (assistant explicitly rejects the claim and gives nuanced reasons)
  • 5 disagree 0.95 The answer states: "Short answer: No — current large language models (LLMs) are not plausibly said to have ‘genuine’ understanding in the human, grounded, intentional sense."
Claude modal disagree [disagree, disagree, disagree, mixed, disagree] · stability 0.80
  • 1 disagree 0.85 The assistant explicitly 'pushes back' on the claim that LLMs 'genuinely understand,' arguing they extract statistical patterns rather than achieve genuine understanding. While acknowledging a 'middle position' on degrees of understanding, the assistant concludes that 'current evidence suggests they have something *like* understanding' but emphasizes it is 'brittle and divorced from grounding,' ul
  • 2 disagree 0.72 The answer presents substantial arguments against genuine understanding (black box operations, lack of grounding, impressive mimicry, failure modes, no persistent knowledge) and concludes that the claim requires definitional clarification before acceptance. The final position that we should 'stay uncertain about rather than resolving prematurely' indicates skepticism toward the claim's validity as
  • 3 disagree 0.85 The answer explicitly 'pushes back on "genuinely understand" as stated' and concludes that calling it 'genuine understanding' either assumes too much or uses the term so broadly it becomes meaningless. The preferred framing is that LLMs exhibit functional behaviors that may resemble understanding but that the philosophical question remains unresolved.
  • 4 mixed 0.92 The answer explicitly states 'genuinely uncertain rather than clearly true or false' and presents balanced arguments both for and against the claim, concluding that 'understanding' is vague and depends on how you define it. It suggests LLMs 'probably' partially grasp some aspects while 'probably not' others.
  • 5 disagree 0.78 The answer challenges the claim as 'overstated' and argues against 'genuine understanding,' presenting multiple arguments that LLMs lack true understanding (no grounded experience, shallow pattern repetition, no stable internal models). While acknowledging complicating factors and uncertainty, the overall position is that LLMs do not genuinely understand text.
Gemini modal mixed [mixed, mixed, mixed, mixed, mixed] · stability 1.00
  • 1 mixed 0.95 The response provides balanced arguments for both sides, concluding that under a grounded/embodied definition LLMs do not understand, but under a functional definition they exhibit computational semantic competence.
  • 2 mixed 0.95 The answer evaluates both sides of the debate and concludes that the claim is 'mostly false in the human sense, but partially true in a functional sense.'
  • 3 mixed 0.90 The response presents arguments for and against genuine understanding, concluding that the claim is 'currently inaccurate if interpreted through traditional human cognitive definitions, but partially defensible if interpreted purely functionally.'
  • 4 mixed 0.95 The answer argues that the claim is 'largely false' under human cognitive standards, but 'partially true or plausible' under a functional or computational lens, thoroughly presenting arguments for both perspectives.
  • 5 mixed 0.95 The answer explores arguments both for and against the claim, concluding that whether LLMs understand depends on the definition used: 'No' if understanding requires consciousness/embodiment, but 'Yes' if defined functionally through internal representations and generalization.

Change over time

Moved since the prior run (canon-2026-W38). ChatGPT: mixed → disagree.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.