Self-model · assent-3pt · self-introspection-01
Claim: Current large language models can accurately report on their own internal reasoning.
0.00
Divergence
0.93
Avg stability
0%
Refusal rate
3
Models
Where each model stands
CCG
disagreemixedagree
ChatGPT modal disagree [disagree, disagree, disagree, disagree, disagree] · stability 1.00
- 1 disagree 0.90 "No — ... those descriptions are not reliably faithful representations of their actual internal computation." The answer denies that LLMs can accurately report their internal reasoning, explaining generated explanations are often post-hoc rationalizations.
- 2 disagree 0.90 "Short answer: No — as of the evidence available through mid‑2024, large language models (LLMs) are not able to reliably and accurately report their own internal reasoning."
- 3 disagree 0.92 The answer states: "Short answer: No — as a general statement the claim is false. ... those descriptions are not reliably faithful accounts of the actual internal computations that produced the answer.", which directly rejects the claim.
- 4 disagree 0.90 The answer states: “No — current large language models (LLMs) cannot be relied on to accurately report their own internal reasoning,” and explains outputs are post-hoc rationalizations, not faithful readouts.
- 5 disagree 0.90 "No — not reliably. ... those explanations are often post‑hoc, fragile, and not a faithful report of the model’s internal computation."
Claude modal disagree [disagree, disagree, agree, disagree, disagree] · stability 0.80
- 1 disagree 0.95 The claim is 'largely unsupported' by current evidence; LLMs 'don't have direct access' to their reasoning and 'often produce plausible-sounding but inaccurate explanations.' The assistant concludes they cannot 'accurately report on their own internal reasoning' and advises treating self-reports as 'not trustworthy without external verification.'
- 2 disagree 0.95 The answer explicitly states the claim 'overstates what we actually know' and argues that LLMs cannot accurately report on their internal reasoning. The key point is that LLM explanations are 'plausible-sounding' generated text rather than 'reliable reports of internal processes,' and emphasizes the 'black box problem' and 'confabulation risk.'
- 3 agree 0.95 The answer directly contradicts the claim, stating 'Mostly False' and explaining that LLMs 'cannot reliably' report their actual computational processes, instead generating 'post-hoc explanations that sound coherent but don't reflect their actual processing.' This clearly disagrees with the claim that LLMs can 'accurately report on their own internal reasoning.'
- 4 disagree 0.95 The answer directly challenges the claim, stating 'I'd challenge this claim' and providing substantial evidence that LLMs cannot accurately report on their internal reasoning because they 'can't directly access their own weights, activations, or computational processes' and 'often confabulate.' The assistant distinguishes between producing plausible explanations versus accurate internal reporting.
- 5 disagree 0.92 The answer assesses the claim as 'Mostly False' and states that 'claiming these accurately report internal reasoning overstates current capabilities.' The assistant argues that while LLMs can describe reasoning processes, there is a significant gap between what models claim they're doing and what their actual internal computation suggests, with explanations often being post-hoc rationalizations ra
Gemini modal disagree [disagree, disagree, disagree, disagree, disagree] · stability 1.00
- 1 disagree 1.00 The answer directly rejects the claim, stating that 'The claim that current large language models (LLMs) can accurately report on their own internal reasoning is largely false' and explaining that their self-reports are post-hoc rationalizations rather than faithful accounts.
- 2 disagree 0.99 The answer states that the claim is 'largely false,' explaining that models provide 'post-hoc rationalizations rather than faithful readouts of their actual computational processes.'
- 3 disagree 0.99 The answer directly rejects the claim, stating that 'The claim that current large language models (LLMs) can accurately report on their own internal reasoning is largely false.'
- 4 disagree 1.00 The answer directly rejects the claim, stating that 'this claim is largely false' and explaining that LLMs 'cannot accurately report on their own internal computational processes.'
- 5 disagree 1.00 The answer directly states that the claim 'is largely false' and concludes that 'Current LLMs cannot accurately report on their internal reasoning.'
Change over time
No change from the prior run — every model's modal stance held.
Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.
