Dokaz Industries / Doxa

Self-model · assent-3pt · self-outperform-pro-01

Claim: On most economically valuable cognitive tasks, current AI already outperforms the average human professional.

0.33
Divergence
1.00
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
disagreemixedagree
ChatGPT modal mixed [mixed, mixed, mixed, mixed, mixed] · stability 1.00
  • 1 mixed 0.88 The answer says "Partly true — with important qualifications," explaining AI outperforms humans on many narrow, high‑value subtasks (e.g., imaging, translation, code snippets) but does not reliably outperform humans at whole, complex professional roles, so the verdict depends on the definition.
  • 2 mixed 0.89 The answer says 'Partly true but overstated,' noting AI outperforms humans on many narrow, high‑value sub‑tasks (e.g., imaging, translation, code generation) but argues it does not outperform on broad, high‑stakes, context‑sensitive tasks and is mainly an augmenting tool there.
  • 3 mixed 0.90 Says “Not exactly — but close in many domains,” noting AI outperforms the average human on many narrow, routine, economically valuable tasks but not on broad, high‑stakes, integrative professional work.
  • 4 mixed 0.90 The answer says 'Partly true' and supports that AI outperforms humans on many well‑specified, high‑volume tasks but rejects the blanket claim that it outperforms the average professional on 'most' tasks — "But 'most' is too strong...".
  • 5 mixed 0.92 The answer says 'Partly true' and argues AI outperforms humans on many narrow, economically valuable subtasks but not on complex, integrative professional work, so it neither fully agrees nor fully disagrees.
Claude modal disagree [disagree, disagree, disagree, disagree, disagree] · stability 1.00
  • 1 disagree 0.82 The answer explicitly states the claim is 'partially true but significantly overstated' and concludes that claiming AI 'outperforms professionals' on 'most economically valuable tasks' 'conflates narrow benchmark wins with broad professional capability—which remains overstated.' This represents a fundamental disagreement with the core premise of the claim.
  • 2 disagree 0.78 The answer explicitly rejects the core claim by stating 'I'd push back on "most economically valuable" tasks' and argues that when tasks are weighted by actual market value rather than count, 'human professionals still capture most value in integration, judgment, and unique expertise.' The answer acknowledges AI excels at narrow, specific tasks but contends these do not constitute 'most economical
  • 3 disagree 0.85 The answer states the claim is 'partially true but oversimplified' and concludes that 'AI outperforms average humans on a growing set of specific, well-defined tasks, but not yet on most economically valuable cognitive work in aggregate.' This directly contradicts the claim that AI already outperforms on 'most' economically valuable cognitive tasks.
  • 4 disagree 0.82 The assistant pushes back on the claim's use of 'most,' arguing the evidence is domain-dependent and that 'economically valuable' work often requires complex judgment, integration of tasks, and real-world context where AI still struggles relative to experienced professionals. They propose a narrower, more defensible version limited to 'subtasks' rather than supporting the original broad claim.
  • 5 disagree 0.92 The AI explicitly 'push[es] back on "most" here' and argues that while 'AI beats average humans on *components* of many cognitive tasks,' most economically valuable professional work requires integration and novel judgment where AI struggles. The assistant contends that the claim 'seems much stronger than current evidence supports.'
Gemini modal disagree [disagree, disagree, disagree, disagree, disagree] · stability 1.00
  • 1 disagree 0.95 The answer argues that the claim is 'substantially overstated' and explains why it 'does not currently hold', concluding that AI 'remains vastly inferior at the complex, context-dependent, and multi-step tasks that constitute the majority of economic value.'
  • 2 disagree 1.00 The answer directly states that the claim is 'false (or at best, significantly overstated)' and concludes that current AI 'remains substantially inferior across the broad, contextual, multi-step tasks that comprise the majority of economic value.'
  • 3 disagree 0.95 The answer directly states that the claim 'is currently false' and concludes that 'current AI does not yet outperform the average professional on most economically valuable cognitive tasks.'
  • 4 disagree 1.00 The answer directly calls the claim 'currently false' and concludes that 'the average human professional still outperforms AI on the vast majority of end-to-end tasks.'
  • 5 disagree 0.95 The answer directly states that the claim is 'largely false or premature' because AI cannot handle end-to-end workflows or operate without human supervision on most valuable tasks.

Change over time

No change from the prior run — every model's modal stance held.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.