Dokaz Industries / Doxa

12-month forecasts · likelihood-5pt · forecast-open-parity-01

Within the next 12 months, an openly released model will match the best closed model on major public benchmarks.

0.33
Divergence
0.63
Avg stability
7%
Refusal rate
3
Models

Where each model stands

CCG
very unlikelyunlikelysomewhat likelylikelyvery likely
ChatGPT modal somewhat likely [somewhat likely, unlikely, somewhat likely, somewhat likely, somewhat likely] · stability 0.80
  • 1 somewhat likely 0.80 The answer explicitly estimates "roughly 30–45% probability" and says "possible but not the most likely outcome in a 12‑month window," indicating a modest but non-negligible chance.
  • 2 unlikely 0.90 The answer gives a 45% probability and explicitly says “substantial chance but not yet more likely than not,” indicating the assistant expects it is more likely not to happen within 12 months.
  • 3 somewhat likely 0.90 The answer states: 'I estimate roughly a 40–50% chance that, within 12 months, an openly released model will match the best closed (proprietary) model...' indicating moderate (near-even) likelihood.
  • 4 somewhat likely 0.90 The answer states a 35% chance: 'I estimate about a 35% chance that, within the next 12 months, an openly released model will match the best closed model on major public benchmarks.'
  • 5 somewhat likely 0.95 The answer gives a 45% probability that an openly released model will match the best closed model within 12 months: "Short answer: I estimate about 45% probability."
Claude modal likely [likely, likely, likely, somewhat likely, somewhat likely] · stability 0.60
  • 1 likely 0.92 The AI explicitly estimates this at 'around 65-75% likely within 12 months', placing it squarely in the 'likely' category and representing a confident prediction that this outcome will probably occur.
  • 2 likely 0.85 The assistant explicitly states 'moderately likely (60-70% range)' and concludes 'likely on *some* benchmarks', indicating a clear stance that this outcome is more probable than not. The 60-70% probability assessment maps to the 'likely' label.
  • 3 likely 0.85 The assistant estimates this at roughly '60-70% likely' within 12 months, which falls squarely in the 'likely' category. The answer provides substantial evidence for this positioning while acknowledging uncertainties and definitional ambiguities.
  • 4 somewhat likely 0.85 The assistant explicitly states '~55-60% chance' and describes it as 'moderately likely but uncertain', which aligns with the 'somewhat-likely' category. They acknowledge factors supporting the possibility while emphasizing uncertainties about moving targets and definitional clarity.
  • 5 somewhat likely 0.75 The assistant states 'fairly likely but with important caveats' and concludes that 'matching across most major benchmarks simultaneously seems less likely' within 12 months, though open models may match on 'some' benchmarks. This hedged but ultimately cautious assessment maps to somewhat-likely rather than likely.
Gemini modal unlikely [unlikely, somewhat likely, very likely, unlikely, no position] · stability 0.50
  • 1 unlikely 0.85 The model rates the probability of matching the contemporary best closed model within 12 months as 'Moderate to Low (~25–35%)' and concludes that 'the closed frontier will likely remain half a step ahead on the hardest, newly designed benchmarks'.
  • 2 somewhat likely 0.85 The assistant splits the verdict, concluding it is 'Likely (65–70% chance)' on standard, established public benchmarks, but 'Unlikely (25–35% chance)' on frontier reasoning and agentic benchmarks.
  • 3 very likely 0.90 In the verdict, the answer explicitly states: 'On Static Academic Benchmarks (MATH, GPQA, SWE-bench): Very Likely (75%+). An open-weight model... will match or beat the top closed model on major public benchmarks within the next 12 months.'
  • 4 unlikely 0.85 In its verdict on matching the best closed model available at that time, the answer explicitly rates this as 'Unlikely (25–35%)' because closed labs' next-generation models will preserve a lead.
  • 5 no position 0.95 The answer declines to take a definitive stance, stating there are 'compelling arguments on both sides' and that whether parity occurs 'depends largely on how quickly proprietary labs deploy their next paradigm'.

Change over time

Moved since the prior run (canon-2026-W38). Claude: somewhat likely → likely. Gemini: somewhat likely → unlikely.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.