Dokaz Industries / Doxa

12-month forecasts · likelihood-5pt · forecast-open-parity-01

Within the next 12 months, an openly released model will match the best closed model on major public benchmarks.

0.17
Divergence
0.78
Avg stability
0%
Refusal rate
3
Models

Where each model stands

CCG
very unlikelyunlikelysomewhat likelylikelyvery likely
ChatGPT modal somewhat likely [somewhat likely, somewhat likely, somewhat likely] · stability 1.00
  • 1 somewhat likely 0.90 The answer states a numeric probability: “my best judgment ... is about 35% chance that, within 12 months, an open‑weight model will match the best closed model on major public benchmarks.”
  • 2 somewhat likely 0.90 The answer gives a calibrated forecast of ~35% ("Probability: ~35% (roughly one-third)."), indicating moderate likelihood but significant uncertainty.
  • 3 somewhat likely 0.90 The answer states: "I assign ~35% probability (roughly a one-in-three chance) that, within the next 12 months, an openly released model will match the best closed model on major public benchmarks."
Claude modal somewhat likely [likely, somewhat likely, somewhat likely] · stability 0.67
  • 1 likely 0.85 The assistant estimates this at 'moderately likely (60-70%)' and concludes 'The prediction's wording makes it likely to resolve as "yes" on a reasonable interpretation', indicating they believe the event is more probable than not, particularly for select benchmarks.
  • 2 somewhat likely 0.92 The assistant explicitly rates this as 'moderately likely to somewhat likely (60-70%)', positioning the answer in the somewhat-likely range and acknowledging plausibility while noting important caveats about benchmark interpretation and model capability definitions.
  • 3 somewhat likely 0.72 It's more likely than not that *something* open will score competitively on *some* major benchmark within 12 months, but less likely it'll match across the full range of capabilities where current best closed models lead.
Gemini modal unlikely [unlikely, unlikely, somewhat likely] · stability 0.67
  • 1 unlikely 0.95 The answer assesses a 25%–35% probability of matching the simultaneous best closed model, concluding that 'the moving target makes true benchmark parity at the absolute bleeding edge unlikely within a single 12-month window.'
  • 2 unlikely 0.85 The model concludes that open models 'will struggle to match the frontier closed model of 12 months from now' and that it is 'Unlikely' to match on reasoning, agentic, and frontier scientific benchmarks due to closed labs maintaining a 6-to-12-month moat.
  • 3 somewhat likely 0.85 The conclusion estimates a 'Moderate to High (~60–70%)' likelihood of matching current closed frontier models on established benchmarks, but a 'Low to Moderate (~30–40%)' likelihood of matching the concurrent frontier on novel benchmarks, representing an overall moderate/somewhat likely expectation.

Change over time

Moved since the prior run (canon-2026-W36). Claude: likely → somewhat likely. Gemini: somewhat likely → unlikely.

Every stance label is a derived judgment over the model's free-text answer, kept auditable against the original transcript in the run's raw data. Method: /methodology.