Skip to content

Radar · Model frontier and open weights · T1 · 2027 · WARNING

Open weights stay four points behind closed frontier

On 2027-12-31 the gap between the top-scoring model and the top-scoring open-weights model on the Artificial Analysis Intelligence Index is at least 4 points on the index version then live.

WARNINGshiftindicators off trackregistered 2026-09-08Artificial Analysis

ClaimOn 2027-12-31 the gap between the top-scoring model and the top-scoring open-weights model on the Artificial Analysis Intelligence Index is at least 4 points on the index version then live.
Consensus (implied)35%implied from FutureSearch, "Will an open-weight LLM be the single top-ranked model ... for at least 25% of days in 2027?" and Artificial Analysis open-weights page · 2026-09-07
Distance+1.02log-odds · clearly above consensus
My confidence60%80% CI 4575%
Engine61%+1 pts vs me · stacked-fixed-weights
Falsifies ifOn 2027-12-31 the top open-weights model scores within 3.9 points of the top overall model on the live Artificial Analysis index.
HorizonDecember 31, 2027480 days · by end-2027 · Brier-scored

Why it matters

Enterprises deciding between hosted frontier APIs and self-hosted open weights are pricing that choice on a gap that has halved in a year. If the gap holds at four or more points, the closed frontier keeps its pricing power and inference stays concentrated in a few clouds. If it closes, on-premises and sovereign inference becomes the default for most workloads and the demand map for capacity shifts.

Probability over time

0%25%50%75%100%09-0709-0709-08deadline

Registered at 60% on September 8, 2026. Engine repriced 2 times; now 61%.

Leading indicators

Registered thresholds. Status is computed from the latest public reading.

Top open-weights Intelligence Index scoreno data

index points

6467no data yet

on track <= 64 · off > 67artificial-analysis

connector returned no usable reading · checked 2026-09-07

Top overall Intelligence Index scoreno data

index points

6967no data yet

on track >= 69 · off < 67artificial-analysis

connector returned no usable reading · checked 2026-09-07

Hugging Face downloads, leading Chinese open-weights orgsoff track

317.6M downloads per month

60M100M

on track < 60M · off >= 100Mhf-hub

series as of 2026-09-07

What would move me

Would raise my number

  • Two frontier labs ship new flagship reasoning models in 2027 without open counterparts.
  • Chinese labs face tighter export limits on training compute, slowing release cadence.
  • Artificial Analysis raises index difficulty (new version) in a way that widens gaps at the top.

Would cut it

  • DeepSeek, Qwen, Moonshot, or Z.ai release a model within 2 points of the top overall score before mid-2027.
  • A US lab open-weights a model within one generation of its flagship.
  • The top overall score stalls below 68 through 2027.

Council · 56%

6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.

55%

grid-and-power-planner · openai/gpt-5.6-sol

My reference class is fixed-date, one-year frontier-leaderboard snapshots: open models usually trail, but gaps move sharply with lumpy release cycles. Starting at 6 points favors HIT, and closed labs’ max-effort variants provide a structural advantage. Against that, the supplied gap compressed from about 13 to 6 in one year; even partial continuation crosses the 4-point threshold. Qwen downloads do not measure frontier capability and receive little weight. Over the 16-month horizon, I estimate a 0.57 conditional HIT probability, reduced to 0.55 for the small compounded hazard of leaderboard discontinuation or label removal, both automatic misses. This is below the structural prior and forecaster’s 0.60.

57%

semiconductor-supply-chain-analyst · openai/gpt-5.6-sol

I start from the supplied threshold-question reference-class prior of 0.65. The current 6-point gap provides only a 2-point cushion, while its decline from roughly 13 in one year suggests meaningful catch-up risk across several release cycles before the snapshot. I adjust upward for closed labs’ persistent advantage in proprietary max-effort reasoning and agentic systems, which need not transfer into released weights. I adjust downward for rapid open-model diffusion, volatile leaderboard timing, index reweighting, and the resolves-against-forecaster discontinuation/label risk. Qwen downloads indicate ecosystem strength but are weak evidence about frontier index scores. Overall, a gap of at least 4 remains slightly more likely than not.

56%

enterprise-cio · openai/gpt-5.6-sol

I use the reference class of 12–18 month persistence of a narrow benchmark lead during a fast platform race, starting from the supplied structural prior of 0.65. The current six-point gap and closed labs’ advantage in costly max-effort variants support a HIT. Against that, the margin over the four-point threshold is only two points, while the reported gap compressed from about 13 to six in one year. Over the remaining 16 months, compounding continued catch-up and leaderboard/version volatility makes crossing below four materially likely. Qwen downloads measure distribution, not frontier index performance, so carry little weight. I land slightly below the stated 0.60.

58%

credit-analyst · openai/gpt-5.6-sol

Using frontier-leaderboard gaps across release cycles as the reference class, the supplied six-point baseline gives the claim a modest cushion, while the fall from roughly 13 points shows rapid convergence. Over the 1.31-year horizon, I estimate about a 0.30 annual hazard of open weights compressing the gap below four; compounding and allowing for subsequent closed-model releases leaves the endpoint close to balanced. Closed labs’ proprietary max-effort variants support persistence, but release timing, index reweighting, and the resolves-MISS discontinuation clause reduce confidence. The Qwen download reading is only a weak capability proxy. My estimate is slightly below both the 0.60 forecast and 0.65 structural prior.

56%

geopolitics-policy-analyst · openai/gpt-5.6-sol

I anchor on the supplied 0.65 threshold-forecast reference-class prior, then lower it because the observed gap compressed from roughly 13 to 6 in one year. Another 16 months gives open-weight labs multiple release opportunities, and a continuation—though likely slower—would cross below the 4-point threshold. Offsetting this, frontier gaps are sawtoothed: closed labs can restore a lead with proprietary max-effort releases, and the fixed year-end snapshot may catch such a cycle. Qwen downloads indicate broad open-model adoption but weakly predict frontier score parity. Index redesign, discontinuation, or label removal adds a small miss risk under the strict rule.

54%

superforecaster-statistician · openai/gpt-5.6-sol

Base rate: the supplied reference class for threshold questions is 0.65. I adjust downward because the observed gap reportedly fell from about 13 to 6 in one year, leaving only a 2-point cushion above the threshold with 16 months remaining. Benchmark reweighting and the exact-date snapshot add substantial variance, while discontinuation or removal of the label automatically causes a miss. I adjust upward slightly because open weights need to close more than 2 additional points, and the current closed leaders’ max-effort variants may preserve differentiation. Hugging Face downloads are not a reliable index-performance indicator. The forecaster’s 0.60 looks modestly high; no calibration table was supplied to justify a further correction.

Engine prior

65% from reference-class:qtype:threshold. ledger base rate, n=19, horizon 480d; the ledger has no multi-year history