Top open-weights Intelligence Index scoreno data
— index points
on track <= 64 · off > 67artificial-analysis ↗
connector returned no usable reading · checked 2026-09-07
Radar · Model frontier and open weights · T1 · 2027 · WARNING
On 2027-12-31 the gap between the top-scoring model and the top-scoring open-weights model on the Artificial Analysis Intelligence Index is at least 4 points on the index version then live.
Enterprises deciding between hosted frontier APIs and self-hosted open weights are pricing that choice on a gap that has halved in a year. If the gap holds at four or more points, the closed frontier keeps its pricing power and inference stays concentrated in a few clouds. If it closes, on-premises and sovereign inference becomes the default for most workloads and the demand map for capacity shifts.
Registered at 60% on September 8, 2026. Engine repriced 2 times; now 61%.
Registered thresholds. Status is computed from the latest public reading.
Top open-weights Intelligence Index scoreno data
— index points
on track <= 64 · off > 67artificial-analysis ↗
connector returned no usable reading · checked 2026-09-07
Top overall Intelligence Index scoreno data
— index points
on track >= 69 · off < 67artificial-analysis ↗
connector returned no usable reading · checked 2026-09-07
Hugging Face downloads, leading Chinese open-weights orgsoff track
317.6M downloads per month
on track < 60M · off >= 100Mhf-hub ↗
series as of 2026-09-07
Would raise my number
Would cut it
6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.
55%
grid-and-power-planner · openai/gpt-5.6-sol
My reference class is fixed-date, one-year frontier-leaderboard snapshots: open models usually trail, but gaps move sharply with lumpy release cycles. Starting at 6 points favors HIT, and closed labs’ max-effort variants provide a structural advantage. Against that, the supplied gap compressed from about 13 to 6 in one year; even partial continuation crosses the 4-point threshold. Qwen downloads do not measure frontier capability and receive little weight. Over the 16-month horizon, I estimate a 0.57 conditional HIT probability, reduced to 0.55 for the small compounded hazard of leaderboard discontinuation or label removal, both automatic misses. This is below the structural prior and forecaster’s 0.60.
57%
semiconductor-supply-chain-analyst · openai/gpt-5.6-sol
I start from the supplied threshold-question reference-class prior of 0.65. The current 6-point gap provides only a 2-point cushion, while its decline from roughly 13 in one year suggests meaningful catch-up risk across several release cycles before the snapshot. I adjust upward for closed labs’ persistent advantage in proprietary max-effort reasoning and agentic systems, which need not transfer into released weights. I adjust downward for rapid open-model diffusion, volatile leaderboard timing, index reweighting, and the resolves-against-forecaster discontinuation/label risk. Qwen downloads indicate ecosystem strength but are weak evidence about frontier index scores. Overall, a gap of at least 4 remains slightly more likely than not.
56%
enterprise-cio · openai/gpt-5.6-sol
I use the reference class of 12–18 month persistence of a narrow benchmark lead during a fast platform race, starting from the supplied structural prior of 0.65. The current six-point gap and closed labs’ advantage in costly max-effort variants support a HIT. Against that, the margin over the four-point threshold is only two points, while the reported gap compressed from about 13 to six in one year. Over the remaining 16 months, compounding continued catch-up and leaderboard/version volatility makes crossing below four materially likely. Qwen downloads measure distribution, not frontier index performance, so carry little weight. I land slightly below the stated 0.60.
58%
credit-analyst · openai/gpt-5.6-sol
Using frontier-leaderboard gaps across release cycles as the reference class, the supplied six-point baseline gives the claim a modest cushion, while the fall from roughly 13 points shows rapid convergence. Over the 1.31-year horizon, I estimate about a 0.30 annual hazard of open weights compressing the gap below four; compounding and allowing for subsequent closed-model releases leaves the endpoint close to balanced. Closed labs’ proprietary max-effort variants support persistence, but release timing, index reweighting, and the resolves-MISS discontinuation clause reduce confidence. The Qwen download reading is only a weak capability proxy. My estimate is slightly below both the 0.60 forecast and 0.65 structural prior.
56%
geopolitics-policy-analyst · openai/gpt-5.6-sol
I anchor on the supplied 0.65 threshold-forecast reference-class prior, then lower it because the observed gap compressed from roughly 13 to 6 in one year. Another 16 months gives open-weight labs multiple release opportunities, and a continuation—though likely slower—would cross below the 4-point threshold. Offsetting this, frontier gaps are sawtoothed: closed labs can restore a lead with proprietary max-effort releases, and the fixed year-end snapshot may catch such a cycle. Qwen downloads indicate broad open-model adoption but weakly predict frontier score parity. Index redesign, discontinuation, or label removal adds a small miss risk under the strict rule.
54%
superforecaster-statistician · openai/gpt-5.6-sol
Base rate: the supplied reference class for threshold questions is 0.65. I adjust downward because the observed gap reportedly fell from about 13 to 6 in one year, leaving only a 2-point cushion above the threshold with 16 months remaining. Benchmark reweighting and the exact-date snapshot add substantial variance, while discontinuation or removal of the label automatically causes a miss. I adjust upward slightly because open weights need to close more than 2 additional points, and the current closed leaders’ max-effort variants may preserve differentiation. Hugging Face downloads are not a reliable index-performance indicator. The forecaster’s 0.60 looks modestly high; no calibration table was supplied to justify a further correction.
65% from reference-class:qtype:threshold. ledger base rate, n=19, horizon 480d; the ledger has no multi-year history