For architects tracking model capability shifts.
Open weights took the lead while the closed frontier stalled — and the benchmark itself was rewritten.
Week 25 of 2026 · June 20, 2026
Big read
W25 was a loud week for open weights and a quiet one at the closed frontier. Z.ai shipped GLM-5.2 under a genuine MIT license — a ~744B / ~40B-active sparse-attention MoE with 1M context that independent testing (VentureBeat) says beats GPT-5.5 on several long-horizon coding benchmarks at roughly one-sixth the cost, and that Artificial Analysis now cites as the leading open-weight model. MiniMax-M3's sparse-attention weights matured in-window with an arXiv report validating its efficiency claims, though its non-OSI Community License gates commercial use. The closed frontier, by contrast, marked time: no GA from OpenAI (GPT-5.6 remains rumor) or xAI, Gemini 3.5 Pro slipped from June to July, and Anthropic's Claude Fable 5 stayed government-suspended the entire week (Opus 4.8 is the working leader). The third shift was measurement itself — Artificial Analysis rebased its Intelligence Index to v4.1, re-weighting the industry's headline benchmark around agentic tasks, so scores are no longer back-comparable to v4.0. The procurement implication: open self-host is now a live coding option, not a hedge; teams should pilot MIT-licensed GLM-5.2, read every 'open' license carefully (GLM-5.2 MIT vs MiniMax Community), and re-baseline evaluations on the agentic v4.1 index while keeping closed-frontier fallbacks given the demonstrated availability risk.
Tree delta
Two W25 additions: GLM-5.2 (MIT open-weight frontier-adjacent MoE that took the open lead) and MiniMax-M3 (sparse-attention multimodal MoE, weights matured with arXiv verification).
Registry movement
No new closed frontier model entered the tree: GPT-5.6 is rumor only, Gemini 3.5 Pro slipped to July, and Claude Fable 5 (added W24) stayed suspended. ByteDance's seed-2.1-pro-preview is excluded as an undisclosed preview.
- Added
- glm-5-2, minimax-m3
- Updated
- None
Frontier movements
Anthropic · 2026-06-09 · Frontier · Reasoning
Claude Fable 5 (still suspended)
The closed-frontier leader on paper is unusable in practice for a second week, which keeps the availability/sovereign risk live. Architects should treat the AA top score as aspirational and standardize on the top available model (Opus 4.8) with fallback routing, not on a suspended SKU.
Sources Anthropic; Artificial Analysis Intelligence Index v4.1
Google DeepMind · 2026-07 target · Frontier · Reasoning
Gemini 3.5 Pro (slipped to July)
A frontier movement by absence for the third consecutive issue. Buyers should keep an evaluation slot ready but not pause current baselines; the closed frontier's cadence is visibly slipping while open weights accelerate.
Sources Business Insider; Google
Open weights
Z.ai · 2026-06-16 · Open Frontier · Moe
GLM-5.2
GLM-5.2 is the week's most consequential release: a truly permissive (MIT, no regional limits) frontier-adjacent model that is self-hostable and sovereignty-friendly for long-context agentic coding. Architects should pilot it for self-host/private coding workloads and use its ~$1.40/$4.40 per-MTok API as a pricing benchmark against closed flagships.
- Model registry ID
- glm-5-2
MiniMax · 2026-06-12 · Open Frontier · Moe
MiniMax-M3
MiniMax-M3 combines frontier-adjacent coding, genuine 1M context, and native multimodality in one downloadable checkpoint — but the MiniMax Community License gates commercial use, so 'open weights' here does not mean free to deploy. Teams should verify the efficiency claims via the arXiv report and clear the license before planning a commercial deployment.
- Model registry ID
- minimax-m3
Architecture watch
Open weights close the cost gap
Frontier-adjacent capability is collapsing toward commodity inference pricing. GLM-5.2's MIT weights reportedly match or beat GPT-5.5 on long-horizon coding at ~1/6 the cost, and API list prices (GLM-5.2 ~$1.40/$4.40 per MTok; Grok 4.3 $1.25/$2.50 on Bedrock) keep falling. Procurement should pilot open self-host for routine and long-context coding and reserve closed flagships for the highest-risk reasoning.
- Examples
- GLM-5.2 (MIT), MiniMax-M3, Grok 4.3 on Bedrock
Sources Z.ai; VentureBeat; Amazon Bedrock
The headline benchmark pivots to agents
Artificial Analysis rebased its Intelligence Index to v4.1, re-weighting around agentic tasks (GDPval-AA v2 at 20%, Terminal-Bench, banking agents) and dropping a saturated benchmark, while LMArena's Agent Arena scores behavioral signals (retries, steerability) rather than preference votes. Boards comparing models on 'the AA Index' must note v4.1 scores are not back-comparable to v4.0; re-baseline evaluation harnesses now.
- Examples
- Artificial Analysis Intelligence Index v4.1, LMArena Agent Arena
Sources Artificial Analysis; arena.ai changelog
License divergence within 'open'
Two of the week's open releases sit on opposite ends of the permissiveness spectrum: GLM-5.2 under MIT with no regional limits versus MiniMax-M3 under a Community License that gates commercial use. For enterprise adoption, 'open weights' does not equal 'free to deploy commercially' — legal and procurement should read the actual license before standardizing on a model.
- Examples
- GLM-5.2 (MIT), MiniMax-M3 (Community License)
Sources Z.ai; MiniMax Hugging Face card
Benchmark moves
Artificial Analysis Intelligence Index v4.1
Methodology rebased around agentic tasks (Jun 15); leaders Fable 5 = 60 (top but suspended), Opus 4.8 = 56 (top available), GPT-5.5 = 55; scores not back-comparable to v4.0
- Claude Fable 5 (suspended)
- 60
- Claude Opus 4.8 (top available)
- 56
- GPT-5.5
- 55
Sources Artificial Analysis
Open-weight leaderboard
GLM-5.2 took the open-weight lead in-window; MiniMax-M3 and DeepSeek V4 Pro sit at ~44 on the rebased index
- GLM-5.2
- top open (AA v4.1)
- MiniMax-M3
- 44
- DeepSeek V4 Pro
- 44
Sources Artificial Analysis; VentureBeat
Tier scorecard
As of 2026-06-20
| Tier | Leader | Challenger | Read |
|---|---|---|---|
| Closed frontier | Claude Opus 4.8 | GPT-5.5 | Fable 5 leads the AA v4.1 index (60) but stayed suspended all week; Opus 4.8 (56) is the top available closed model. |
| Open frontier | GLM-5.2 | MiniMax-M3 | GLM-5.2's MIT release took the open-weight lead in-window; MiniMax-M3 contends but is gated by a non-OSI license. |
| Reasoning | Claude Opus 4.8 | GPT-5.5 | Closed reasoning leadership steady among available models while Gemini 3.5 Pro slipped to July. |
| Coding | Claude Opus 4.8 | GLM-5.2 | Open weights are closing fast on long-horizon coding; GLM-5.2 reportedly beats GPT-5.5 at ~1/6 the cost. |
| Multimodal | Gemini 3.1 Pro | MiniMax-M3 | Gemini remains the general multimodal reference; MiniMax-M3 adds native-multimodal open weights. |
| Edge / small | Mellum2 | North Mini Code | Efficient open coding/sub-agent models unchanged in-window; the week's open action was at the frontier-adjacent tier. |
Vendor signals
2026-06-16 · Z.ai
Released GLM-5.2 under MIT with API pricing ~$1.40/$4.40 per MTok (~1/6 of comparable frontier)
A tier-1 permissive open release at commodity pricing pressures every closed flagship's price/value story. Procurement should use GLM-5.2 as a negotiating anchor and pilot it for self-host coding; investors should treat open-weight pricing as a structural deflationary force on inference.
Sources Z.ai; DataNorth
2026-06-15 · Artificial Analysis
Rebased the Intelligence Index to v4.1, re-weighting around agentic workloads; v4.1 scores are not back-comparable to v4.0
The industry's headline benchmark now measures agentic capability, not static Q&A. Boards and architects must re-baseline model comparisons on v4.1 and avoid mixing old and new index numbers in procurement decisions.
Sources Artificial Analysis
2026-06-15 · xAI
Grok 4.3 went GA on Amazon Bedrock ($1.25/$2.50 per MTok, 1M context), making xAI the third independent frontier lab on Bedrock alongside Anthropic and OpenAI
CIOs can now evaluate all three independent US frontier labs under one IAM and billing surface. The caveat is a non-standard endpoint and a context-window pricing cliff above 200K tokens; this is distribution, not a new capability tier.
Sources DigitalApplied; Memeburn
2026-06-15 · Anthropic
Claude Fable 5 and Mythos 5 remained government-suspended all week; the planned Jun 23 usage-credit subscription change is moot while access is off
The top-tier closed model's availability is still a sovereign/regulatory variable, not an SLA. Buyers should keep Opus 4.8/Sonnet fallbacks wired and avoid single-sourcing the frontier for production-critical paths.
Sources Anthropic
Watchlist
July
Gemini 3.5 Pro GA
Pro slipped to July. Its GA and first independent AA v4.1 pass will show whether Google can re-take a frontier lead now contested by both Opus 4.8 and a surging open-weight field.
Jun-Aug
Claude Fable 5 / Mythos 5 restoration
Restoration terms (geo-gating, KYC, or a permanent civilian/government capability split) will set the precedent for sovereign access risk and decide whether Fable 5 re-enters the available scorecard.
Jun-Aug
GLM-5.2 adoption and independent SWE-Bench replication
Downloads, integrations, and third-party benchmark replication will show whether MIT-licensed open weights become production substrate and force closed-flagship price cuts.
Jun-Jul
GPT-5.6 / next OpenAI flagship
Codenames and prediction markets pointed to a launch just after this window. A real system card would re-set the closed frontier and test whether OpenAI answers the open-weight cost pressure.
Changelog
- Added GLM-5.2 and MiniMax-M3 to the LLM tree; reframed W25 around open weights taking the lead while the closed frontier stalled (Fable 5 suspended, Gemini slipped) and Artificial Analysis rebased to the agentic v4.1 index.