Skip to content

Model layer

The Model Pulse

For architects tracking model capability shifts.

Open weights took the lead while the closed frontier stalled — and the benchmark itself was rewritten.

Big read

W25 was a loud week for open weights and a quiet one at the closed frontier. Z.ai shipped GLM-5.2 under a genuine MIT license — a ~744B / ~40B-active sparse-attention MoE with 1M context that independent testing (VentureBeat) says beats GPT-5.5 on several long-horizon coding benchmarks at roughly one-sixth the cost, and that Artificial Analysis now cites as the leading open-weight model. MiniMax-M3's sparse-attention weights matured in-window with an arXiv report validating its efficiency claims, though its non-OSI Community License gates commercial use. The closed frontier, by contrast, marked time: no GA from OpenAI (GPT-5.6 remains rumor) or xAI, Gemini 3.5 Pro slipped from June to July, and Anthropic's Claude Fable 5 stayed government-suspended the entire week (Opus 4.8 is the working leader). The third shift was measurement itself — Artificial Analysis rebased its Intelligence Index to v4.1, re-weighting the industry's headline benchmark around agentic tasks, so scores are no longer back-comparable to v4.0. The procurement implication: open self-host is now a live coding option, not a hedge; teams should pilot MIT-licensed GLM-5.2, read every 'open' license carefully (GLM-5.2 MIT vs MiniMax Community), and re-baseline evaluations on the agentic v4.1 index while keeping closed-frontier fallbacks given the demonstrated availability risk.

Tree delta

Two W25 additions: GLM-5.2 (MIT open-weight frontier-adjacent MoE that took the open lead) and MiniMax-M3 (sparse-attention multimodal MoE, weights matured with arXiv verification).

Registry movement

No new closed frontier model entered the tree: GPT-5.6 is rumor only, Gemini 3.5 Pro slipped to July, and Claude Fable 5 (added W24) stayed suspended. ByteDance's seed-2.1-pro-preview is excluded as an undisclosed preview.

Added
glm-5-2, minimax-m3
Updated
None

Frontier movements

Anthropic · 2026-06-09 · Frontier · Reasoning

Claude Fable 5 (still suspended)

The closed-frontier leader on paper is unusable in practice for a second week, which keeps the availability/sovereign risk live. Architects should treat the AA top score as aspirational and standardize on the top available model (Opus 4.8) with fallback routing, not on a suspended SKU.

Sources Anthropic; Artificial Analysis Intelligence Index v4.1

Google DeepMind · 2026-07 target · Frontier · Reasoning

Gemini 3.5 Pro (slipped to July)

A frontier movement by absence for the third consecutive issue. Buyers should keep an evaluation slot ready but not pause current baselines; the closed frontier's cadence is visibly slipping while open weights accelerate.

Sources Business Insider; Google

Open weights

Z.ai · 2026-06-16 · Open Frontier · Moe

GLM-5.2

GLM-5.2 is the week's most consequential release: a truly permissive (MIT, no regional limits) frontier-adjacent model that is self-hostable and sovereignty-friendly for long-context agentic coding. Architects should pilot it for self-host/private coding workloads and use its ~$1.40/$4.40 per-MTok API as a pricing benchmark against closed flagships.

Model registry ID
glm-5-2

Sources Z.ai blog; VentureBeat; Hugging Face

MiniMax · 2026-06-12 · Open Frontier · Moe

MiniMax-M3

MiniMax-M3 combines frontier-adjacent coding, genuine 1M context, and native multimodality in one downloadable checkpoint — but the MiniMax Community License gates commercial use, so 'open weights' here does not mean free to deploy. Teams should verify the efficiency claims via the arXiv report and clear the license before planning a commercial deployment.

Model registry ID
minimax-m3

Sources TechTimes; Hugging Face; arXiv:2606.13392

Architecture watch

Open weights close the cost gap

Frontier-adjacent capability is collapsing toward commodity inference pricing. GLM-5.2's MIT weights reportedly match or beat GPT-5.5 on long-horizon coding at ~1/6 the cost, and API list prices (GLM-5.2 ~$1.40/$4.40 per MTok; Grok 4.3 $1.25/$2.50 on Bedrock) keep falling. Procurement should pilot open self-host for routine and long-context coding and reserve closed flagships for the highest-risk reasoning.

Examples
GLM-5.2 (MIT), MiniMax-M3, Grok 4.3 on Bedrock

Sources Z.ai; VentureBeat; Amazon Bedrock

The headline benchmark pivots to agents

Artificial Analysis rebased its Intelligence Index to v4.1, re-weighting around agentic tasks (GDPval-AA v2 at 20%, Terminal-Bench, banking agents) and dropping a saturated benchmark, while LMArena's Agent Arena scores behavioral signals (retries, steerability) rather than preference votes. Boards comparing models on 'the AA Index' must note v4.1 scores are not back-comparable to v4.0; re-baseline evaluation harnesses now.

Examples
Artificial Analysis Intelligence Index v4.1, LMArena Agent Arena

Sources Artificial Analysis; arena.ai changelog

License divergence within 'open'

Two of the week's open releases sit on opposite ends of the permissiveness spectrum: GLM-5.2 under MIT with no regional limits versus MiniMax-M3 under a Community License that gates commercial use. For enterprise adoption, 'open weights' does not equal 'free to deploy commercially' — legal and procurement should read the actual license before standardizing on a model.

Examples
GLM-5.2 (MIT), MiniMax-M3 (Community License)

Sources Z.ai; MiniMax Hugging Face card

Benchmark moves

Artificial Analysis Intelligence Index v4.1

Methodology rebased around agentic tasks (Jun 15); leaders Fable 5 = 60 (top but suspended), Opus 4.8 = 56 (top available), GPT-5.5 = 55; scores not back-comparable to v4.0

Claude Fable 5 (suspended)
60
Claude Opus 4.8 (top available)
56
GPT-5.5
55

Sources Artificial Analysis

Open-weight leaderboard

GLM-5.2 took the open-weight lead in-window; MiniMax-M3 and DeepSeek V4 Pro sit at ~44 on the rebased index

GLM-5.2
top open (AA v4.1)
MiniMax-M3
44
DeepSeek V4 Pro
44

Sources Artificial Analysis; VentureBeat

Tier scorecard

As of 2026-06-20

TierLeaderChallengerRead
Closed frontierClaude Opus 4.8GPT-5.5Fable 5 leads the AA v4.1 index (60) but stayed suspended all week; Opus 4.8 (56) is the top available closed model.
Open frontierGLM-5.2MiniMax-M3GLM-5.2's MIT release took the open-weight lead in-window; MiniMax-M3 contends but is gated by a non-OSI license.
ReasoningClaude Opus 4.8GPT-5.5Closed reasoning leadership steady among available models while Gemini 3.5 Pro slipped to July.
CodingClaude Opus 4.8GLM-5.2Open weights are closing fast on long-horizon coding; GLM-5.2 reportedly beats GPT-5.5 at ~1/6 the cost.
MultimodalGemini 3.1 ProMiniMax-M3Gemini remains the general multimodal reference; MiniMax-M3 adds native-multimodal open weights.
Edge / smallMellum2North Mini CodeEfficient open coding/sub-agent models unchanged in-window; the week's open action was at the frontier-adjacent tier.

Vendor signals

2026-06-16 · Z.ai

Released GLM-5.2 under MIT with API pricing ~$1.40/$4.40 per MTok (~1/6 of comparable frontier)

A tier-1 permissive open release at commodity pricing pressures every closed flagship's price/value story. Procurement should use GLM-5.2 as a negotiating anchor and pilot it for self-host coding; investors should treat open-weight pricing as a structural deflationary force on inference.

Sources Z.ai; DataNorth

2026-06-15 · Artificial Analysis

Rebased the Intelligence Index to v4.1, re-weighting around agentic workloads; v4.1 scores are not back-comparable to v4.0

The industry's headline benchmark now measures agentic capability, not static Q&A. Boards and architects must re-baseline model comparisons on v4.1 and avoid mixing old and new index numbers in procurement decisions.

Sources Artificial Analysis

2026-06-15 · xAI

Grok 4.3 went GA on Amazon Bedrock ($1.25/$2.50 per MTok, 1M context), making xAI the third independent frontier lab on Bedrock alongside Anthropic and OpenAI

CIOs can now evaluate all three independent US frontier labs under one IAM and billing surface. The caveat is a non-standard endpoint and a context-window pricing cliff above 200K tokens; this is distribution, not a new capability tier.

Sources DigitalApplied; Memeburn

2026-06-15 · Anthropic

Claude Fable 5 and Mythos 5 remained government-suspended all week; the planned Jun 23 usage-credit subscription change is moot while access is off

The top-tier closed model's availability is still a sovereign/regulatory variable, not an SLA. Buyers should keep Opus 4.8/Sonnet fallbacks wired and avoid single-sourcing the frontier for production-critical paths.

Sources Anthropic

Watchlist

July

Gemini 3.5 Pro GA

Pro slipped to July. Its GA and first independent AA v4.1 pass will show whether Google can re-take a frontier lead now contested by both Opus 4.8 and a surging open-weight field.

Jun-Aug

Claude Fable 5 / Mythos 5 restoration

Restoration terms (geo-gating, KYC, or a permanent civilian/government capability split) will set the precedent for sovereign access risk and decide whether Fable 5 re-enters the available scorecard.

Jun-Aug

GLM-5.2 adoption and independent SWE-Bench replication

Downloads, integrations, and third-party benchmark replication will show whether MIT-licensed open weights become production substrate and force closed-flagship price cuts.

Jun-Jul

GPT-5.6 / next OpenAI flagship

Codenames and prediction markets pointed to a launch just after this window. A real system card would re-set the closed frontier and test whether OpenAI answers the open-weight cost pressure.

Changelog

  • Added GLM-5.2 and MiniMax-M3 to the LLM tree; reframed W25 around open weights taking the lead while the closed frontier stalled (Fable 5 suspended, Gemini slipped) and Artificial Analysis rebased to the agentic v4.1 index.