For architects tracking model capability shifts.
No new closed frontier shipped; the open/local agent substrate widened underneath it.
Week 23 of 2026 · June 6, 2026
Big read
W23 did not produce the expected Gemini 3.5 Pro GA or a fresh Anthropic/OpenAI frontier release. That absence is the story: Claude Opus 4.8 remains the public closed-frontier leader for coding and agentic work, while Google kept Pro in the June watch window and the model layer's actual shipping activity moved down-stack. JetBrains released Mellum2, an Apache-2.0 12B/2.5B-active MoE designed for low-latency routing, RAG, summarization, validation, and sub-agent calls; NVIDIA released Cosmos 3 as an open physical-AI omni-model with Nano 16B and Super 64B variants; H Company released Holo3.1 with local computer-use sizes and quantized checkpoints. The procurement implication is sharper than another leaderboard reshuffle: production agent systems are becoming portfolios of models. Keep Opus/GPT/Gemini-class models for high-risk reasoning and codebase-scale orchestration, but push cheap, private, repeated sub-agent work into specialized open/local models. The tree delta therefore adds efficient-agent and physical-AI nodes rather than another general chatbot crown.
Tree delta
Three W23 additions: Mellum2 for efficient text/code sub-agent workloads, Cosmos 3 for physical-AI omni-modeling, and Holo3.1 for local computer-use agents.
Registry movement
Gemini 3.5 Pro remains pending; no tree row is added until Google publishes the GA model card or API identifier.
- Added
- mellum2, cosmos-3-nano, holo-3-1
- Updated
- None
Frontier movements
Google DeepMind · 2026-06 target · Frontier · Reasoning
Gemini 3.5 Pro
This is a frontier movement by absence. Buyers waiting for Gemini 3.5 Pro should keep the launch on the June watchlist but should not pause current coding-agent baselines: the public board still has Opus 4.8 leading the closed frontier, and Pro's economics and benchmark profile remain unverified. The likely enterprise split is task routing — Gemini for huge-context/multimodal work, Opus/GPT for coding and agentic reliability — not a universal replacement.
Sources Google Gemini 3.5 announcement; AI Tool Bolt June comparison
Open weights
JetBrains · 2026-06-01 · Edge · Moe
Mellum2
Mellum2 is not trying to win the frontier leaderboard; it is trying to lower the cost of the thousands of routine model calls inside agent systems. Routing, RAG, summarization, validation, and lightweight code tasks are exactly where private deployments want an efficient open model. If the benchmark claims hold, this is a practical procurement node for teams trying to cut orchestration cost without sending every step to a closed flagship.
- Model registry ID
- mellum2
NVIDIA · 2026-06-01 · Specialist · Multimodal
NVIDIA Cosmos 3
Cosmos 3 widens the model tree away from text-only agents and into physical AI. The important shift is a unified model that can reason over world state and action generation rather than stitching separate generation and control pipelines together. Robotics, simulation, and industrial automation teams should evaluate it as synthetic-data and reasoning infrastructure, not as a chatbot substitute.
- Model registry ID
- cosmos-3-nano
H Company · 2026-06-02 · Specialist · Agentic
Holo3.1
Holo3.1 pushes computer-use agents toward deployability: multiple sizes, quantized checkpoints, and local inference targets matter more than a single headline score. Enterprises automating browser, desktop, and internal-tool workflows can now separate privacy-sensitive UI action from hosted frontier reasoning. That supports a two-layer architecture: local CUA for execution, closed frontier for planning and verification.
- Model registry ID
- holo-3-1
Sources Hugging Face Holo3.1 launch
Architecture watch
Cheap specialist sub-agents
The agent stack is splitting into high-reasoning planners and cheap repeated workers. Mellum2 and Holo3.1 are purpose-built for the calls that happen hundreds or thousands of times inside a workflow: routing, validation, summarization, UI action, and local execution. Model routers should now budget by step type rather than treating one flagship as the default for every agent call.
- Examples
- Mellum2, Holo3.1-0.8B / 4B / 9B, Claude Opus 4.8 fast mode
Sources Hugging Face Mellum2 and Holo3.1 launches
Physical-AI omni-models
Open model activity is expanding from language and code into physical-world simulation, video rendering, and action generation. Cosmos 3's combined world generation, physical reasoning, and action generation points toward a branch where synthetic data and robotics workflows become first-class model workloads. That is a different buyer and deployment path than enterprise chat.
- Examples
- NVIDIA Cosmos 3 Nano, NVIDIA Cosmos 3 Super, Bernini-R renderer
Sources Hugging Face NVIDIA Cosmos 3 launch; ByteDance Bernini-R Hugging Face card
Frontier release gaps measured in weeks
The absence of a new frontier release this week matters because expectations have compressed. Buyers are now tempted to delay procurement for a model that may arrive in days. The practical answer is to separate infrastructure choices from model choice: standardize evaluation harnesses, routers, and cost controls so a June GA can be tested and slotted without freezing current deployments.
- Examples
- Claude Opus 4.8, Gemini 3.5 Pro pending, GPT-5.6 speculation
Sources Google Gemini 3.5 announcement; Anthropic Opus 4.8 announcement
Benchmark moves
Artificial Analysis Intelligence Index
No new W23 leaderboard reset; Claude Opus 4.8 remains the public #1 at 61.4 while Gemini 3.5 Pro is still pending GA
- Claude Opus 4.8
- 61.4
- GPT-5.5
- ~60
- Gemini 3.1 Pro
- ~57
Sources Artificial Analysis summaries via June model comparisons
SWE-Bench Pro
Closed frontier still leads coding: Opus 4.8's 69.2% remains the public bar; no new open release in W23 changes the top coding score
- Claude Opus 4.8
- 69.2%
- GPT-5.5
- ~66-67%
- Gemini 3.1 Pro
- ~62%
Local computer-use deployability
Holo3.1 shifts the measurable axis from one top-line CUA score to model size and quantization availability for local execution
- Holo3.1-0.8B
- ultra-light local
- Holo3.1-9B
- balanced local
- Holo3.1-35B-A3B
- state-of-the-art tier
Sources Hugging Face Holo3.1 launch
Tier scorecard
As of 2026-06-06
| Tier | Leader | Challenger | Read |
|---|---|---|---|
| Closed frontier | Claude Opus 4.8 | GPT-5.5 | No W23 reset; Gemini 3.5 Pro remains the watched June challenger rather than a published benchmark row. |
| Open frontier | DeepSeek V4-Pro | GLM-5.1 | No new open frontier text model displaced the April leaders; W23 open activity shifted to specialist/local models. |
| Reasoning | Claude Opus 4.8 | GPT-5.5 | Closed reasoning leadership steady while Gemini 3.5 Pro remains pending. |
| Coding | Claude Opus 4.8 | GPT-5.5 | Opus 4.8 still owns the visible SWE-Bench Pro lead; Mellum2 matters for cheap sub-agent code/text calls. |
| Multimodal | Gemini 3.1 Pro | Cosmos 3 | Gemini remains the general multimodal reference; Cosmos 3 creates a specialist physical-AI branch. |
| Edge / small | Mellum2 | Holo3.1-9B | Efficient local/sub-agent models were the week's real release activity. |
Vendor signals
2026-06 · Google DeepMind
Gemini 3.5 Pro remains promised for June with no public API model ID, pricing, or third-party benchmark row by W23 close
Procurement teams should prepare an eval slot but avoid freezing current deployments on an unreleased model. The operating pattern is rapid re-baselining, not launch-date speculation.
Sources Google Gemini 3.5 announcement and June developer guidance
2026-06-01 · JetBrains
Released Mellum2 under Apache 2.0 with an explicit low-latency production-workload positioning
IDE and enterprise-platform vendors can now point to a plausible open/private model for background text-code tasks. That increases pressure on hosted copilots to justify every closed-frontier call by risk or quality, not habit.
2026-06-01 · NVIDIA
Released Cosmos 3 on Hugging Face while also announcing Vera Rubin production at GTC Taipei
NVIDIA is binding the model and infrastructure stories together: physical-AI models create demand for the simulation, synthetic-data, and rack-scale compute stack it sells. Buyers should evaluate model capability and deployment substrate together.
Watchlist
Jun 7-30
Gemini 3.5 Pro GA
The first public model card, price row, API ID, and Artificial Analysis pass will determine whether June becomes a true frontier reset or just a routing expansion.
Jun-Jul
Mythos-class Anthropic availability
Anthropic has publicly framed stronger gated models as pending cyber safeguards. A wider release would change the closed-frontier scorecard more than another Opus point release.
Jun-Aug
Open/local agent model adoption
Downloads, integrations, and benchmark replications for Mellum2, Cosmos 3, and Holo3.1 will show whether specialist open models are becoming production substrate or just launch-week noise.
Changelog
- Added Mellum2, Cosmos 3 Nano, and Holo3.1 to the LLM tree and reframed W23 around efficient/local specialist models rather than a closed-frontier release.