This was the heaviest release week of the year, and the interesting thing about it is not the capability. Meta shipped Muse Glimmer under Apache 2.0 on Monday, OpenAI shipped GPT-5.6-Cyber behind a new access tier the same day, SpaceXAI shipped Grok 4.6 on Wednesday, DeepSeek swapped its flagship endpoint to the 0813 build on Wednesday without saying so, Google shipped Gemini 3.7 Flash on Thursday at half the prior price, and Z.ai announced GLM-5.3 on Friday. Six models. Four vendors claiming some version of openness. And in three of the six cases, the artifact a buyer can actually obtain is not the artifact anyone measured.
DeepSeek is the cleanest example. On August 12 the version string on its pricing page changed from the April preview to DeepSeek-V4-Pro-0813, and the vendor-reported gains are enormous: Terminal Bench 2.1 from 72.1 to 87.9, DeepSWE from 12.8 to 62.7, CyberGym from 52.7 to 83.3, the Artificial Analysis Intelligence Index from 45 to 53. There was no blog post. The Hugging Face repositories continued to host the April preview. For a stretch this week, every independent evaluation of 'DeepSeek V4 Pro' was measuring a build nobody could download, and every downloadable build was one nobody was serving. Z.ai did a softer version of the same thing: GLM-5.3 was announced with a claimed best-in-open-source Terminal Bench 3.0 result and a CyberGym score ahead of Claude Mythos 5, reachable at announcement only through a paid coding subscription, with weights promised on Hugging Face 'within two weeks.'
This publication has recorded both models as gated rather than open-weight in the tree this week. That is not pedantry. Last week's issue documented Artificial Analysis measuring identical open weights losing up to 40% of reference accuracy depending on which endpoint served them. Stack the two findings and the procurement problem is concrete: the score belongs to a specific build served in a specific configuration, and a buyer downloading weights under the same product name is acquiring neither. The gap used to be closed-versus-open. The gap that matters now is served-versus-downloadable, and it opens and closes on a vendor's release schedule rather than on a license.
Meta went the other direction and it is worth being precise about what it did. Muse Glimmer is a 30B dense multimodal model under Apache 2.0 — no 700-million-user gate, no naming rule, no acceptable use policy, and no mechanism for Meta to withdraw the grant. Artificial Analysis puts it at 35 on the Intelligence Index and 44 on the Openness Index. But Meta published weights and withheld training data and training code, and it is not serving the model on its own API, so every price and latency number a buyer sees for Glimmer comes from a third party. Meta gave away the terms and kept the frontier: Muse Spark 1.2 stays closed, with its weights promised later.
The pricing story runs the same shape. Headline rates fell — Gemini 3.7 Flash at $0.75 / $3.75 is half of 3.6 Flash, and Grok 4.6 ties GPT-5.6 Sol Max on the Intelligence Index at $2 / $6 against Sol's $5 / $30. But Google's own pricing page says the Gemini rate is introductory through December 31 and that $1.50 / $7.50 applies from January 1, and DeepSeek's flat $0.435 / $0.87 gives way on August 16 to reported peak and off-peak tiers of $1.32 / $3.96 and half that off-peak. Both increases were disclosed in advance, in writing, by the vendor. Anyone modelling agent unit economics on this week's rate card is modelling a promotional period with a published expiry date, and the expiries are inside the planning horizon for anything being procured this quarter.
The one number that survives all of this is Grok 4.6's turn count. Artificial Analysis measured roughly 53 turns and 0.5B input tokens to resolve long-horizon agentic tasks against roughly 103 turns and 2.0B tokens for Claude Opus 5 at maximum settings. Turn efficiency is a property of the model that does not reprice on August 16 and does not depend on which endpoint serves it. For anyone running agents at scale, that is the more durable procurement input than the rate card — and it is the number the fewest buyers are tracking.