Two things happened in the software layer this week that only make sense read together. On Thursday Ornith AI, a lab nobody had heard of a fortnight ago, published Ornith-1.5 under the MIT license — a 397B mixture-of-experts flagship with 35B active parameters, a 35B dense sibling, and a 9B distilled variant, benchmarked on the lab's own harness at 89.7% on Terminal-Bench 2.1 and 61.4% on DeepSWE. The next day OpenAI announced a token price cut of more than 20% on GPT-5.6 Sol, held for three months. One vendor said the top of the coding leaderboard is now open, released under a permissive license, and available to anyone. The other vendor said the current price of frontier reasoning tokens is negotiable and it is willing to say so out loud, for a fixed period, in writing.
The honest reading is that neither the leaderboard claim nor the price claim is settled. Ornith AI's Terminal-Bench figure comes from the lab's own harness, which has not been audited by any third party, and no independent tracker — Artificial Analysis, Vals, LMSYS — had ranked Ornith-1.5 by publish time. This publication is adding the model to the tree because the artifact is real and downloadable, and it is recording it at medium placement confidence with a notable field saying so plainly. If Artificial Analysis or Vals reproduces the score inside the next four weeks, the classification is trivial to lift. If they do not, that itself is the finding, and it belongs in a procurement document rather than a headline.
OpenAI's move is the more consequential one for buyers this quarter. A promotional 20%-plus cut on the flagship reasoning model, published as a dated three-month change, is the same behaviour Google put on the Gemini 3.7 Flash pricing page last week: cheap inference sold with an expiry. This publication's precision rule is that a price with a published end date is a promotional variable, not the durable rate — the durable rate is whatever the vendor's willing to write on a contract that spans the reversal. Every agent business case whose payback crosses November should model the post-promo price, and every vendor cost comparison built on this week's per-token headline should carry the date attached.
A third release rounds out the layer without changing it. DeepSeek published DeepSeek-V4-Flash-Vision-Exp on Friday, an experimental API-only vision-capable variant of its Flash tier with no weights and no full-quality Pro comparator, aimed at multimodal harness developers testing the API surface. This publication is not adding it to the tree — models.yaml does not have a pattern for experimental API-only rows, and creating one for a variant with no weights would set a precedent the taxonomy would then have to defend against every future API-only preview. The Pulse notes it and moves on.
The absent number is the Artificial Analysis Intelligence Index. No new AA Index release landed in-window, so Grok 4.6, GPT-5.6 Sol Max and Claude Opus 5 are still separated by two points of a measurement last refreshed on August 6. Coverage will keep quoting positions on the board this week; none of them changed. And CyberGym remains where GLM-5.3 and DeepSeek-V4 Pro 0813 left it last week — vendor-reported, unreproduced, and cited interchangeably in trade press.
One finding on the served-versus-downloadable split from last week worth pulling forward: Z.ai has still not published GLM-5.3 weights to Hugging Face. The two-week window announced with the model closes at the end of this issue's coverage period. If nothing lands in the next seven days, the benchmark-first-weights-later pattern this publication flagged in W33 gets one more instance and moves toward being a durable vendor behaviour rather than a one-off. The tree keeps GLM-5.3 recorded as gated until weights are confirmed.