Skip to content

Model layer

The Model Pulse

For architects tracking model capability shifts.

Same list price, three different bills

Big read

A 'model' is now a policy-wrapped system, and its price is a set of meters rather than a list rate: that is what this week's three mid-tier releases show, and it changes how you evaluate and how you budget. Start with the gates, because each lab shipped a different kind. Google's is an access gate: Gemini 4 Argon exists, scores on independent indices, and cannot be bought. Anthropic's is a routing gate: Sonnet 5.5 can hand a higher-risk cyber request to Sonnet 5 mid-session, visibly, through classifiers in the request path. OpenAI's is a classification gate: Sol inherits GPT-6 Astra's safeguards stack with a Critical rating in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment controls), and the planned Astra successor was withheld with no numbers. The September 25 pause covers an unreleased most-capable tier; GPT-6 Astra itself still serves dots and the Ultrafast tier. An evaluation harness (the software that wraps the model, supplies its tools and runs the benchmark) that assumes a fixed model behind an endpoint will produce numbers that do not reproduce in production.

The meters are the second half. For a session that rereads a large context many times, the cache-read meter alone can make the Sol and Sonnet 5.5 bills diverge by 2x behind the same $2 / $10 list price; for a single long-document pass, Sol's long-context repricing does the opposite. Argon's $2 / $10 is a 50% introductory discount for at least one month, and TokenCost noted that its post-intro card is Opus 5.5's card line for line; DataCamp and TokenCost both published the meters-over-list-price read before this issue, and synthorai added the multiple the meter table does not show, that Sonnet 5.5 writes 1.7-3.1x the output tokens of Sol on comparable tasks. The full per-meter table, with the cache-read rates and the 272K threshold, is in the architecture watch below. Above the mid tier on the independent index, Opus 5.5 at $4 / $20 is the only buyable model that outscores Sonnet 5.5, and the question for most agent budgets is whether its two extra index points are worth doubling the bill; GPT-6 Astra at $10 / $50 is buyable and scores below Sonnet 5.5, which is the stronger point for a buyer.

One number to carry: Anthropic's own Terminal-Bench 4.0 harness puts Sonnet 5.5 at 70.6% and Artificial Analysis's independent run (its September 30 Argon article) at 64%, a 6.6-point gap that is the size of the vendor-versus-independent difference to expect in your own evaluation; the independent table, with the effort settings each model was run at, is in the benchmark moves.

What to do: rerun the agent cost model with cache-read rates, long-context thresholds and output tokens per task as inputs, not list price; benchmark Sonnet 5.5 against Sol on your own task set rather than theirs, with the real request mix including the requests the gates are designed to catch; and treat Argon as a Q4 evaluation item, not a Q4 deployment option. The capital and power consequences are in AI Stack Weekly; the permission-layer consequences for agents are in Agent Techniques.

Tree delta

Six rows added, one updated. Three closed frontier-adjacent models at $2 / $10 (Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 4 Argon as gated), two open-weight MoE releases (mixture-of-experts, a model that activates only a few of its sub-networks per token; MiMo-V2.6-Pro now verified, Kolibri-1), and Holo4 under the Holo3.1 precedent (W23 placed H Company's computer-use family as a tree node because it ships in multiple sizes with quantized checkpoints, reduced-precision weights, for local inference). The decision-model fine-tunes are reviewed and deferred rather than placed: the tree's 3-of-6 rule creates a new branch only when at least three of six tests hold (durable architectural distinction, three or more independent vendors, changed infrastructure behaviour, changed user-facing capability, not representable as a tag, and three to five deployed models). This week supplied the vendor count; the architectural-distinction test is the open one, because a classifier head on a shared base may be a fine-tuning recipe rather than a new build.

Registry movement

Gemini 4 Argon carries status gated and placement confidence medium because the public cannot run it. MiMo-V2.6-Pro lifts last week's deferral: the Hugging Face repo now carries MIT license metadata and a model card. Grok 4.7's row now carries a 500K context, which the Amazon Bedrock listing prints citing xAI's September 21 launch announcement; last week's row had no context figure because the house had not recorded one.

Added
claude-sonnet-5-5, gpt-6-1-sol, gemini-4-argon, mimo-v2-6-pro, holo-4, kolibri-1
Updated
grok-4-7

Frontier movements

Anthropic · 2026-09-28 · Frontier · Reasoning

Claude Sonnet 5.5

Artificial Analysis places Sonnet 5.5 (max) at 56, two points behind Opus 5.5 at $2 / $10 against $4 / $20. It leads Terminal-Bench 4.0 on both the vendor harness and Artificial Analysis's independent run, 6.6 points apart (the table is in the benchmark moves). The visible fallback to Sonnet 5 on higher-risk cyber requests, driven by classifiers in the request path, is a behaviour your security team should test, because it changes which model answers mid-session; it is also the week's first-party example of a classifier sitting in a production approval path, which Agent Techniques picks up.

Model registry ID
claude-sonnet-5-5

Sources Anthropic

OpenAI · 2026-09-29 · Frontier · Reasoning

GPT-6.1 Sol

Sol has the cheapest cache meter at the tier and the only long-context surcharge (the meter table is in the architecture watch). OpenAI's DeepSWE v1.1 75.2% and $5.47 per Terminal-Bench Science task are vendor-run, and the cost-per-task advantage over Astra in OpenAI's table follows from Sol's list price being one-fifth of Astra's. Artificial Analysis's independent 52 puts Sol one point below Astra. Sol is classified Critical in cybersecurity (OpenAI's highest internal capability tier, which triggers its strictest deployment safeguards) and ships under Astra's safeguards stack, so expect the same refusal surface as the top tier. It replaced GPT-6 Sol after seven days; pin model ids.

Model registry ID
gpt-6-1-sol

Sources OpenAI

Google DeepMind · 2026-09-30 · Frontier · Reasoning

Gemini 4 Argon

Artificial Analysis scored Argon (high) at 53, equal to Astra, at about 60% of Astra's cost per task, and attributes the gain to two things: lower hallucinations and stronger agentic results (AutomationBench-AA 78%, first). The hallucination half rests partly on abstention. Per AA's September 30 article, Argon's 15% rate on AA-Omniscience (AA's test of whether a model answers wrongly or declines when it does not know) is the lowest of the three models AA compared (GPT-6 Astra 51%, GPT-6.1 Sol 54%; Sonnet 5.5 was not in that comparison and Astra is not a week's release), but its accuracy is 50% against Astra's 63%, so it declines more rather than knowing more. Google's table leads DeepSWE v1.1 at 77.9% (vendor-run) and trails Opus 5.5 on Terminal-Bench 4.0. The $2 / $10 rate is a 50% discount for at least one month, after which the card is $4 / $20. Access is limited to Fairwind Program cyber defenders; developers are 'next' with no date; note the 1M-token output limit for long-document generation when it arrives.

Model registry ID
gemini-4-argon

Sources Google

SpaceXAI · 2026-09-28 · Frontier · Reasoning

Grok 4.7

For AWS-committed enterprises this is a way to run a SpaceXAI model under existing Bedrock IAM, logging and private-link controls. AWS's listing prints the 500K context figure citing xAI's September 21 launch announcement, which the house had not recorded; the price on Bedrock should be checked against the $2 / $6 API rate before assuming parity.

Model registry ID
grok-4-7

Sources AWS Machine Learning Blog

Open weights

Xiaomi · 2026-09-21 · Open Frontier · Moe

MiMo-V2.6-Pro

Last week's deferral lifts: the repo carries MIT metadata and a model card for a 1.02T / 42B-active multimodal MoE ('active' is the parameters used per token). Artificial Analysis lists it at 46, ahead of GLM-5.3 (45) and Kimi K3 (44) and tied with Grok 4.7; its cost and minutes per task are in the benchmark moves. For a buyer the number to weigh is the latency: the cheapest open frontier model is also one of the slowest, which decides whether it fits interactive or batch work.

Model registry ID
mimo-v2-6-pro

Sources Hugging Face model card

Aleph Alpha · 2026-10-03 · Open Frontier · Moe

Kolibri-1

The fully open training recipe (20T pretraining tokens in 21 days, about 392k GPU-hours, roughly 6.4e23 FLOPs) is the useful artifact for anyone budgeting a sovereign model; the benchmark table (AIME 2025 96.9%, SWE-Bench Verified 66.4%) is vendor-run with no independent evaluation. At about 78 GB in FP8 (8-bit floating-point weights, half the memory of 16-bit) it fits a single H200 or B200; among European open-weight releases it is the first the house has recorded that combines a single-GPU footprint with a 1M context, though Mistral has shipped Apache 2.0 models with on-prem footprints at shorter contexts.

Model registry ID
kolibri-1

Sources Aleph Alpha

H Company · 2026-09-28 · Specialist · Agentic

Holo4

H Company's own figures are 45.4% on AutomationBench (business-workflow automation in simulated apps) at $0.05 per task for the 27B, read from its chart, and 61.7% on OSWorld 2.0 (desktop control) at $1.22 per task against Opus 5.5's 81.8%. The per-task costs are cheap in absolute terms; no frontier cost per AutomationBench or OSWorld task exists in the sources, so there is no frontier cost comparison to draw, and the vendor tables are from different versions. CC BY-NC 4.0 is more restrictive than the Apache 2.0 Qwen base, so enterprises need a commercial license from H Company before production use. Publishing all trajectories is the right precedent and lets you audit the score.

Model registry ID
holo-4

Sources Hugging Face blog (H Company)

Cloudflare · 2026-10-01 · Specialist · Dense

Clef

Clef returns calibrated probabilities over 1-64 typed questions in one forward pass, which is the shape an agent's allow/flag/block gate needs. Cloudflare's benchmark table (BANKING77 94.20 macro-F1, the average of per-class accuracy scores, against Jev's 79.74) is vendor-run, and the one community head-to-head posted on Hacker News, a 250-sample test, found Clef 5.2x more expensive and roughly five times slower than Jev for a 1.2-point gain. Read it as a reviewer model you can self-host under Apache 2.0, not as a leaderboard result; the technique read and the decision-model price list are in Agent Techniques.

Sources Cloudflare Blog

Architecture watch

List-price convergence with meter divergence

Three labs chose the same headline number and set every other meter independently. This is the fact home for the mid-tier meter packet. Cache reads (re-sending context the provider already holds): GPT-6.1 Sol $0.10 per million, Sonnet 5.5 $0.20, Argon 95% off list. Long-context repricing: Sol moves to $4 / $15 above 272K tokens; Sonnet 5.5 has no surcharge across its 1M context. Introductory pricing: Argon's $2 / $10 is a 50% discount for at least one month, reverting to $4 / $20, and nobody outside the Fairwind Program can pay either rate yet. As a heuristic rather than a measurement (the house has no token-mix data for production agents), the cache-read rate is the meter most likely to dominate agent sessions, which tend to reread context more than they emit tokens; the long-context threshold catches document and codebase workloads, and only Sol has it; the introductory reversion catches budget planning, and only Argon has it. A router that keys on list price will pick the wrong model for most agent workloads; key it on cache-read rate, context threshold and effective price after the intro period ends, and measure your own token mix to replace the heuristic.

Examples
Claude Sonnet 5.5 ($2 / $10, cache $0.20, no long-context tier), GPT-6.1 Sol ($2 / $10, cache $0.10, $4 / $15 above 272K), Gemini 4 Argon ($2 / $10 intro, $4 / $20 list, cached 95% off)

Sources OpenAI API pricing; Anthropic pricing; Google Gemini 4 Argon announcement

Task heads on a shared 27B base

Four releases in one week post-trained the same Qwen3.8-27B checkpoint into a specialist: three decision heads and one computer-use agent. The pattern is that a 27B dense open base has become the default substrate for task-specific heads, in the way BERT-base once was for classifiers, and that the differentiation is in the head, the loss and the data rather than the trunk. For buyers this means provenance and training-data disclosure matter more than parameter counts, and the Perplexity case, where the published shards match a community checkpoint byte for byte, shows how thin the provenance can be. For the tree it raises the question of whether decision models earn a branch, which the placement rules answer with a 3-of-6 independent-vendor test that this week put in play.

Examples
Cloudflare Clef (Qwen3.8-27B), Perplexity pplx-decider-v1-27b (Qwen3.8-27B), AutoTrust JEV-27B (Qwen3.8-27B), Holo4-27B (Qwen3.8-27B)

Sources Cloudflare, Perplexity, AutoTrust and H Company model cards on Hugging Face

Gating as a release stage

Each of the three labs shipped or announced a model this week with a safety gate that changes what a buyer actually gets. Google's is an access gate: the model exists, scores on independent indices, and is unavailable. Anthropic's is a routing gate: the model you called may hand the request to its predecessor mid-session, visibly, on the decision of reasoning-extraction classifiers in the request path. OpenAI's is a classification gate: Sol inherits Astra's Critical-cyber safeguards, and the planned Astra successor was withheld entirely. The architectural consequence is that 'model' now denotes a policy-wrapped system whose behaviour depends on request content, and evaluation harnesses that assume a fixed model behind an endpoint will produce numbers that do not reproduce in production. Test with your real request mix, including the requests the gate is designed to catch. The same architecture, a classifier deciding what the acting model may do, is the week's technique in Agent Techniques.

Examples
Gemini 4 Argon (Fairwind Program only), Claude Sonnet 5.5 (visible fallback to Sonnet 5 on higher-risk cyber tasks), GPT-6.1 Sol (Critical in cyber, Astra's safeguards stack), GPT-6.1 Astra (withheld)

Sources Google Gemini 4 Argon announcement; Claude Sonnet 5.5 System Card; OpenAI Deployment Safety Hub

Verbosity and latency as hidden costs of open frontier scores

The top open-weights model on the independent index is also one of the slowest: Artificial Analysis clocks MiMo-V2.6-Pro at 19.5 minutes per Intelligence Index task at $0.13, roughly a fifteenth of GLM-5.3's $2.01 per task for one more index point. AA's cost per task is price times tokens emitted, so a model that scores by generating long reasoning traces can be cheap per token and slow per task at the same time, and the comparison that matters is cost and minutes per task, not list price. Kolibri-1's 3.46B active on 78B total is the extreme case of a small active footprint, with no independent latency or tokens-per-task figure yet. An architect choosing an open model for interactive work should weight tokens per second and tokens per task as heavily as the index score; for batch work the index-per-dollar figure is the right one.

Examples
MiMo-V2.6-Pro (46 on the index, 19.5 minutes per task, $0.13 per task), GLM-5.3 (45 on the index, $2.01 per task), Kolibri-1 (3.46B active, vendor-run only)

Sources Artificial Analysis model pages

Benchmark moves

Artificial Analysis Intelligence Index v4.3.2

Four new entries in one week; Opus 5.5 holds first, Sonnet 5.5 enters second, Argon ties Astra, MiMo-V2.6-Pro confirmed as top open weights. The setting in parentheses is the reasoning-effort level the model was run at; scores are not comparable across settings

Claude Opus 5.5 (max)
58
Claude Sonnet 5.5 (max)
56
Gemini 4 Argon (high)
53
GPT-6 Astra (max)
53
GPT-6.1 Sol (max)
52

Sources Artificial Analysis

Artificial Analysis Intelligence Index v4.3.2, open weights

MiMo-V2.6-Pro confirmed first among open-weights models, tying Grok 4.7 at a fraction of the cost per task; this table is the fact home for the open-weights cost and latency figures

MiMo-V2.6-Pro
46 ($0.13 per task, 19.5 minutes per task, both read from AA's chart)
GLM-5.3 (max)
45 ($2.01 per task, read from AA's chart)
Kimi K3 (max)
44
Grok 4.7 (closed, for reference)
46 ($2.73 per task, read from AA's chart)

Sources Artificial Analysis

Terminal-Bench 4.0 (independent Artificial Analysis runs, with the vendor figure for comparison)

Artificial Analysis's independent run puts Sonnet 5.5 first at 64%, 6.6 points below Anthropic's own 70.6%; Opus 5.5, Astra and Argon follow within seven points. The AA figures are read from AA's September 30 Argon article and sit outside the week's graded research set; the vendor figures are graded

Claude Sonnet 5.5 (max) (Artificial Analysis, independent)
64%
Claude Opus 5.5 (max) (Artificial Analysis, independent)
60%
GPT-6 Astra (Artificial Analysis, independent)
59%
Gemini 4 Argon (Artificial Analysis, independent; Google's own run 57.4%)
57%
Claude Sonnet 5.5 (Anthropic, vendor-run, for comparison)
70.6%

Sources Anthropic; Google DeepMind; Artificial Analysis

DeepSWE v1.1 (vendor-run tables)

Argon claims the top score; Sol matches Astra at one-fifth the price; every figure is from the vendor's own table

Gemini 4 Argon (Google)
77.9%
GPT-6.1 Sol at high (OpenAI)
75.2%
GPT-6 Astra (OpenAI / Google tables)
~74.8% / 74.1%
Claude Opus 5.5 (Google table)
74.2%
GPT-6 Sol best (OpenAI)
68.8%

Sources OpenAI; Google DeepMind

AutomationBench (vendor-run, mixed versions)

Argon leads on Google's table, 8.8 points ahead of Opus 5.5 on the same table; Holo4-27B reaches 45.4% at $0.05 per task on H Company's chart; Sol's 31.7% on OpenAI's table is a different version and is not comparable with Google's rows

Gemini 4 Argon (Google table)
51.3%
Holo4-27B (H Company, v1.0.6, read from chart)
45.4% at $0.05 per task
Claude Opus 5.5 (Google table)
42.5%
Holo4-35B-A3B (H Company, v1.0.6, read from chart)
34.5% at $0.02 per task
GPT-6.1 Sol at medium (OpenAI table; OpenAI states a 2.2-point lead over Opus 5.5 on its own run, which is not the Google figure)
31.7%

Sources Google DeepMind; H Company; OpenAI

LMArena text leaderboards (published Sep 30 and Oct 2)

Argon debuts first on Hard Prompts (preliminary) and Coding; MiMo-V2.6-Pro enters Coding at 1540

Gemini 4 Argon, Coding
1560 ± 17
Gemini 4 Argon, Hard Prompts (preliminary, 3,166 votes)
1551 ± 11
MiMo-V2.6-Pro, Coding
1540 ± 18

Sources LMArena

Tier scorecard

As of 2026-10-03

TierLeaderChallengerRead
Closed frontierClaude Opus 5.5GPT-6 Astra (tied on the index by Gemini 4 Argon, which cannot be bought)Opus 5.5 leads the independent index at 58, five points clear of Astra and Argon at 53. Fable 5.1 held the slot through W39 on buyer-trace grounds (the house kept it as the default until a buyer-side trace showed Opus 5.5 matching it on real work); no such trace has arrived, and the house has no independent index score for Fable 5.1 in this window, so the handover rests on Opus 5.5's index lead and on Anthropic positioning Opus as its flagship, not on a measured Fable gap. Argon's tie with Astra is real and unpurchasable.
Open frontierMiMo-V2.6-ProGLM-5.3Verified weights under MIT and an independent 46 make MiMo the open leader; GLM-5.3 at 45 is the challenger. DeepSeek-V4.1-Flash gives up the slot it held on serving cost; Kolibri-1 enters the watch column pending any independent score.
ReasoningClaude Opus 5.5Claude Sonnet 5.5Two points separate them on the independent index at a 2x price difference. Sonnet 5.5 is the default for most reasoning workloads this quarter; Opus 5.5 is the ceiling. GPT-6.1 Sol at 52 is the cross-vendor alternative with the cheaper cache meter.
CodingClaude Sonnet 5.5GPT-6.1 SolSonnet 5.5 leads Terminal-Bench 4.0 on both the vendor and the independent run (table above) and takes the slot from Opus 5.5 at half the price; Sol replaces GPT-6 Astra as challenger because it matches Astra on OpenAI's DeepSWE table at one-fifth the cost. Argon's 77.9% DeepSWE is the highest claim and the least testable.
MultimodalGemini 3.8 Live Extended ThinkingGemini 4 ArgonThe buyable Google multimodal model holds the slot; Argon's 1M-token output and LVBench 91.7% (vendor-run) would take it when developers can call it. MiMo-V2.6-Pro's native text, image, video and audio input makes it the open alternative.
Edge / small4B decision heads via llama.cpp (Kev-4B and peers)Clef-flash (9B)This tier covers models meant to run on one device or one CPU, judged on a shipped task they do well rather than on the index. The day-0 decision heads, from 144 million parameters to 4B, answering allow/flag/block in 3-43 ms per question (author-reported; 3 ms is Julia-1 at 144 million parameters, and the 4B heads are 12-36 ms) through llama.cpp's typed endpoint take the slot from last week's leader, the vendor claim that a 30B mixture-of-experts model runs locally on the Snapdragon 8 Elite Extreme Gen 6 handset, which still has no independent throughput figure; Clef-flash at 9B under Apache 2.0 replaces Nex-N2.5 mini as challenger. Holo4-27B at $0.05 per AutomationBench task (vendor-run) is noted, but a 27B model served through an API is not an edge footprint.

Vendor signals

2026-09-28 · OpenAI

Shelves GPT-6.1 Astra after internal alignment tests; the pause on an unreleased most-capable tier remains; publishes a 'towards safety cases' post with no numbers

GPT-6 Astra remains in production behind dots and the Ultrafast tier; what OpenAI withheld is its successor, and what it paused is a tier above Astra that never shipped. Architects should keep the cross-vendor fallback for top-tier tool use through Q4 because the roadmap above Astra is now empty. The post's absence of quantified results means the external evidence for the decision is the UK AI Security Institute's evaluation of GPT-6 Astra, which is about the model that stayed; the three-model table and its caveats are in Agent Techniques.

Sources OpenAI

2026-09-29 · OpenAI

Adds an Ultrafast service tier for GPT-6 Astra at $60 / $300 per million (6x) for up to 8x token speed in Codex and 6x in the API, per OpenAI; introduces Pro 500 at $500 a month; halves the Pro 200 Codex allowance from October 30

This is the first retail price increase on a frontier coding tier since the mid-tier convergence, and it prices speed as a premium product rather than a model property. Teams that depend on Codex throughput should model the move to Pro 500 or Sol Ultrafast against moving the workload to Sonnet 5.5 before October 30.

Sources OpenAI API pricing

2026-09-30 · Google DeepMind

Announces Gemini 4 Argon with access limited to Fairwind Program cyber defenders; publishes introductory pricing with no end date and no developer availability date

Google chose to claim the frontier on independent indices before letting anyone buy the model. For a buyer the signal is that Google's Q4 roadmap is public and its Q4 product is not. The reversion to $4 / $20 after the introductory month means the $2 / $10 figure in comparison pieces is a promotional rate, not Google's price.

Sources Google

2026-09-28 · Anthropic

Ships Sonnet 5.5 to all platforms including free Claude.ai; sets retirement no sooner than September 28, 2027; Haiku 5.5 slips to 'coming weeks' for a second week

A one-year retirement floor is a deployment guarantee most labs do not give and should go into vendor-selection scorecards. The Haiku slip matters for the cheap tier: there is still no Anthropic model below $2 / $10 in the 5.5 family, which leaves the sub-dollar tier to GPT-6 Luna, Gemini Flash and the open models.

Sources Claude Platform docs

2026-10-01 · Perplexity

Launches a Decisions API at $0.04 per million input tokens with free output, backed by pplx-decider-v1-27b; community analysis finds the weights byte-identical to a 12-day-old community checkpoint that the launch post does not mention

The price is the lowest in the decision-model category and the provenance is the weakest. Enterprises evaluating per-action review should treat the vendor benchmark table as describing someone else's model until Perplexity explains the match in its launch materials, and should prefer a vendor whose training provenance is documented.

Sources Hugging Face (perplexity-ai/pplx-decider-v1-27b); ModelSystem.One

2026-10-01 · Black Forest Labs

Launches FLUX 3 Image via API and a commercial-weights license, with open weights promised 'in the coming weeks'

The open-weights promise is a calendar item, not a release; FLUX 3 is tracked here only because its weights, when published, would be the first open image model of the generation. Until then it is an API product.

Sources Black Forest Labs

Watchlist

October

Claude Haiku 5.5 model id on the Claude Platform models page

Two weeks of 'coming weeks'. Haiku 5.5 would be Anthropic's first sub-$2 model with the 5.5 safeguards and the obvious reviewer model for per-action gating.

Q4 2026

Gemini 4 Argon developer availability and the introductory pricing end date

Independent benchmarks on a model nobody can run are a preview, not a result. Developer access resets the top-tier comparison and the $4 / $20 reversion tests the mid-tier convergence.

October

A second independent Terminal-Bench 4.0 harness on Claude Sonnet 5.5

Anthropic's 70.6% and Artificial Analysis's 64% are 6.6 points apart on the number most likely to move coding-agent defaults this quarter. A second independent harness would show whether the gap is AA's setup or the vendor's.

October

Perplexity response on pplx-decider-v1-27b provenance

Byte-identical shards to a community checkpoint, unmentioned in the launch post, is either an oversight or a disclosure problem. Which one it is decides whether the Decisions API belongs in an enterprise evaluation.

By Oct 30

Nvidia InferenceX submission for Vera Rubin NVL72

The lapsed Q3 commitment and its consequences are tracked in AI Stack Weekly; a Rubin submission would give the first independent cost-per-million figures for serving this week's models.

Q4 2026

First independent evaluation of Kolibri-1 and FLUX 3 open weights

Kolibri-1 pairs a single-GPU footprint with a 1M context, a combination no other European open-weight release in the house record has; its table is vendor-run. FLUX 3's weights decide whether the open image tier moves this generation.

Changelog

  • Authored for the September 28 to October 4, 2026 window from vendor launch posts, model cards and Artificial Analysis and LMArena publications.
  • Six tree rows added and one updated; the decision-model fine-tunes are reviewed and deferred pending the placement rules' 3-of-6 test rather than placed.
  • Scorecard: Opus 5.5 takes the closed-frontier and reasoning leader slots from Fable 5.1 on the independent index; MiMo-V2.6-Pro takes the open leader slot from DeepSeek-V4.1-Flash on verified weights and an independent score; the Open frontier challenger moves from Atria Dawn Preview to GLM-5.3; the Coding leader and challenger move from Opus 5.5 and GPT-6 Astra to Sonnet 5.5 and GPT-6.1 Sol; the Edge / small leader moves from the Snapdragon 8 Elite Extreme on-device claim to the 4B llama.cpp decision heads and the challenger from Nex-N2.5 mini to Clef-flash; the Multimodal challenger moves from Gemini 3.8 Flash TTS to Gemini 4 Argon.
  • Every benchmark figure is labelled vendor-run or independent at the point of use; vendor tables from different labs are not compared against each other except where flagged as such.
  • Revision cycle 1 (editorial board): the Terminal-Bench 4.0 table now carries Artificial Analysis's independent runs for Sonnet 5.5 (64%), Opus 5.5 (60%), GPT-6 Astra (59%) and Argon (57%), which the first draft missed; Argon's hallucination figure is corrected from highest to lowest in the set, achieved by abstention; Argon's tier is corrected to frontier; the claim that OpenAI attributes Sol's cost advantage to cache reads is removed; the Opus 5.5 AutomationBench row is labelled derived; the bigRead is restructured around the three meters.
  • Revision cycle 2 (editorial board): the Opus 5.5 AutomationBench row now carries Google's printed 42.5% instead of a derived figure and the Sol row is marked non-comparable; Argon's Omniscience comparison is restricted to the models in Artificial Analysis's September 30 article; the AA Terminal-Bench and Omniscience figures are labelled as outside the week's graded research set at the point of use (the fact audit verified them against the article; research.json still needs rows for them); the NOTICE-file provenance clause is removed as unverified; the Grok 4.7 context note no longer says 'undisclosed'; the Holo4 cost claim is narrowed to absolute per-task cost with no frontier comparison; the bigRead leads with the gating thesis, credits DataCamp, TokenCost and synthorai for the meter frame, and drops the 'house contribution' line; the mid-tier meter packet's fact home is the first architecture-watch entry and the AISI table's fact home is Agent Techniques; the Edge / small tier is defined and re-ranked; the Coding and Open frontier changes are recorded in the scorecard changelog line.