Skip to content

Model layer

The Model Pulse

For architects tracking model capability shifts.

Model governance now starts at the deployable package: entitlement, retention, runtime version, and price window

Big read

The differentiated procurement move this week is not the already-circulating claim that the harness matters. It is that a model stock-keeping unit (SKU) now includes administrator enablement, retention route, provider-adapter version, reasoning effort, and promotion expiry. Claude Fable 5.1 is broadly available in GitHub Copilot but administrator-enabled for enterprises and follows a documented default-retention route outside approved zero-data-retention arrangements. Gemini 3.8 Flash enters the same surface under introductory provider pricing, while Astra enterprise access is off by default.

ARC Prize supplies the evidence for why runtime version belongs in that record: its controlled maximum-reasoning comparison produced a 35.9-point harness-associated spread. The best-observed adapter spread is larger but changes reasoning effort, so it is descriptive rather than causal. This launch-week finding belongs to ARC Prize and prior coverage; the operational extension here is to version the whole deployable SKU and rerun it without provider-private state before claiming portability.

Open weights were quiet after W35's five-model wave. NVIDIA's Apache-2.0 Muse-Glimmer-30B low-precision checkpoint reduced storage substantially, but it is a quantization update to an existing model, not evidence that the open frontier closed the gap this week. The model tree adds Astra, Fable 5.1, and Gemini 3.8 Flash without manufacturing a new open-frontier leader from a footprint optimization.

Tree delta

3 model rows added: GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash. No existing model row changed; Muse-Glimmer's low-precision checkpoint is an update, not a new lineage node.

Registry movement

W36's delta is closed and distribution-heavy. The durable change is that runtime state, administrator entitlement, retention, and promotional pricing now sit beside architecture in model diligence.

Added
gpt-6-astra, claude-fable-5-1, gemini-3-8-flash
Updated
None

Frontier movements

OpenAI · 2026-09-03 · Frontier · Agentic

GPT-6 Astra

Evaluate Astra as a model-plus-runtime release. ARC Prize's matched-effort comparison associates a 35.9-point spread with the harness; portability reviews should reproduce it without provider-private reasoning state.

Model registry ID
gpt-6-astra

Sources OpenAI, ARC Prize

Anthropic · 2026-09-01 · Frontier · Reasoning

Claude Fable 5.1 in GitHub Copilot

Engineering leaders get a broadly distributed coding model, not an automatically enabled enterprise SKU. Put administrator entitlement and the documented 30-day default retention route into the architecture record alongside model version and benchmark results.

Model registry ID
claude-fable-5-1

Sources GitHub

Google · 2026-09-03 · Frontier · Reasoning

Gemini 3.8 Flash in GitHub Copilot

Treat current economics as a promotional observation, not a 2027 run rate. GitHub describes recovery from actionable terminal failures but publishes no reproducible benchmark or exact multiplier, so buyers should measure completed-task cost on their own harness.

Model registry ID
gemini-3-8-flash

Sources GitHub

Open weights

NVIDIA · 2026-09-01 · Open Frontier · Multimodal

Muse-Glimmer-30B NVFP4

This is a deployment-footprint update to an existing W33 model, not a new frontier row. Blackwell operators should test it for lower memory pressure, but NVIDIA's small Terminal-Bench gain over the 16-bit floating-point (BF16) baseline may be sampling noise and should not drive a capability claim.

Model registry ID
muse-glimmer-30b

Sources NVIDIA on Hugging Face

H Company · 2026-09-03 · Specialist · Multimodal

NeoMME-800M

Relevant for multimodal retrieval and agent memory, not generative-model substitution. Day-zero Transformers support makes it testable now, but no production latency, retrieval cost, or independent benchmark is public, so keep it out of production scorecards.

Sources Hugging Face, NeoMME paper

Architecture watch

Runtime version becomes a governed SKU property

ARC Prize's 35.9-point same-effort spread is too large to leave the runtime as an implementation footnote. Evaluation records should version context compression, memory visibility, state persistence, tool surface, retry budget, entitlement, and retention route with the model.

Examples
GPT-6 Astra Standard harness, GPT-6 Astra Provider Adapter

Sources ARC Prize, OpenAI

Distribution policy is part of the model SKU

Three closed-model movements arrived with material entitlement, retention, or price-window conditions. Procurement catalogs should stop recording only model family and API rate; they need surface, administrator state, retention route, promotion expiry, and provider adapter version.

Examples
Claude Fable 5.1 administrator enablement, Gemini 3.8 Flash promotional pricing, Astra enterprise access off by default

Sources GitHub, OpenAI

Open-weight progress shifted from new lineage to deployment compression

W36 did not repeat W35's open-frontier release wave. NVIDIA cut an existing checkpoint's footprint 2.4× and H Company opened compact multimodal encoders; both lower implementation friction, but neither supplies an independent frontier comparison. Buyers should value deployability without relabeling it capability convergence.

Examples
Muse-Glimmer-30B NVFP4, NeoMME-260M, NeoMME-800M

Sources NVIDIA on Hugging Face, H Company

Benchmark moves

ARC-AGI-3 Semi-Private

At matched maximum reasoning, Astra moved from 62.7% on the Standard harness to 98.6% on OpenAI's Provider Adapter, a 35.9-point system spread.

GPT-6 Astra, Provider Adapter maximum reasoning
98.6% / $17,332
GPT-6 Astra, Standard harness maximum reasoning
62.7% / $26,098

Sources ARC Prize

Artificial Analysis Intelligence Index v4.1.1

Astra entered at 61 at maximum effort, reinforcing that one interactive benchmark should not define general capability.

Claude Fable 5.1
66
Claude Opus 5
63
GPT-6 Astra, maximum effort
61

Sources Artificial Analysis

Terminal-Bench 2.1, vendor-reported checkpoint comparison

NVIDIA reports 47.05% for Muse-Glimmer NVFP4 versus 45.22% BF16, while cautioning that the small gain may be sampling noise.

Muse-Glimmer-30B NVFP4 (vendor-reported)
47.05%
Muse-Glimmer-30B BF16 (vendor-reported)
45.22%

Sources NVIDIA on Hugging Face

Tier scorecard

As of 2026-09-05

TierLeaderChallengerRead
Closed frontierClaude Fable 5.1Claude Opus 5Artificial Analysis Index remains the broad capability anchor; Astra's adapter result is harness-specific and should not replace it.
Open frontierGLM-5.3-FlashIBM Granite 4.2 30BNo new open-frontier release displaced W35's leaders; Muse-Glimmer NVFP4 changes footprint, not lineage or independent rank.
ReasoningClaude Fable 5.1GPT-6 AstraAstra is the system-design challenger; persistent state materially changes ARC results, while Fable retains the broader independent Index lead.
CodingClaude Fable 5.1Gemini 3.8 FlashBoth gained GitHub distribution; compare completed-task cost after administrator policy and year-end Gemini pricing are included.
MultimodalGemini 3.8 FlashGPT-6 AstraAstra adds 1M context and computer use; Gemini's Copilot rollout broadens surface reach, but GitHub publishes no neutral task benchmark.
Edge / smallNeoMME-800MNeoMME-260MSpecialist retrieval encoders, not generative leaders; test for multimodal memory only after measuring latency and retrieval quality.

Vendor signals

2026-09-03 · OpenAI

Astra enterprise access off by default; Standard API priced at $10/$50 per million input/output tokens

Astra is a premium governed rollout, not an automatic fleet replacement. Require explicit enablement, provider-adapter versioning, reasoning-effort capture, and a portability test before using its adapter result in procurement.

Sources OpenAI

2026-09-01 · Anthropic / GitHub

Fable 5.1 reaches general availability but Business and Enterprise administrators must enable it

Public availability does not equal enterprise entitlement. Catalog the selected safety and retention route, including the documented 30-day default outside approved zero-data-retention arrangements.

Sources GitHub

2026-09-03 · Google / GitHub

Gemini 3.8 Flash enters Copilot under introductory provider pricing through December 31

Do not annualize current economics into 2027. Measure completed-task cost now, then rerun the same workload when post-promotion pricing is published.

Sources GitHub

2026-09-01 · NVIDIA

Muse-Glimmer-30B NVFP4 ships Apache 2.0 for Blackwell with a 24.7 GB checkpoint

Blackwell self-hosters gain a smaller deployable artifact. Treat the vendor-reported benchmark delta as noise until reproduced, and value the release on footprint and runtime compatibility.

Sources NVIDIA on Hugging Face

Watchlist

Sep 9

GLM-5.3-Flash promotion expires

Recalculate the W35 open-frontier cost case on steady-state list pricing and independent provider variance.

September

Neutral Astra memory-harness reproductions

A visible-state harness closing part of the 35.9-point matched-effort gap would make the runtime gain more portable.

Q4 2026

Anthropic Enterprise Frontier Safeguards rollout

Look for precision, recall, rolling-window, alert-volume, audit, and customer-cloud cost evidence before approval.

Dec 31

Gemini 3.8 Flash introductory pricing expiry

The post-promotion rate determines whether today's Copilot task economics survive into 2027.

Next AA refresh

Astra and W35 open-model independent rankings

Use a common harness to test whether Astra's system advantage and W35's vendor-reported open scores generalize.

Changelog

  • Added GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash to the LLM tree.
  • Held the open-frontier scorecard after a deployment-compression week; Muse-Glimmer NVFP4 remains an update to the W33 lineage.
  • Credited ARC Prize for the harness finding and reframed the issue around SKU governance: entitlement, retention route, runtime version, effort, and price window.