Skip to content

Cross-stack flywheel

AI Stack Weekly

For officers tracking AI market movement.

The cheapest capability and the cheapest capacity both came from contracts and harnesses, not from chips or weights

Executive summary

7 minute read

Key takeaways

  • The week's biggest capability gain and its biggest capacity gain both arrived without new silicon or new weights: OpenAI tripled its ARC-AGI-3 score by retaining reasoning and compacting context on the same model, and Meta moved a 1 GW campus off its balance sheet into an 80/20 BlackRock venture carrying $12.5B of debt.
  • OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 and Terra 20% to $2/$12 while holding Sol at $5/$30 — on a fixed 1:4 workload that widens the cheapest-to-premium spread from roughly 5x to 25x, making tier routing a larger cost lever than vendor choice.
  • Moonshot shipped Kimi K3's full 2.8T weights on July 27, but the revenue-tiered license moves open-weight diligence from model evaluation to license classification: commercial hosting above $20M trailing revenue needs a separate deal.
  • Memory, not accelerators, set the price signal: SK hynix and Samsung both printed records and described HBM shortage into 2027-2028, and Amazon attributed a $20B increase in CY2026 capex guidance primarily to memory costs.
  • The networking hypothesis strained for the first time — Equinix posted record interconnection adds while interconnection revenue grew about 9% against roughly 16% total revenue growth, which is the specific falsification test named on the thesis page.
  • No accelerator launched, no closed frontier model shipped, and no frontier lab closed a funding round; the week's movement was entirely in contracts, licenses, harnesses, and memory supply.

By the numbers

GPT-5.6 cheapest-to-premium price spread on a fixed 1:4 workload, before and after the July 30 cut
5x to 25x — House measurement: Luna's representative bill fell from $25 to $5 while Sol held at $125
OpenAI's ARC-AGI-3 public-set score from harness memory settings alone, same model
13.3% to 38.3% — Vendor-reported, with roughly 6x fewer output tokens
Debt financing in the Meta-BlackRock El Paso venture — 1 GW, online from 2028
$12.5B — BlackRock funds take 80%, Meta retains 20% and leases the whole campus
Amazon CY2026 cash capex guidance, raised from ~$200B
~$220B — Earnings-call disclosure; Jassy attributed the increase primarily to memory prices
Microsoft commercial remaining performance obligation, up 84% year over year
$678B — Only about 30% is expected to convert inside twelve months
Kimi K3 total and active parameters, released under a revenue-tiered license
2.8T / 104B — 896 experts, 16 selected, 1M context

Big story

Nothing new was built this week. No accelerator launched, no closed frontier model shipped, and no frontier lab closed a round. Yet both the largest capability gain and the largest capacity gain on the board arrived anyway, and they arrived from the same place: the structures wrapped around assets that already existed.

On the capability side, OpenAI reported tripling its ARC-AGI-3 public-set score from 13.3% to 38.3% by changing two harness settings — retaining private reasoning between tool calls and compacting old context instead of truncating it — while consuming roughly six times fewer output tokens. Same weights, same model, different memory policy. That is a vendor-measured result on one benchmark, and it deserves the discount that implies. It is still the week's most consequential number, because it says a large fraction of what teams currently attribute to model quality is actually attributable to the scaffold they wrapped around it.

On the capacity side, Meta and BlackRock closed a venture for a 1 GW El Paso campus in which BlackRock-managed funds take 80% and Meta keeps 20%, contributes land and construction in progress, leases the entire campus back for an initial four years extendable toward twenty, and sits behind roughly $12.5B of debt against a residual-value guarantee threshold near $13B. Core Scientific separately converted AMD's rack roadmap into about 530 MW of contracted, third-party-owned capacity with a path toward 2.5 GW — before a single Helios rack has shipped. Neither transaction added a megawatt this week. Both changed who owns the asset, who carries the debt, and what a gigawatt costs the operator that uses it.

OpenAI's price cut belongs in the same frame rather than in a price-war frame. Luna fell 80% to $0.20/$1.20 and Terra 20% to $2/$12 while Sol held at $5/$30. On a fixed workload of one million input and four million output tokens, Luna's bill drops from $25 to $5 while Sol stays at $125 — the distance between the cheapest and most expensive tier of the same family widens from about 5x to 25x in a single day. The competitive story is secondary. The operational story is that tier selection now carries 25x of cost leverage inside one vendor, which makes routing policy a larger financial decision than vendor selection.

The decision implication runs three ways. Model buyers should re-run their agent evaluations with identical weights and different memory and routing policies before attributing any gap to the model, and should instrument per-tier token consumption and completion rates before moving volume to Luna on list price alone. Infrastructure buyers should read the El Paso and Core Scientific structures as the emerging default and price residual-value guarantees, lease terms, and third-party ownership as part of capacity cost. And anyone modeling 2027 unit economics should stop indexing to accelerator generations: with two memory suppliers describing shortage into 2028 and Amazon naming memory as the primary driver of a $20B capex revision, HBM contracts now set the clearing price that GPU roadmaps used to.

Flywheel arc · all-three

A large fraction of what teams currently attribute to model quality is actually attributable to the scaffold they wrapped around it.

  • No accelerator launched, no closed frontier model shipped, and no frontier lab closed a round — yet the week produced both a large capability gain and a large capacity gain, from harnesses and contracts rather than assets.
  • OpenAI tripled its ARC-AGI-3 score from 13.3% to 38.3% on the same model by retaining reasoning between tool calls and compacting context instead of truncating, with roughly 6x fewer output tokens.
  • Meta and BlackRock closed an 80/20 venture for a 1 GW El Paso campus with $12.5B of debt and a leaseback of up to twenty years; Core Scientific contracted ~530 MW of AMD capacity before any Helios rack shipped.
  • The Luna cut widened the cheapest-to-premium GPT-5.6 spread from about 5x to 25x on a fixed 1:4 workload, which makes routing policy a bigger cost decision than vendor choice.
  • What to do: vary harness memory policy before swapping models, price residual-value guarantees and leasebacks into capacity cost, and index 2027 unit economics to HBM contracts rather than to GPU generations.

Software lens

What this means

The model layer stopped competing on weights this week and competed on scaffolds and rate cards instead. OpenAI's harness result and DeepSeek's ten-point Index gain on unchanged architecture both say the same thing: post-training and context policy are now moving capability faster than parameter counts, so teams should re-baseline agents against memory and routing changes before buying a bigger model. On the open side, K3 and Inkling-Small expand the deployable frontier while relocating the real gate to licensing and VRAM footprint.

  • Hold the weights fixed and vary the memory policy first: this week's largest measured capability gain came from the scaffold rather than the model. Agent Techniques Weekly carries the method and its caveats.
  • Open weights now split into two diligence tracks: Kimi K3 needs license classification by counsel, Inkling-Small needs roughly 600 GB of BF16 VRAM or 180 GB at NVFP4.
  • See The Model Pulse for the full architecture and tree read on K3, Inkling-Small, and the DeepSeek re-post-train.

Jul 27

Moonshot published Kimi K3's full 2.8T MoE weights (104B active, 896 experts, 1M context) under a revenue-tiered license requiring a separate commercial deal for hosting above $20M trailing revenue

Sources Moonshot AI; Hugging Face; GitHub

Jul 30

OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 and Terra 20% to $2/$12, held Sol at $5/$30, and replaced Priority Processing with a Fast mode at 2x price for up to ~2.5x speed

Sources OpenAI

Jul 29

OpenAI reported tripling its ARC-AGI-3 public-set score from 13.3% to 38.3% using retained reasoning plus context compaction on the same model, with roughly 6x fewer output tokens

Sources OpenAI

Jul 31

DeepSeek put V4-Flash-0731 into public beta at unchanged $0.14/$0.28 pricing and unchanged 284B/13B architecture, and Artificial Analysis measured its Intelligence Index up 10 points to 50

Sources DeepSeek API changelog; Artificial Analysis

Jul 30

Thinking Machines released Inkling-Small under Apache 2.0 — 276B total, 12B active, multimodal — and Artificial Analysis scored it 40 against Inkling's 41 at under a third the parameters

Sources Thinking Machines Lab; Artificial Analysis

Hardware lens

What this means

This was a memory week, not a silicon week, and that matters more for planning. With SK hynix and Samsung both reporting records while describing shortage persisting into 2028, and Amazon attributing a $20B capex increase primarily to memory prices, the binding input on 2027 capacity has visibly moved from accelerator supply to HBM contracts. Buyers negotiating multi-year reserved capacity should assume memory pass-through in pricing and should treat 2027 delivery commitments that are not backed by disclosed HBM allocation as schedule risk rather than schedule.

  • Index 2027 accelerator delivery risk to HBM allocation, not GPU roadmaps: two suppliers described shortage into 2028 and 2027 pricing is still being negotiated now.
  • Memory inflation is already inside cloud unit economics — Amazon named it as the primary driver of a $20B capex revision, so reserved-capacity pricing should be expected to follow.
  • Core Scientific's 530 MW turns AMD's rack story into contracted US capacity with a named landlord, but deliveries begin in 2027 and nothing is energized yet.

Jul 29

SK hynix reported record Q2 revenue of KRW 79.3T at a 76% operating margin, began HBM4 mass shipments, and said long-term agreements with about ten customers are set while 2027 volumes and prices remain under negotiation

Sources SK hynix; Yonhap

Jul 30

Samsung posted record semiconductor results, shipped industry-first HBM4E samples, and guided Q3 HBM4 sales to more than triple sequentially while describing shortage worsening in 2027 and persisting into 2028

Sources Samsung Electronics; CNBC

Jul 30

Amazon raised CY2026 cash capex guidance to about $220B from about $200B, attributing the increase primarily to higher memory prices, and said capacity will still fall short of 2026 demand

Sources Amazon; CNBC earnings-call coverage

Jul 28

Core Scientific and AMD signed 15-year agreements covering about 530 MW across five sites with more than $14B of potential base contracted revenue and AMD rights to reserve toward roughly 2.5 GW, with deliveries starting 2027

Sources Core Scientific

Jul 29

Microsoft spent $41B on capex in the June quarter with roughly two-thirds on short-lived compute assets, added about 1 GW and 31 data centers in the quarter, and reached 88 data centers online across FY26

Sources Microsoft; Datacenter Dynamics

Networking lens

What this means

Optics, switch silicon and the standards bodies were quiet in-window, so the honest networking read comes from a colocation print rather than a product launch — and it cuts against this publication's own hypothesis. Record interconnection adds at Equinix coincided with interconnection revenue growing roughly 9% while total revenue grew about 16%. Investors and architects should not treat unit adds as proof that interconnect is the fastest-growing layer; the revenue mix is the test, and this quarter it went the other way.

  • Equinix is the week's most important negative result: record interconnection units with interconnection revenue growing slower than the total business is precisely the pattern that would falsify the durable-networking thesis.
  • Watch normalized interconnection revenue against total revenue on the 10-Q and in Q3 prints before treating one quarter as a trend.
  • Campus announcements continue to quantify power and leave fiber and long-haul capacity undisclosed, which understates the interconnect requirement of distributed training.

Jul 29

Equinix reported a record 9,700 net interconnection adds and $453M of interconnection revenue, but interconnection revenue grew about 9% against roughly 16% total revenue growth, and raised FY2026 guidance to $10.205-10.285B

Sources Equinix

Jul 29

Microsoft said dock-to-live time for new GPUs in its largest regions fell nearly 50% year over year, with servers, networking and software assets rising to $215.9B from $132.8B

Sources Microsoft earnings call; Datacenter Dynamics

Jul 29

The DOE Paducah selection was marketed on existing transmission, water, fiber and road access with NextEra to upgrade transmission near the campus, but disclosed no interconnect or long-haul capacity figures

Sources DOE; Brookfield; NextEra Energy

Capital flow

Capital in, revenue out, and the direction of travel.
CategoryCapital inRevenue outBurn to revenueMovement
Frontier Labs — OpenAI, Anthropic, Google DeepMind, xAI~$95B · prior ~$95B · flat~$21B · prior ~$21B · flat~4.5xNo frontier lab closed primary financing in-window. The reported Nvidia backstop of up to ~$250B for OpenAI's Ohio campus is talks-only and is deliberately excluded from capital in.
Hyperscaler-Hosted — Azure-OpenAI, AWS-Anthropic, Google Cloud-Gemini, Oracle-OCI~$250B · prior ~$230B · up~$70B · prior ~$62B · up~3.6xAmazon raised CY2026 cash capex guidance by about $20B to ~$220B and Meta lifted its full-year floor to $130-145B, while Azure passed $100B of annual revenue and AWS re-accelerated to +37%.
Neoclouds — CoreWeave, Nscale, Crusoe, Lambda, Fluidstack, IREN~$17B · prior ~$17B · flat~$5B · prior ~$5B · flat~3.4xNo neocloud financing closed in-window and CoreWeave's backlog remains the March figure until its August 11 print, but Core Scientific contracted about 530 MW to AMD with a path toward 2.5 GW.
On-Prem / Hybrid — Enterprise GPU clusters, sovereign and national programs, Cisco / Dell / HPE~$101B · prior ~$101B · flat~$36B · prior ~$36B · flat~2.8xNo change to the aggregate. Kimi K3's downloadable weights expand what a private cluster can actually run, while the license's revenue trigger adds a legal gate that did not previously exist for open models.

Frontier Labs detail

The structural story is that the next tranche of frontier funding may not be equity at all. A vendor credit wrap lets a lab borrow against a supplier's balance sheet, which changes who bears the risk without changing the lab's cash position. Until definitive documentation exists, runway analysis should treat it as unfunded.

Capital in value
$95B
Revenue out value
$21B
  • 2026-07-27 · Nvidia reported in advanced talks to guarantee up to ~$250B of OpenAI Ohio lease and construction debt; Nvidia declined to comment and no definitive agreement is documented · Up to ~$250B (reported, unsigned)

Sources CNBC; WSJ

Hyperscaler-Hosted detail

This is the first week in several where the revenue denominator moved as visibly as the capex numerator. Azure crossing $100B annual and AWS reaching a company-reported ~$169B annualized run-rate improve observability, but neither company breaks out AI-attributable revenue, so the ratio remains an estimate. Backlog figures are contracted, not recognized: Microsoft's $678B RPO converts only about 30% inside twelve months.

Capital in value
$250B
Revenue out value
$70B
  • 2026-07-30 · Amazon raised CY2026 cash capex guidance to ~$220B from ~$200B, attributing the increase primarily to memory prices; AWS backlog cited at $496B on the call · ~$220B guide
  • 2026-07-29 · Microsoft reported $41B of quarterly capex, Azure annual revenue above $100B, and commercial RPO of $678B (+84% YoY) with roughly 30% converting inside twelve months · $41B quarter / $678B RPO
  • 2026-07-30 · Meta reported $31.1B of quarterly capex and raised the floor of its FY2026 range to $130-145B while free cash flow fell 91% to $784M · $31.1B quarter / $130-145B guide

Sources Amazon; CNBC earnings-call coverage · Microsoft investor relations · Meta SEC Exhibit 99.1

Neoclouds detail

The category's most important number is still unobserved: whether contracted backlog converts to revenue at the pace its financing assumes. Core Scientific's AMD agreement is the more interesting structural event, because it gives an accelerator vendor a named landlord and gives operators a second rack supply chain to negotiate against — both contracted, neither energized.

Capital in value
$17B
Revenue out value
$5B
  • 2026-07-28 · Core Scientific and AMD signed 15-year agreements for ~530 MW across five sites with >$14B potential base contracted revenue and AMD rights toward ~2.5 GW; deliveries begin 2027 · ~530 MW / >$14B base term
  • 2026-07-27 · CoreWeave scheduled Q2 results for August 11; the $99.4B March 31 backlog remains the last official print with no in-window financing disclosure

Sources Core Scientific · CoreWeave investor relations

On-Prem / Hybrid detail

Open weights got materially more useful and materially more conditional in the same release. Private-cluster owners can now run a 2.8T MoE with a million-token context, but procurement has to classify intended use against the license before deployment, and Inkling-Small's roughly 600 GB BF16 footprint sets a hardware floor that most enterprise clusters will meet only at NVFP4.

Capital in value
$101B
Revenue out value
$36B
  • 2026-07-27 · Moonshot published Kimi K3 weights under a license requiring a separate commercial agreement for hosting above $20M trailing-twelve-month revenue, with attribution duties above 100M MAU

Sources Moonshot AI; VentureBeat analysis

See the full capital-flow breakdown

Signal vs noise

Signal score 5/5

Microsoft's Azure passed $100B of annual revenue while commercial remaining performance obligation reached $678B, up 84% year over year.

Both figures are primary and audited-grade. The discipline is in the second number: only about 30% of RPO converts inside twelve months, so this is contracted backlog rather than near-term revenue and should not be compared against annual capex as though it were.

Sources
Microsoft investor relations press release and earnings call

Signal score 5/5

Meta and BlackRock closed a venture for a 1 GW El Paso campus at roughly $14B development cost, with $12.5B of debt and BlackRock funds holding 80%.

Structure, ownership split, lease term and residual-value threshold are all disclosed. This is the week's most transferable fact: it is a template for moving gigawatt-scale AI capacity off a hyperscaler balance sheet while retaining operational control.

Sources
Meta newsroom and investor relations; advisors Morgan Stanley and J.P. Morgan

Signal score 4/5

Moonshot shipped Kimi K3's full 2.8T weights, converting last week's open-frontier watch item into a deployment fact.

The artifacts are real and dated, so the deployment claim holds. The license is the part buyers will feel: hosting above $20M trailing revenue needs a separate agreement, which moves the decision from an ML evaluation to a legal classification.

Sources
Hugging Face and GitHub artifacts with a published license; secondary enterprise analysis

Signal score 4/5

OpenAI cut GPT-5.6 Luna 80% and Terra 20% while holding Sol unchanged, widening the intra-family price spread.

The rate card is a primary fact and the spread arithmetic follows from it. What remains unmeasured is completed-task cost: Luna may consume more tokens or attempts on hard work, so the 25x list-price gap is an upper bound on realized savings, not a forecast.

Sources
OpenAI product announcement and pricing documentation; multi-outlet confirmation

Signal score 3/5

HBM supply will remain short through 2027 and into 2028, keeping memory pricing elevated.

Two independent suppliers pointing the same direction in the same week is meaningful, but the specific shortage language comes from call commentary rather than an audited disclosure. Treat it as a strong planning assumption, not a confirmed supply curve.

Sources
SK hynix and Samsung Q2 releases plus earnings-call paraphrase in Korean business press

Signal score 2/5

Nvidia will guarantee up to $250B of OpenAI's data-center financing, effectively underwriting the Ohio campus.

Noise as stated. There is no filing, no definitive agreement and no confirmation from either party. Vendor credit-wrapping is a genuinely important emerging structure, which is exactly why it should be scored on documentation rather than on reporting.

Sources
WSJ and CNBC reporting citing anonymous sources; Nvidia declined to comment

Signal score 2/5

Microsoft's disclosure of more than 30 million paid Copilot seats demonstrates that enterprise AI is delivering measurable productivity returns.

Seats are a distribution metric, not a value metric. Neither average revenue per seat, weekly active use, nor workflow attach was disclosed, and a 60% sequential Copilot revenue figure is vendor-characterized rather than a reported line item. Procurement should ask for the usage data before treating seat growth as ROI evidence.

Sources
Microsoft earnings call; vendor-reported seat and interaction counts

House measurement

Filing-Derived

OpenAI's July 30 price change widened the cost distance between its cheapest and most expensive GPT-5.6 tier from roughly 5x to 25x on an identical workload.

Method: Take a fixed representative agentic workload of 1,000,000 input tokens and 4,000,000 output tokens — the same 1:4 ratio this publication has used in prior issues for continuity. Compute the list-price bill for each GPT-5.6 tier before and after the July 30 change, using published per-million-token rates: Luna $1.00/$6.00 before and $0.20/$1.20 after; Terra $2.50/$15.00 before and $2.00/$12.00 after; Sol $5.00/$30.00 unchanged. Bill equals input rate plus four times output rate. Divide the Sol bill by the Luna bill to get the intra-family spread. No caching discounts, Fast-mode premium, batch pricing or retry overhead is applied.

Implication: Tier routing is now a larger financial lever than vendor selection inside a single model family. A team that can classify 60% of its traffic as Luna-eligible saves more than it would by switching providers, which means routing policy deserves the scrutiny normally reserved for procurement. It also means per-tier token telemetry stops being an engineering nicety and becomes a finance requirement: without it, nobody can tell whether a 25x list-price gap is producing real savings.

Caveats: This is list-price arithmetic on a synthetic ratio, not a measured workload. It does not show that Luna completes the same tasks — cheaper models frequently consume more tokens, take more attempts, or fail outright, and any of those compress the realized ratio well below 25x. It excludes prompt caching, batch pricing, the Fast-mode premium, and the cost of retries and human review. A real comparison requires completed-task measurement, which no independent party has published.

Luna bill on the fixed workload
$25.00 to $5.00 — An 80% reduction, matching the headline rate cut exactly because the workload ratio is fixed
Terra bill on the fixed workload
$62.50 to $50.00 — A 20% reduction
Sol bill on the fixed workload
$125.00 unchanged — The premium tier did not move, which is what widens the spread
Cheapest-to-premium spread within GPT-5.6
5.0x to 25.0x — Sol divided by Luna, before and after July 30

Sources OpenAI, advancing the price-performance frontier with GPT-5.6 · OpenAI API pricing documentation

Synthesis · Connecting the dots

Inductive · 74% confidence

The marginal unit of AI advantage moved from the asset to the contract and the harness: the week's largest capability gain and its largest capacity gain both arrived without new silicon or new weights.

Steel-man: The strongest objection is that neither move creates anything: the harness gain is a vendor-measured result on a single puzzle benchmark with obvious incentive to publish it, and an SPV is an accounting rearrangement that does not add a megawatt or a FLOP. Both points are fair, and the claim narrows accordingly. What survives is that buyers do not procure the frontier, they procure a given capability or capacity at a given cost, and both of those costs moved substantially this week through structure alone. That is a real change in what a buyer pays even if the underlying frontier stood still.

  • OpenAI tripled its ARC-AGI-3 public-set score from 13.3% to 38.3% by retaining private reasoning between tool calls and compacting old context rather than truncating it, on the same model and with roughly 6x fewer output tokens.
  • Meta moved a 1 GW campus off its own balance sheet into an 80/20 BlackRock venture carrying $12.5B of debt, a residual-value guarantee threshold near $13B, and a leaseback extendable toward twenty years.
  • Core Scientific converted AMD's rack roadmap into about 530 MW of contracted, third-party-owned capacity with a path toward 2.5 GW, before a single Helios rack has shipped.

Sources OpenAI on retained reasoning and context compaction, ARC-AGI-3 13.3% to 38.3% · Meta and BlackRock El Paso venture, 1 GW, ~$14B development cost, $12.5B debt · Core Scientific and AMD, ~530 MW across five sites, deliveries from 2027

Abductive · 71% confidence

Memory, not accelerators, is now setting the clearing price for 2027 AI capacity, and the signal is arriving through hyperscaler operating guidance rather than through chip roadmaps.

Steel-man: The memory attribution comes from an earnings call rather than a filing line item, and memory is a conveniently external explanation for capex growth that might instead reflect volume or a bad negotiation. The inference holds because two independent suppliers described the same shortage direction in the same week while both were reporting records — a coincidence that is harder to explain as narrative management than as supply reality. It would be weakened considerably if Q3 shows memory prices flat while capex guidance rises again.

  • SK hynix and Samsung both printed record memory quarters and described HBM shortage persisting through 2027 and into 2028, with 2027 volumes and prices still under negotiation across roughly ten customers.
  • Amazon raised CY2026 cash capex guidance by about $20B to ~$220B and attributed the increase primarily to higher memory prices, not to additional capacity.
  • No new accelerator silicon launched in-window, so the entire cost signal for next year's capacity reached the market through the memory supply chain and a cloud provider's guidance revision.

Sources SK hynix Q2 2026 results: HBM4 mass shipments, 2027 volumes and prices under negotiation · Amazon Q2 2026: CY2026 capex guidance raised to ~$220B on memory prices · Samsung Q2 2026: first HBM4E samples, Q3 HBM4 sales to more than triple, shortage into 2028

Inductive · 68% confidence

Open weights and open protocols moved in opposite directions on enterprise usability this week: the weights became more conditional while the connector standard became more portable.

Steel-man: A revenue-tiered license is irrelevant to the majority of enterprises that self-host for internal use, so describing the weights as more conditional overstates the friction for a typical buyer. The distinction still matters because it relocates diligence from model evaluation to license classification, which is a different team on a slower clock. The claim would break if K3's license turns out to be read as unambiguously permissive for internal enterprise deployment without counsel involvement.

  • Moonshot published Kimi K3's full 2.8T weights but gated commercial hosting above $20M trailing revenue behind a separate agreement, with attribution duties above 100M monthly active users.
  • The MCP 2026-07-28 specification went final with a stateless, cacheable, gateway-routable core, enterprise-managed authorization, and dynamic client registration deprecated.
  • Snowflake and GitHub shipped governance and orchestration surfaces in the same week that assume MCP is the connector substrate rather than one integration option among several.

Sources Kimi K3 weights and revenue-tiered license on Hugging Face · MCP specification 2026-07-28: stateless core, cacheable tool listings, DCR deprecated

Synthesis · Thesis test

Hypothesis 1 · Supported

The cycle is accelerating, not slowing.

The cadence held even in a week with no new silicon and no new closed frontier model. DeepSeek gained ten Artificial Analysis Index points on unchanged architecture and unchanged price through re-post-training alone, Moonshot shipped 2.8T open weights, Thinking Machines matched its flagship within one Index point at under a third of the parameters, and OpenAI tripled a benchmark score through harness policy. Four independent capability or cost improvements inside seven days is acceleration, and notably none of them required a hardware generation to arrive first — which is a faster loop than the flywheel model assumes.

Counter-evidence: No accelerator launched in-window, and Gemini 3.5 Pro missed a general-availability deadline this publication had predicted for July 31, remaining in partner testing. If the hardware side of the flywheel is stalling while software compounds, the coupled acceleration this hypothesis describes is not what is actually happening.

Sources DeepSeek V4-Flash-0731 Intelligence Index up 10 points at unchanged $0.14/$0.28 · Inkling-Small at Index 40 versus Inkling at 41, under a third the parameters, Apache 2.0

Hypothesis 2 · Supported

Capital is concentrated, returns are diffuse.

The earnings prints made the gap unusually legible. Microsoft's commercial RPO reached $678B, up 84%, while only about 30% converts inside twelve months. Meta spent $31.1B in a quarter and reported free cash flow of $784M, down 91%. Amazon's trailing free cash flow was negative $7.6B while it raised capex guidance by $20B. Capital is concentrating at four balance sheets and the returns are showing up as contracted future obligations and as enterprise seat counts rather than as current cash generation at the spender.

Counter-evidence: Azure crossing $100B of annual revenue and AWS re-accelerating to 37% growth are the strongest evidence yet that concentrated capex is converting into revenue at the spender rather than diffusing away from it. If that continues for two more quarters, the hypothesis needs rewriting rather than defending.

Sources Microsoft commercial RPO $678B, up 84%, roughly 30% inside twelve months · Meta Q2 2026: $31.1B capex, free cash flow $784M, down 91% year over year

Hypothesis 3 · Strained

Networking is the durable layer.

This publication named the falsification test in advance: interconnect growth lagging compute growth at the colocation operators. Equinix delivered exactly that pattern. A record 9,700 net interconnection adds coincided with interconnection revenue of $453M growing roughly 9% on a normalized basis while total revenue grew about 16%. Unit demand for interconnect is genuinely at a record, but the pricing power the hypothesis predicts should show up in revenue mix, and this quarter the mix moved toward power and space instead. One quarter from one operator is not the two consecutive quarters required for a revision, so the hypothesis is flagged rather than rewritten.

Sources Equinix Q2 2026: record 9,700 net interconnection adds, interconnection revenue growth below total

Hypothesis 4 · Supported

Open weights pull the floor up.

Three releases in one week expanded what a private cluster can run. Kimi K3's 2.8T weights with a million-token context are downloadable, Inkling-Small is Apache 2.0 at an Index of 40, and DeepSeek's cheapest tier gained ten Index points without a price change. All three re-route compute demand toward self-hosted and sovereign capacity rather than reducing it, which is the mechanism this hypothesis describes. The demand-catalyst reading is intact.

Counter-evidence: Both releases carry gates the word open obscures. K3 requires a separate commercial agreement for hosting above $20M trailing revenue, and Inkling-Small needs roughly 600 GB of VRAM at BF16 or 180 GB at NVFP4. If open frontier models keep arriving with revenue triggers and cluster-class footprints, the floor they pull up is only reachable by organizations that already had a floor.

Sources Kimi K3: 2.8T total, 104B active, 896 experts, 1M context, weights published July 27 · Inkling-Small released under Apache 2.0 at 276B total and 12B active parameters

Hypothesis 5 · Supported

Power is the binding constraint for the next 24 months.

The Paducah selection is the cleanest illustration yet of the constraint's shape. Delivering roughly 1.8 GW to a campus required pairing it with up to 2 GW of new dedicated natural gas and up to 2.6 GW of storage, with completion targeted for 2031 and the power service agreement still subject to Kentucky regulatory approval. Meta's fully financed 1 GW El Paso campus does not come online until 2028. In both cases capital was available immediately and electricity was not, which is precisely the ordering this hypothesis asserts.

Counter-evidence: Both projects found a path. Dedicated generation, federal land, and third-party capital together produced credible multi-gigawatt commitments in a single week, which suggests power is an expensive and slow constraint rather than a binding one. A binding constraint stops projects; this week it only delayed and repriced them.

Sources DOE Paducah selection: up to 2 GW dedicated gas plus 2.6 GW storage, completion 2031 · NextEra Energy on transmission upgrades and dedicated generation for the Paducah campus

Synthesis · Pattern watch

Inductive · 5 weeks observed

Frontier model economics are shifting from headline list price toward route-aware completed-task cost, and the publicly available evidence is not keeping up with the pricing changes.

Next expectation: Within two weeks an independent evaluator publishes a completed-task cost comparison for GPT-5.6 Luna that reports token consumption and retry or failure rates, not list price alone. If nothing appears by mid-August, the gap between pricing changes and public measurement is widening rather than closing, and buyers are routing on marketing.

  • W27-W28: headline per-token cuts arrived alongside separate reasoning tiers, making a single advertised price ambiguous for the first time.
  • W29: Moonshot raised Kimi pricing while DeepSeek introduced time-of-day differentials, so the same model could carry different costs by clock.
  • W30: Opus 5 was marketed on comparable quality at fewer output tokens, and Gemini 3.6 Flash on lower cost per completed task, both vendor-measured.
  • W31: OpenAI cut Luna 80% and held Sol flat, widening the intra-family spread on a fixed workload from roughly 5x to 25x, with no independent completed-task measurement published for any of it.

Inductive · 3 weeks observed

Large AI capacity is migrating off operator balance sheets into third-party ownership structures backed by contractual credit support rather than by the operator's own equity.

Next expectation: A second hyperscaler discloses a data-center joint venture or special-purpose vehicle with third-party majority ownership and some form of residual-value guarantee before the end of Q4 2026. If instead the next large campus is financed on balance sheet, the El Paso structure is a Meta-specific response to its own free-cash-flow compression rather than an industry template.

  • W29-W30: Hut 8 disclosed a 15-year, $9.8B base-term agreement for 352 MW, and OpenAI's Camellia campus was structured with customer-funded terms including curtailment obligations.
  • W30: AMD and Anthropic paired up to $5B of equity with up to 2 GW of deployment, tying vendor capital directly to capacity commitments.
  • W31: Meta and BlackRock closed an 80/20 El Paso venture with $12.5B of debt and a residual-value guarantee near $13B, Core Scientific contracted ~530 MW to AMD, and Nvidia was reported in talks to guarantee up to ~$250B of OpenAI financing.

Synthesis · Second-order effects

Next two to three quarters

OpenAI's Luna cut widens the cheapest-to-premium price spread within one model family from roughly 5x to 25x on a fixed workload.

Model-tier selection becomes a larger cost lever than vendor selection, which shifts effective budget authority toward whoever owns routing policy and makes per-tier token telemetry a finance requirement rather than an engineering nicety. Teams without that instrumentation cannot tell whether a 25x list-price gap is producing any real savings, so the first casualty is the credibility of AI cost forecasts built on blended token assumptions.

Who moves
FinOps and platform engineering teams, AI gateway and routing vendors, and model providers that compete on headline list price rather than completed-task cost

Six to twelve months

The MCP 2026-07-28 specification removes sessions, makes tool listings cacheable, and deprecates dynamic client registration.

Connector traffic becomes ordinary cacheable HTTP that any gateway can route, audit and cache, which turns the enterprise agent control point into an infrastructure purchase rather than a platform choice. That erodes differentiation for agent platforms whose principal value was connector breadth, and it moves the buying conversation toward the vendors who already own the gateway and identity layers.

Who moves
Agent platform vendors, API gateway and data-cloud vendors, and enterprise security and identity teams now inheriting connector authorization

Through the 2027 contracting cycle

Amazon attributes a $20B capex guidance increase primarily to memory prices while two suppliers describe HBM shortage persisting into 2028.

Memory pricing becomes an explicit pass-through in cloud unit economics, so GPU-hour prices stop tracking accelerator generations and start tracking HBM contract cycles. That inverts the usual advice on capacity commitments: waiting for the next accelerator generation no longer reliably lowers cost, because the binding input reprices on a different and less visible schedule.

Who moves
Cloud buyers negotiating multi-year reserved capacity, neoclouds with fixed-price backlog, memory suppliers, and CFOs modeling AI unit cost

Synthesis · Strategic outlook

The useful way to read this week is that AI got cheaper and more available without anything new being built, and that should change where leaders look for advantage. When a benchmark score triples on harness settings alone, when a gigawatt changes hands through an ownership structure rather than a construction milestone, and when the cheapest tier of a model family drops 80% while the most expensive holds flat, the differentiated decisions are no longer about which model or which chip. They are about routing policy, license classification, contract structure, and memory allocation — four things that sit with legal, finance, and platform engineering rather than with a research team.

The honest counterweight is that this week also produced our first strained hypothesis in several. Equinix posted record interconnection adds while interconnection revenue grew more slowly than the total business, which is the pattern we named in advance as the test that would falsify our networking thesis. We are not revising on one quarter from one operator, but we are watching it rather than explaining it away, and the Q3 colocation prints in November are the decision point.

For the next quarter, three things carry the most information. CoreWeave's August 11 print is the first real read on whether neocloud backlog converts. Nvidia's August 26 guide is the first read on the Rubin ramp against AMD's newly contracted 530 MW. And the 2027 HBM settlements now under negotiation will set delivery schedules more tightly than any published roadmap. Everything else this week was structure; those three are evidence.

Where we differ

Differ

The week's defining story is an AI price war, with OpenAI's 80% Luna cut as evidence that competition has shifted decisively toward cost.

Framing this as a price war points attention at the wrong decision. The cut left Sol untouched, so on a fixed 1:4 workload the distance between OpenAI's own cheapest and most expensive tier widened from about 5x to 25x. That makes internal routing policy the dominant cost lever and vendor choice secondary — a different action than shopping providers.

Sources VentureBeat

Extend

Kimi K3's weights are genuinely open, with a licensing caveat that enterprises should be aware of.

The caveat is the story, not a footnote to it. A revenue-tiered license moves open-weight diligence off the ML team and onto counsel, because the question is no longer whether the model performs but how the intended deployment classifies. Add Inkling-Small's roughly 600 GB BF16 footprint and the practical gate on open frontier models this week became legal review plus cluster capacity.

Sources VentureBeat

Differ

Record interconnection adds at Equinix confirm that interconnection is the durable, fastest-growing layer of AI infrastructure.

Record unit adds coincided with interconnection revenue growing roughly 9% while total revenue grew about 16%. That is the exact pattern this publication named in advance as the falsification test for its own networking hypothesis, so we mark the hypothesis strained rather than claim confirmation from the unit count. Units are demand; revenue mix is the test.

Sources Equinix Q2 2026 results

Open

Nvidia will guarantee up to $250B of OpenAI's data-center financing, effectively underwriting the Ohio campus.

Nvidia declined to comment, no definitive agreement exists, and the reporting rests on anonymous sources. Vendor credit-wrapping may well become the binding capital-markets product for non-investment-grade labs, which is precisely why it should be scored on a filing rather than a headline. We have written a prediction against documentation instead of booking the number.

Sources WSJ, confirmed by CNBC

Levers

MetricCurrentPriorDirectionThreshold
Frontier lab cash runway at current burn~30-40 months, unchanged — no lab closed primary financing in-window, while a reported ~$250B Nvidia guarantee for OpenAI's Ohio campus would move funding from equity to vendor-guaranteed debt if it is ever signed~30-40 months; unchanged cash estimate, but OpenAI's reported infrastructure obligation rises to $750B through 2030flatBelow 18 months for any top-four lab
Hyperscaler AI capex to disclosed AI revenue ratio~3.6x on the publication's committed-capital estimate, improved from ~3.7x as Azure crossed $100B annual revenue and AWS re-accelerated to +37% even while Amazon added ~$20B to CY2026 capex guidance~3.7x on committed capital of ~$230B against a revenue proxy of ~$62BdownAbove 6x sustained for two consecutive quarters
CoreWeave contracted revenue backlog$99.4B as of March 31, unchanged — the Q2 print lands August 11, so backlog-to-revenue conversion stays unobserved while AWS disclosed a $496B contracted backlog for scale comparison$99.4B as of March 31, with the next disclosure pendingflatSequential decline, or conversion below 15% annually
NVIDIA quarter-over-quarter data center revenue$75.2B for Q1 FY27, unchanged with no earnings event in-window, though Core Scientific's ~530 MW AMD commitment and the reported OpenAI financing backstop both bear on the August 26 guide$75.2B for Q1 FY27, +12% sequentiallyflatTwo consecutive quarters of sequential decline
Open-weight to closed-model capability gap on codingNarrowing into a deployment fact: Kimi K3's weights shipped July 27 at an Artificial Analysis Index of 57 against roughly 60 for the closed leaders, and Inkling-Small reached 40 under Apache 2.0 at under a third of Inkling's parametersWidening on paper — Kimi K3's API scored 57.1 against Claude Fable 5 at 59.9 while K3's weights remained unreleaseddownOpen weights within 2 Index points of the closed leader
Sovereign AI program commitments~15 programs and ~$186B, unchanged — the DOE Paducah selection is roughly $100B of private capital on federal land rather than a new sovereign appropriation~15 programs and ~$186B in announced commitmentsflatAbove 20 programs or $250B committed
PJM capacity auction clearing price$325.00 per MW-day for 2028/29, unchanged with no new auction — Paducah's up to 2 GW of dedicated gas plus 2.6 GW of storage shows large load routing around auction scarcity rather than bidding into it$325.00 per MW-day for 2028/29, at the administrative capflatA second consecutive auction clearing at the cap
Time from interconnection request to energization60-84 months, unchanged — Paducah targets 2031 completion and Meta's 1 GW El Paso campus comes online from 2028, both multi-year despite fully committed capital60-84 months in the major US interconnection queuesflatBelow 48 months in two or more major queues
Cost per task, frontier reasoning modelThe cheap tier reset: GPT-5.6 Luna fell 80% to $0.20/$1.20 and DeepSeek V4-Flash-0731 gained 10 Artificial Analysis Index points at unchanged $0.14/$0.28, while Sol held at $5/$30Falling on list price, with Opus 5 claiming comparable quality at fewer output tokens and Gemini 3.6 Flash claiming lower per-task costdownA frontier-tier reasoning model below $1 per million output tokens
Custom silicon share of hyperscaler AI compute~34-37%, unchanged — Amazon said its AI and Chips businesses each passed a company-reported $25B annualized run-rate without disclosing the mix, and Core Scientific's AMD deal is merchant-accelerator competition rather than in-house silicon~34-37% of hyperscaler AI compute on internally designed acceleratorsflatAbove 45% share

Frontier lab cash runway at current burn

Measures how long the frontier labs can sustain current burn without new capital. A vendor credit wrap changes who bears the risk but not the lab's cash position, so unsigned backstops do not extend runway.

Hyperscaler AI capex to disclosed AI revenue ratio

The denominator got more observable this week for the first time in months: Azure past $100B annual and AWS at a company-reported ~$169B annualized run-rate. Neither company discloses AI-attributable revenue, so the ratio remains an estimate rather than a measurement.

CoreWeave contracted revenue backlog

Backlog is the neocloud category's core collateral, and its conversion rate is the number its financing implicitly assumes. Until August 11 the category is priced on an estimate rather than a print.

NVIDIA quarter-over-quarter data center revenue

The single cleanest read on whether AI infrastructure demand is still compounding. This week added competitive context rather than data: AMD now has contracted US capacity with a named landlord.

Open-weight to closed-model capability gap on coding

Measures whether a self-hosted model can substitute for a frontier API. The gap narrowed in deployability rather than in score, and the remaining friction is license classification and VRAM footprint rather than capability.

Sovereign AI program commitments

Tracks state-directed AI capital as distinct from corporate capex. Paducah is the pattern to watch: governments are contributing land and permitting rather than money, which keeps the sovereign total flat while adding real capacity.

PJM capacity auction clearing price

The clearest market price for grid scarcity in the largest US market. The more informative development is behavioral: campuses are increasingly procuring dedicated generation instead of competing in the capacity market.

Time from interconnection request to energization

The hard limit on how fast AI capacity can actually arrive. This week's announcements demonstrate that capital is no longer the binding constraint on gigawatt-scale campuses; schedule is.

Cost per task, frontier reasoning model

Tracks the real unit economics of agentic work rather than headline token prices. Luna now sits below the threshold at $1.20 output, but completed-task cost including retries remains unmeasured by any independent party.

Custom silicon share of hyperscaler AI compute

Measures how much of hyperscaler AI compute escapes merchant accelerator pricing. Amazon's $25B chips run-rate is the largest disclosed custom-silicon figure to date but is not broken out against its Nvidia fleet.

Predictions

Software · 62% confidence

An independent evaluator publishes completed-task cost showing GPT-5.6 Luna at least 60% cheaper per completed agentic task than GPT-5.6 Terra by September 30, 2026.

ID
p71-luna-task-cost-sep30
Deadline
By September 30, 2026
Trigger
A published third-party harness result comparing Luna and Terra on the same agentic task set, reporting total completed-task cost including retries, with Luna at least 60% lower.

Software · 44% confidence

An independent party reproduces at least a 15-point ARC-AGI-3 improvement from harness memory and compaction settings alone, holding model weights fixed, by October 31, 2026.

ID
p72-harness-memory-reproduction-oct31
Deadline
By October 31, 2026
Trigger
A published non-OpenAI result on the ARC-AGI-3 public set showing at least a 15 percentage-point gain attributable to retained reasoning or context compaction with the same underlying model.

Hardware · 81% confidence

SK hynix or Samsung states in a primary release or earnings transcript that 2027 HBM capacity is substantially committed or sold out by October 31, 2026.

ID
p73-hbm-2027-committed-oct31
Deadline
By October 31, 2026
Trigger
Company press release or official transcript containing an explicit statement that 2027 HBM supply is sold out, fully allocated, or substantially committed.

Networking · 46% confidence

No publicly listed global colocation operator reports Q3 2026 interconnection revenue growing faster than total revenue on a normalized basis by November 15, 2026.

ID
p74-interconnect-revenue-q3-nov15
Deadline
By November 15, 2026
Trigger
Q3 2026 results from at least two publicly listed global colocation operators, with none disclosing normalized interconnection revenue growth exceeding normalized total revenue growth.

Capital · 34% confidence

A definitive agreement of at least $100B in vendor-guaranteed AI data-center financing is publicly documented in a filing or company release by December 31, 2026.

ID
p75-vendor-backstop-documented-dec31
Deadline
By December 31, 2026
Trigger
An SEC filing or company press release describing an executed guarantee, credit support or backstop of at least $100B for a named AI data-center project.

Power · 63% confidence

A power service agreement for the Paducah AI campus is filed with the Kentucky Public Service Commission by December 31, 2026.

ID
p76-paducah-psc-filing-dec31
Deadline
By December 31, 2026
Trigger
A docketed Kentucky PSC filing containing a power service agreement naming the Paducah campus and the serving utility.

Prior predictions scored

Pending · Software

Artificial Analysis publishes an Opus 5 Intelligence Index result within three points of Claude Fable 5 by August 15, 2026.

Interim: an Artificial Analysis write-up dated July 24 puts Opus 5 at an Intelligence Index of 61 against roughly 60 for Fable 5, which would satisfy the trigger. Two problems keep this pending rather than scored. The trigger specifies a model page and the available evidence is an article, and the article predates the prediction itself — a calibration flaw worth naming, since a forward prediction should not be satisfiable by evidence that already existed when it was written.

ID
p67-opus5-aa-gap-aug15
Confidence
74%
Deadline
By August 15, 2026
Trigger
A public Artificial Analysis model page scoring Opus 5 no more than 3.0 Index points below Fable 5 on the then-current methodology.

Pending · Hardware

At least one named customer reports receiving a production AMD Helios rack for workload qualification by June 30, 2027.

Interim: Core Scientific and AMD signed 15-year agreements on July 28 covering about 530 MW with deliveries beginning in 2027, which strengthens the path to a named delivery. No rack has been delivered and no customer has reported qualification workloads.

ID
p68-helios-production-rack-q2-2027
Confidence
68%
Deadline
By June 30, 2027
Trigger
Customer or AMD announcement naming a delivered 72-GPU MI455X Helios production rack running customer qualification workloads.

Pending · Power

A second US multi-gigawatt AI campus publicly commits to at least 250 MW of utility-directed peak curtailment by January 31, 2027.

Interim: two large campuses were announced in-window — Paducah and Meta's 1 GW El Paso site — and neither disclosed curtailment terms or utility dispatch rights. The Kentucky PSC filing is the most likely place such terms would surface.

ID
p69-camellia-curtailment-template-jan31
Confidence
49%
Deadline
By January 31, 2027
Trigger
Utility, developer, or customer filing naming a second campus, curtailment amount of at least 250 MW, and utility dispatch rights.

Pending · Software

An independent benchmark finds Gemini 3.6 Flash at least 12% cheaper per completed agentic task than Gemini 3.5 Flash by August 31, 2026.

Interim: no independent completed-task harness result has been published. OpenAI's Luna cut changes the comparison set that evaluators are likely to prioritize, which may reduce the chance anyone runs this specific Gemini-to-Gemini test before the deadline.

ID
p70-flash-task-cost-aug31
Confidence
66%
Deadline
By August 31, 2026
Trigger
Independent published harness results comparing total completed-task cost on the same agentic task set and reporting at least 12% savings.

Watchlist

Aug 2

EU AI Act Article 50 transparency obligations become enforceable

Interaction disclosure, deepfake labelling and public-interest text labelling apply with fines up to €15M or 3% of worldwide turnover. Only machine-readable marking for systems placed on the market before August 2 gets a grace period, running to December 2.

Aug 11

CoreWeave Q2 2026 results

The first neocloud print since the $99.4B March backlog and the cleanest available read on whether contracted backlog converts to revenue at the pace the category's financing assumes.

Aug 26

NVIDIA Q2 FY27 results, and GitHub Copilot default-model enablement

Nvidia's guide is the first hard read on the Rubin ramp against AMD's newly contracted 530 MW. The same date is when Copilot's new default model turns on for organizations that have not set an explicit opt-out, which is an administrative deadline rather than a product choice.

By Aug 15

Independent completed-task cost for GPT-5.6 Luna

List price fell 80%, but nothing yet measures tokens consumed, retry rates or completion quality at the cheap tier. The entire routing case rests on evidence that does not exist yet.

By Sept 30

Kentucky PSC filing for the Paducah power service agreement

The campus is contingent on regulatory approval of its power agreement. The filing will show whether dedicated generation for AI load clears on ratepayer-neutral terms, and whether curtailment rights are part of the bargain.

H2 2026

2027 HBM volume and price settlements

SK hynix says 2027 volumes and prices are under negotiation with roughly ten customers now. Those contracts will set accelerator delivery schedules more tightly than any published GPU roadmap.

Changelog

  • The networking hypothesis moves to strained. Equinix posted record interconnection adds while interconnection revenue grew about 9% against roughly 16% total revenue growth — the specific pattern named on the thesis page as the falsification test. One quarter is not the two required for a revision, so the hypothesis is flagged rather than rewritten.
  • The house measurement switches from a model-substitution scenario to a filing-derived tier-spread calculation: the cost distance between the cheapest and most expensive GPT-5.6 tier on a fixed 1:4 workload.
  • Kimi K3, Inkling-Small and DeepSeek V4-Flash-0731 enter the LLM Evolutionary Tree through this week's Model Pulse tree delta; K3 is the first entry whose open-weight status carries a revenue-tiered license condition.
  • No AI Shockwave Timeline or Market Reference Architecture change this week. Financing structures and harness technique both moved materially, but no dated event reset a frontier assumption.
  • The capital-flow table raises hyperscaler-hosted capital in to ~$250B and the revenue proxy to ~$70B on Amazon's guidance increase, Meta's raised floor, and the first disclosure of Azure above $100B annual. The reported Nvidia-OpenAI backstop is deliberately excluded from capital in until definitive documentation exists.