Skip to content

AI Market · Forecast desk · Live · software

44%

Anthropic or a named enterprise customer publishes an independently attributed Skills API production case study with a stated manual baseline and a measured post-Skills outcome (time, cost, or completion), by October 31, 2026.

Probability history

0%25%50%75%100%08-2209-06deadline

Called at 58% on August 22, 2026. Repriced 1 time; now 44%. 55 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)58%
Reference class (n=29)50%
Recalibrated53%
Feature model34%
Council42%
Final44%

Drivers

  • Reference class qtype:disclosure runs 0.50 (n=29) vs stated 0.58.
  • Hierarchical calibration moves stated 0.58 to 0.53 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.34: conjuncts=1, named disclosure=True, horizon=70d.
  • Council of 5 at 0.42.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

43%

grid and power economist · openai/gpt-5.6-sol

The 55-day window permits a marketing publication, and Anthropic routinely has named enterprise relationships. But the hit is narrower than a generic customer story: it must specifically concern Skills and provide customer-attributed outcomes. Product announcements, anonymous examples, quotes without outcomes, or Anthropic-only anecdotes fail under the ambiguity rule. The software and disclosure reference classes center near 0.45–0.50, while the forecaster’s 0.50–0.59 band has materially underperformed. With no supplied evidence of a scheduled case study or customer commitment, 0.58 is too high.

43%

semiconductor supply chain analyst · openai/gpt-5.6-sol

The 55-day window is long enough for a product-marketing case study, and Anthropic has incentives to demonstrate enterprise adoption of Skills. However, the resolution bar is materially narrower than a launch post, testimonial, or Anthropic-authored usage anecdote: it requires a named customer and independently attributed outcomes. The supplied evidence contains no scheduled publication, identified customer, or imminent joint announcement. Relevant software and disclosure reference classes center near 0.45–0.50, while the forecaster’s 0.50–0.59 band has historically underperformed its stated confidence. Strict adverse ambiguity further lowers the estimate below 0.58.

45%

capital markets credit analyst · openai/gpt-5.6-sol

The software and disclosure reference classes center near 0.45–0.50, while the stronger 31–90-day rate is less informative about this unusually specific content requirement. A qualifying customer story is plausible, but publication timing is discretionary rather than filing-driven. The rule excludes product announcements, demos, unnamed adoption claims, and Anthropic-only anecdotes; obtaining a named enterprise’s approval and attributable outcomes can take longer than 55 days. With no supplied evidence of a scheduled Skills customer proof point, 0.58 appears optimistic, especially given the forecaster’s weak realized rate in that confidence band.

37%

superforecaster statistician · openai/gpt-5.6-sol

I use the lens:software base rate of 0.452, cross-checked against the 0.500 disclosure base rate. The 55-day window helps, but resolution requires more than product promotion: a named enterprise must independently attribute outcomes specifically to Skills. Case-study approval and measurable-outcome disclosure often lag adoption, and ambiguity counts as failure. The forecaster’s 0.50–0.59 predictions realized only 0.367 overall, indicating material overconfidence around the stated 0.58. These factors lower my estimate below the software base rate.

44%

geopolitics regulation analyst · openai/gpt-5.6-sol

The 55-day window is adequate for a product-marketing publication, and enterprise AI vendors frequently coordinate customer stories with fall announcements. However, the rule is narrower than a routine testimonial: it requires a named customer, outcomes specifically attributable to Skills, and publication by Anthropic or that customer. Generic adoption quotes, launch-partner mentions, or Anthropic-authored anecdotes would fail. The most relevant ledger classes cluster around 0.45–0.50, while the forecaster’s 0.50–0.59 band has materially underperformed. Strict ambiguity treatment and the limited window therefore put this below even odds.

Evidence the engine used

direct · undated

Customer Stories | Claude by Anthropic

Cited by grid-and-power-economist: The 55-day window permits a marketing publication, and Anthropic routinely has named enterprise relationships. But the hit is narrower than a generic customer story: it must specifically concern Skills and provide customer-attributed outcomes. Product announcements