Skip to content

AI Market · Forecast desk · Live · software

44%

An independent evaluator publishes completed-task cost showing GPT-5.6 Luna at least 60% cheaper per completed agentic task than GPT-5.6 Terra by September 30, 2026.

Probability history

0%25%50%75%100%08-0109-06deadline

Called at 62% on August 1, 2026. Repriced 1 time; now 44%. 24 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)62%
Reference class (n=8)39%
Recalibrated54%
Feature model61%
Council32%
Final44%

Drivers

  • Reference class qtype:timing runs 0.39 (n=8) vs stated 0.62.
  • Hierarchical calibration moves stated 0.62 to 0.54 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.61: conjuncts=1, named disclosure=False, horizon=60d.
  • Council of 5 at 0.32.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

32%

grid and power economist · openai/gpt-5.6-sol

Only 24 days remain, and the digest provides no evidence that GPT-5.6 Luna is publicly available or undergoing independent evaluation. Resolution requires a completed-task cost comparison, a stricter and less common artifact than token pricing or vendor benchmark claims. The software and timing reference classes are only 0.452 and 0.389, while the forecaster’s 0.60–0.69 predictions have realized below stated confidence. An evaluator could test a newly released API quickly, but release, testing, publication, and beating the specified bar must all occur by September 30; ambiguity counts as a miss.

31%

semiconductor supply chain analyst · openai/gpt-5.6-sol

Only 24 days remain, and the required evidence is narrower than a model release or vendor pricing: an independent evaluator must publish a completed-task cost comparison and clear the specified bar. Evaluators can benchmark API models within days, but completed-task metrics require repeat runs and are less routinely published than token prices or quality scores. The event was forecast on August 1 and remains unmet by September 6, which materially reduces a timing forecast’s conditional probability. The forecaster’s software and timing base rates are also weak, with overprediction in the 0.60–0.69 band. Ambiguity resolving against the forecast further penalizes borderline comparisons.

32%

capital markets credit analyst · openai/gpt-5.6-sol

Only 24 days remain, and the supplied digest shows no qualifying independent completed-task cost result. The resolution is stricter than ordinary model pricing or benchmark coverage: GPT-5.6 Luna must be evaluated, the metric must represent cost per completed task, and the result must beat the specified bar. Independent evaluators can add prominent models quickly, but testing and publishing a methodology-specific comparison creates timing risk. The forecaster’s software and timing reference classes are below its stated 0.62, and its 0.60–0.69 forecasts have been overconfident. Ambiguity resolving against the forecaster further reduces the chance that a near-match qualifies.

24%

superforecaster statistician · openai/gpt-5.6-sol

I use the 0.389 timing-question base rate, cross-checked against the 0.452 software rate. This requires several linked developments within 24 days: GPT-5.6 Luna must be evaluable, an independent evaluator must run a completed-task cost comparison, the result must beat the bar, and publication must occur by the deadline. The digest supplies no model- or evaluator-specific evidence supporting those steps. The strict publication requirement and ambiguity-against policy materially reduce the estimate. The forecaster’s 0.60–0.69 band has realized only 0.556, confirming systematic overconfidence and making 0.62 too high.

42%

geopolitics regulation analyst · openai/gpt-5.6-sol

The 24-day window is feasible for an evaluator update, but the rule requires completed-task cost—not token pricing, vendor claims, or a general benchmark score—and the result must beat a specific bar. That creates both publication and performance hurdles. The forecaster’s timing and software reference classes are below 0.5, while its 0.60–0.69 forecasts have realized below their stated confidence. No supplied evidence confirms that GPT-5.6 Luna is already evaluator-accessible or that a qualifying cost study is scheduled. Ambiguity resolving against the forecast further reduces the chance of a marginal comparison counting.

Evidence the engine used

direct · undated

AI Model &amp; API Providers Analysis | Artificial Analysis

Cited by grid-and-power-economist: Only 24 days remain, and the digest provides no evidence that GPT-5.6 Luna is publicly available or undergoing independent evaluation. Resolution requires a completed-task cost comparison, a stricter and less common artifact than token pricing or vendor benchmark c