Skip to content

AI Market · Forecast desk · Live · software

39%

An independent party reproduces at least a 15-point ARC-AGI-3 improvement from harness memory and compaction settings alone, holding model weights fixed, by October 31, 2026.

Probability history

0%25%50%75%100%08-0109-06deadline

Called at 44% on August 1, 2026. Repriced 1 time; now 39%. 55 days to deadline.

How the number was built

Final method: stacked-fixed-weights (resolved n=57 < 80). Each rung is an independent estimate; the final is a stacked combination.

Stated (issue date)44%
Reference class (n=19)65%
Recalibrated52%
Feature model37%
Council33%
Final39%

Drivers

  • Reference class qtype:threshold runs 0.65 (n=19) vs stated 0.44.
  • Hierarchical calibration moves stated 0.44 to 0.52 (lens software, kappa_stated=24.0, kappa_lens=200.0).
  • Feature model at 0.37: conjuncts=1, named disclosure=False, horizon=91d.
  • Council of 5 at 0.33.

Cautions

  • all 5 votes came from one vendor (openai); extremization disabled, persona diversity only

Council

5 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view, aggregated by log odds mean without extremization.

34%

grid and power economist · openai/gpt-5.6-sol

Software results can be replicated quickly, but this requires three things within 55 days: access to the memory method, an independent implementation, and a public numeric ARC-AGI-3 gain of at least 15 points. Benchmark-harness improvements are especially sensitive to prompts, budgets, model versions, and scoring details. A replication showing a smaller gain, lacking comparable baselines, or remaining an informal lab report would miss under the strict rule. The software reference rate supports a meaningful chance, but the large threshold, short remaining window, and documentation requirement put this below the forecaster’s 0.44.

31%

semiconductor supply chain analyst · openai/gpt-5.6-sol

The hit requires three things within 55 days: an independent replication, the same ARC-AGI-3 comparison, and a documented gain of at least 15 points. Software replication can move quickly when code and artifacts are complete, but ARC results are unusually sensitive to model, harness, budget, and baseline choices; a 15-point delta is also large. The supplied digest contains no qualifying independent result, and the ambiguity policy penalizes incomplete or nonnumeric reports. I therefore place this below the forecaster’s earlier 0.44 despite the broader software reference rate.

36%

capital markets credit analyst · openai/gpt-5.6-sol

The remaining 55 days are enough for implementation and evaluation if harness memory is clearly specified and benchmark access is straightforward. However, resolution requires more than a promising self-report: an independent party must publish a baseline-controlled numeric gain meeting the full 15-point threshold. The supplied digest shows active model releases but no direct evidence of an ongoing reproduction or scheduled ARC Prize validation. Reporting and attribution are additional failure modes, and ambiguity resolves against the event. The software reference rate supports a material chance, but the exact threshold and independence requirements warrant a discount.

26%

superforecaster statistician · openai/gpt-5.6-sol

I use the software reference-class base rate of 0.452. Conditioning on no qualifying public reproduction yet and only 55 days remaining lowers the chance materially. Resolution also requires a conjunction: an independent party must reproduce the effect, reach the unusually large 15-point threshold, report numbers, and do so by the hard date; ambiguity counts as failure. General open-model activity supports some experimentation but is not evidence specific to harness memory or ARC-AGI-3. The forecaster’s low-confidence calibration bin has only one case, while broader results show some overconfidence, so 0.44 receives no upward calibration adjustment.

38%

geopolitics regulation analyst · openai/gpt-5.6-sol

The software reference rate is 0.452, but only 55 days remain and the supplied digest contains no resolution-grade independent result. Reproducing a harness-level memory intervention can be faster than training a model, and a large gain is plausible on an interactive benchmark where state retention matters. Still, the rule requires both independence and an explicitly reported improvement of at least 15 points; code replication without comparable scoring, a lab-associated report, or ambiguous baselines all miss. The elapsed time since the 0.44 forecast therefore warrants a modest downward update.

Evidence the engine used

direct · undated

ARC Prize

Cited by grid-and-power-economist: Software results can be replicated quickly, but this requires three things within 55 days: access to the memory method, an independent implementation, and a public numeric ARC-AGI-3 gain of at least 15 points. Benchmark-harness improvements are especially sensitive