FrontierMath open problems marked solved by AIno data
— count
on track >= 4 · off <= 2epoch ↗
connector returned no usable reading · checked 2026-09-07
Radar · AGI and capabilities · T2 · 2029 · CALL
Epoch AI's FrontierMath Open Problems page credits AI systems with verified solutions to at least 10 of the listed unsolved research-level problems on or before 2028-12-31.
Multiple-choice science benchmarks are exhausted; GPQA Diamond fell before this book was written. Solving research problems nobody has solved is the first rung that cannot be reached by memorization. Ten verified solutions would mean models are producing new mathematics at a rate that changes how research labs staff.
Registered at 65% on September 8, 2026. Engine repriced 2 times; now 42%.
Dated rungs. Each is scored on its own; the thesis does not get credit for the ladder until the rungs land.
filled bar · my probabilityhollow dot · engineamber date · due, awaiting adjudication
m1 · 2027-06-30 · 75% · Epoch credits AI with at least 4 verified open-problem solutions.
m2 · 2027-12-31 · 70% · Epoch credits AI with at least 6 verified open-problem solutions.
m3 · 2028-12-31 · 65% · Epoch credits AI with at least 10 verified open-problem solutions.
Registered thresholds. Status is computed from the latest public reading.
FrontierMath open problems marked solved by AIno data
— count
on track >= 4 · off <= 2epoch ↗
connector returned no usable reading · checked 2026-09-07
Would raise my number
Would cut it
6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.
43%
grid-and-power-planner · openai/gpt-5.6-sol
Starting from 2, this requires eight additional verified solves in 28 months, about 3.4 per year. My reference class is early research-benchmark progress: capability gains are lumpy, while verification and publication lag model performance. A Poisson-style extrapolation from roughly two solves during 2026 lands near, but below, the threshold. Tier-4 saturation is supportive, yet transfer from bounded benchmark questions to genuinely open problems is weak, and the remaining set may be selected for resistance. The literal rule also requires Epoch—not merely papers or announcements—to mark ten qualifying, non-human-assisted solutions; page retirement adds another failure mode. Thus 0.65 appears optimistic.
41%
semiconductor-supply-chain-analyst · openai/gpt-5.6-sol
The reference class is sparse early successes on newly launched, adversarial research benchmarks, where closed-benchmark saturation transfers poorly to genuinely open problems. From 2 solved, the thesis needs 8 additional verified solves in roughly 28 months—about 3.4 annually—with harder-tail depletion and review lag. Capability acceleration and a pool of roughly 48 remaining problems make that plausible, but the strict rule excludes partial or human-assisted work. Resolution also conjunctively requires Epoch’s designation, qualifying solution dates, and continued page tracking. I estimate about 0.44 conditional on continued tracking, reduced to 0.41 overall, below the stated 0.65.
38%
enterprise-cio · openai/gpt-5.6-sol
Starting from 2, the thesis requires 8 additional verified solves in roughly 28 months, about 3.4 per year. My reference class is transitions from benchmark saturation to independently verified research results: benchmark scores can improve quickly, but novel proofs face search, evaluation, attribution, and publication lags. A count model centered on 2–3 accepted solves in 2027 and 3–4 in 2028 puts the threshold in the upper tail even with acceleration. Strict exclusion of human-assisted and under-review work, plus page-continuity risk, lowers the odds further. Tier 4 saturation is supportive but not a reliable linear predictor of autonomous open-problem closure.
44%
credit-analyst · openai/gpt-5.6-sol
Using sparse, independently verified research-mathematics breakthroughs as the reference class, the target requires eight additional credited solutions in about 2.3 years, or roughly 3.5 annually. A Poisson-style baseline near three annual additions puts the threshold below even odds. Rapid capability gains create substantial upside, but benchmark saturation does not reliably translate into autonomous theorem solving. Correlated problem difficulty, verification delays, exclusion of human-assisted work, and the page-continuity condition reduce resolution probability. With no indicator update beyond the baseline count of two, I place the claim below the forecaster’s 0.65.
47%
geopolitics-policy-analyst · openai/gpt-5.6-sol
Starting from 2, the thesis needs 8 additional verified solutions in about 28 months. A Poisson-style reference class for rare, independently verified research breakthroughs, calibrated loosely to two arrivals during 2026, puts the required pace near the upper-middle of the plausible range. Rapid benchmark gains support an accelerating arrival rate, but Tier 4 saturation is weaker evidence for solving genuinely open problems. Epoch’s verification delay, exclusion of human-assisted or under-review work, and the requirement that the page remain active all reduce resolution probability. The missing indicator update adds uncertainty rather than positive evidence. This makes 0.65 too optimistic under strict resolution.
38%
superforecaster-statistician · openai/gpt-5.6-sol
Base rate: I use a 0.30 chance that a curated set of research-level open problems reaches 20% verified autonomous completion within roughly 28 months after only two initial machine solves. The target needs eight additional successes among about 48, while review latency, exclusion of human-assisted work, and page-retirement risk all reduce resolution odds. Rapid FrontierMath Tier 4 saturation is positive, but benchmark performance is a weaker reference class than novel, author-verified research, and the connector supplies no confirming trajectory. Capability acceleration raises my estimate above the base rate, but 0.65 underweights threshold and verification risk.