Skip to content

Radar · AGI and capabilities · T1 · 2027 · CALL

SWE-bench Pro public set passes 90% by 2027

A model posts a resolved rate of at least 90.0% on the SWE-bench Pro public leaderboard (Scale AI, public problem set) on or before 2027-12-31.

CALLupindicators pendingregistered 2026-09-08Scale AI

ClaimA model posts a resolved rate of at least 90.0% on the SWE-bench Pro public leaderboard (Scale AI, public problem set) on or before 2027-12-31.
Consensus (implied)55%implied from Metaculus, "When will AIs hit 90% on SWE-bench Verified" (resolved 2026-04-07) and OpenAI, "Why SWE-bench Verified no longer measures frontier coding capabilities" · 2026-04-07
Distance+1.19log-odds · clearly above consensus
My confidence80%80% CI 6590%
Engine68%-12 pts vs me · stacked-fixed-weights
Falsifies ifOn 2027-12-31 no entry on the SWE-bench Pro public leaderboard shows a resolved rate at or above 90.0%.
HorizonDecember 31, 2027480 days · by end-2027 · Brier-scored

Why it matters

This is the first rung of the capability ladder: can a model close real, uncontaminated repository issues at the rate a strong senior engineer would. Passing it moves the enterprise question from whether coding agents work to how much software headcount they displace, which is what rung four measures.

Probability over time

0%25%50%75%100%09-0709-0709-08deadline

Registered at 80% on September 8, 2026. Engine repriced 2 times; now 68%.

Leading indicators

Registered thresholds. Status is computed from the latest public reading.

Top SWE-bench Pro public resolved rateno data

percent

8683no data yet

on track >= 86 · off < 83manual

connector returned no usable reading · checked 2026-09-07

Artificial Analysis Terminal-Bench 2.1 top scoreno data

percent

8578no data yet

on track >= 85 · off < 78artificial-analysis

connector returned no usable reading · checked 2026-09-07

What would move me

Would raise my number

  • Top public score exceeds 86% before April 2027.
  • Scale AI reports that remaining failures are concentrated in environment setup rather than code reasoning.
  • Two labs publish SWE-bench Pro scores as headline model-card metrics.

Would cut it

  • Scale AI publishes an audit showing a hard ceiling from ambiguous tasks below 90%.
  • Top score stalls within 2 points for two consecutive quarters.
  • Labs shift reporting to a successor benchmark and stop submitting to SWE-bench Pro.

Council · 71%

6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.

74%

grid-and-power-planner · openai/gpt-5.6-sol

Using the reference class of public coding benchmarks already above 80%, I estimate a 0.79 chance of gaining the remaining 9.7 points within 16 months. The supplied rise from roughly 23% to 80.3% supports rapid improvement, and accepted scaffolds plus a public problem set increase benchmark-specific optimization. However, progress usually slows near saturation because remaining tasks are unusually difficult or evaluation-sensitive. I apply about a 0.94 probability that a qualifying result is actually dated, listed under Scale’s standard evaluation, and the leaderboard or same-metric successor remains available. Multiplying gives approximately 0.74, below the forecaster’s 0.80.

72%

semiconductor-supply-chain-analyst · openai/gpt-5.6-sol

My reference class is threshold crossings on rapidly improving but saturating software-agent benchmarks, anchored to the 0.65 structural prior. Moving from 80.3% to 90.0% requires roughly halving the remaining error rate over 16 months. New model generations, inference scaling, and accepted scaffolds make that plausible, while the public set permits optimization and possible contamination. However, the hardest residual tasks can create a long tail, and the observed one-year gain should not be extrapolated linearly. I estimate about 0.77 for a qualifying technical result, discounted for leaderboard retirement, methodology changes, submission delays, or failure to list the result by the deadline.

68%

enterprise-cio · openai/gpt-5.6-sol

My reference class is frontier software-engineering benchmarks after broad lab adoption: headline scores often saturate within 12–24 months, especially when public tasks and accepted scaffolds permit targeted optimization. Moving from 80.3 to 90.0 requires eliminating roughly half the remaining failures, harder than a straight-line extrapolation suggests because of irreducible, flaky, or environment-sensitive tasks. The 16-month window still allows multiple model and agent generations. I discount for diminishing returns, absent confirming indicators, leaderboard submission discretion, and the conjunctive requirement that Scale continue listing the qualifying result. This is above the 0.65 structural prior but below the stated 0.80.

68%

credit-analyst · openai/gpt-5.6-sol

My reference class is public coding benchmarks after frontier-lab adoption: rapid early gains followed by sharply diminishing returns near saturation. From 80.3%, reaching 90% requires halving the unresolved-task rate, materially harder than adding 9.7 linear points. Still, 16 months, public-task optimization, and accepted tool scaffolds make crossing more likely than not. I estimate a 0.62 chance of crossing in the next 12 months and a 0.24 conditional chance in the final four, then discount roughly 0.03 for leaderboard retirement, methodology, or qualifying-entry risks. The 0.65 structural prior is consistent; without fresh indicators, 0.80 is too confident.

72%

geopolitics-policy-analyst · openai/gpt-5.6-sol

Reference class: frontier software-engineering benchmarks after broad lab adoption, which typically show rapid early gains followed by diminishing returns near saturation. Advancing from 80.3 to 90.0 within roughly 16 months is plausible because accepted tool scaffolds, public tasks, and competition among several labs create repeated attempts. The prior year’s unusually large gain warrants an uplift from the 0.65 structural prior. I discount for increasingly difficult residual tasks, evaluation variance, possible rule tightening, and literal-resolution risk from leaderboard retirement or metric replacement. A roughly 0.62 annual success hazard compounded across 1.32 years yields about 0.72, below the forecaster’s 0.80.

70%

superforecaster-statistician · openai/gpt-5.6-sol

Base rate: 0.65 from the supplied reference class for threshold questions. I adjust upward because the stated leaderboard gain from roughly 23% to 80.3% in one year leaves only 9.7 points over nearly 16 months, and accepted scaffolds broaden the paths to success. I limit that adjustment because the remaining improvement requires about halving the current error rate, benchmark progress commonly slows near saturation, and both qualifying performance and timely listing on the named public leaderboard are required. Retirement without a same-metric successor also causes failure. With no fresh indicator data, the forecaster’s 0.80 relies too heavily on straight-line extrapolation.

Engine prior

65% from reference-class:qtype:threshold. ledger base rate, n=19, horizon 480d; the ledger has no multi-year history