Longest METR 50% time horizon for a public modelno data
— hours
on track < 40 · off >= 60manual ↗
connector returned no usable reading · checked 2026-09-07
Radar · AGI and capabilities · T2 · 2029 · WARNING
METR does not publish a 50%-success time-horizon point estimate of 160 work-hours (one working month) or more for any publicly released model on its Time Horizons page on or before 2028-12-31.
A month-long autonomous horizon is the rung where agents stop being tools and become staff. Everything downstream, including the labor-market rung and the capacity to run always-on agents per employee, depends on whether this is reached in 2028 or measured only in retrospect. The thesis bets that measurement lags capability, which matters for how anyone reads AGI timelines.
Registered at 65% on September 8, 2026. Engine repriced 2 times; now 56%.
Dated rungs. Each is scored on its own; the thesis does not get credit for the ladder until the rungs land.
filled bar · my probabilityhollow dot · engineamber date · due, awaiting adjudication
m1 · 2027-06-30 · 60% · METR publishes a 50% time horizon of at least 40 hours for a public model.
m2 · 2027-12-31 · 50% · METR publishes a task suite with human-baseline tasks longer than 80 hours.
m3 · 2028-12-31 · 35% · METR publishes a 50% time horizon of at least 160 hours for a public model (rung 3 achieved).
Registered thresholds. Status is computed from the latest public reading.
Longest METR 50% time horizon for a public modelno data
— hours
on track < 40 · off >= 60manual ↗
connector returned no usable reading · checked 2026-09-07
Longest human-baseline task in METR's suiteno data
— hours
on track < 80 · off >= 200manual ↗
connector returned no usable reading · checked 2026-09-07
Would raise my number
Would cut it
6 personas on one model (openai/gpt-5.6-sol); other vendors were unavailable this run, so read this as one model's persona-averaged view.
47%
grid-and-power-planner · openai/gpt-5.6-sol
I use rapidly iterated AI benchmarks facing measurement saturation as the reference class. Moving from 17.4 to 160 hours requires about 3.2 doublings; the quoted 4–7-month trend makes underlying capability by end-2028 plausible. Certification is harder: METR must add sufficiently long tasks, obtain costly human baselines, evaluate a public model, deem the result reliable and in-range, and publish the point estimate. A longer task suite alone does not resolve MISS. With no new indicator data, I assign certification hazards of roughly 0.03 in late 2026, 0.20 in 2027, and 0.38 in 2028. Compounding, plus about 0.02 risk that publication ceases, leaves 0.47 for HIT.
54%
semiconductor-supply-chain-analyst · openai/gpt-5.6-sol
Starting from the supplied 17.4-hour estimate, 160 hours requires about 3.2 doublings. At METR’s cited 4–7-month pace, raw extrapolation crosses the threshold around late 2027 to mid-2028. Capability alone is insufficient, however: METR must extend its roughly 16-hour measurement ceiling, validate long-baseline tasks, test a public model, and publish an in-range 50% point estimate by the cutoff. Using METR benchmark-refresh cycles and capability extrapolations as the reference class, I assign conditional certification hazards near 0.18 through 2027 and 0.31 in 2028, compounded, plus a small adverse publication-discontinuation risk. Non-certification remains slightly more likely.
62%
enterprise-cio · openai/gpt-5.6-sol
I use the reference class of frontier-benchmark transitions: capability often outruns a benchmark, while redesign, validation, and publication take roughly 1–2 years. A 17.4-to-160-hour increase is about 3.2 doublings, so the quoted 4–7-month trend makes underlying capability plausible before 2029. But certification is conjunctive: METR must deploy longer-baseline tasks, obtain a reliable point estimate, test a publicly released model, and publish by the deadline. I assign about 0.35 cumulative probability to that sequence, plus roughly 0.03 to METR ceasing publication, yielding 0.62 for HIT. This is slightly below the forecaster’s 0.65 because over two budget years a suite expansion is quite feasible.
58%
credit-analyst · openai/gpt-5.6-sol
I use METR’s own frontier-model horizon progression and benchmark-refresh lag as the reference class. The supplied 4–7-month doubling trend can reach 160 hours before the cutoff, but certification additionally requires a substantially extended task suite, reliable in-range estimation, a public model, and timely publication. Treating qualifying certification as a rising hazard averaging roughly 0.20 annually over the remaining 2.3 years gives about 0.40 cumulative risk; a small probability that METR stops publishing also resolves against the statement. Missing indicator data provides no basis to raise confidence. The conjunction therefore makes HIT modestly more likely than MISS, but below the forecaster’s 0.65.
56%
geopolitics-policy-analyst · openai/gpt-5.6-sol
Using METR’s reported 4–7-month doubling trend as the capability reference class, moving from 17.4 to 160 hours requires about 3.2 doublings, implying a raw crossing around late 2027 to mid-2028. That weighs against the thesis. But resolution requires a conjunction: a public model, an upgraded suite capable of reliable month-scale measurement, a qualifying 50% point estimate, and publication by the cutoff. Benchmark redesign and long human baselines can lag capability by years; the current >16-hour unreliability warning is therefore material. With 27 months remaining, I put certification slightly below even odds. The indicators provide no update.
57%
superforecaster-statistician · openai/gpt-5.6-sol
Base rate: I use 0.45 for a fast-moving capability benchmark to publish a ninefold threshold crossing within 27 months when its fitted trend projects crossing inside that window (about 0.26 annual hazard). The 17.4-hour baseline and 4–7-month doubling trend increase certification risk. However, a MISS requires several timed legs: a substantially expanded task suite, a qualifying public model, a reliable 50%-success estimate, and METR publication by the deadline. The current ~16-hour reliability ceiling makes that measurement chain a major constraint. METR discontinuation adds a small adverse tail. With no new indicator data, 0.57 is below the forecaster’s 0.65.