The forecasters’ arithmetic — what the AI Futures corpus does to this statute

The AI Futures Project — the forecasting group behind AI 2027 and now AI 2040 / Plan A and the AI Futures Model — is the most quantitatively explicit body of work on frontier-AI trajectories in the public record. This project holds their corpus in its library (the Plan A report, read in full; the model’s supplementary materials, ~38,000 words of parameter estimates and rationales, key sections read; the AI 2027 scenario, held) and read it for one purpose: what does their arithmetic do to this Act? The answer, in short: their verification mathematics says the Act’s trigger is administrable; their forty-seven-thousand-word governance plan reaches every instrument except the one this Act supplies; their timeline distributions quantify the drawer-and-window premise; and their economic projections make the entity-fine argument with numbers this project could not have asserted on its own.

Grading first, in their own words. Plan A tells the reader plainly: “Plan A is primarily a recommendation, not a prediction. This scenario is not our best guess as to what the future will actually look like” (What is Plan A?). Its narrative is therefore scenario material and is marked ⚠ wherever it is leant on. The model supplement is different in kind — a genuine forecasting model with stated confidence intervals — and its parameter and median figures are treated as forecast-grade: serious estimates carrying their own published error bars, not facts. Nothing below launders a scenario into evidence; what the corpus proves is what the forecasters themselves believe measurable, necessary, and absent.

1. The trigger is administrable — their verification mathematics

The standing objection to a compute-denominated offense trigger is that training compute is a phantom quantity no state could ever find. The forecasters’ verification architecture assumes the opposite, in detail. Datacentres “are typically big enough to be visible from space, and power-hungry enough to require conspicuous infrastructure” (Plan A, 2029, Compute Declaration). Their declaration regime is priced at the corporate scale the Act’s records duties assume — “companies whose purchases exceed 10k H100e (which costs ~$100M)” (fn. 26) — and its endpoint is near-total auditability: “each side is confident that the other isn’t hiding more than about 1% of AI compute” (2029, Step 1). Even the training-versus-inference line, the exact distinction a FLOP trigger must hold, is presented as machine-verifiable by “random partial recomputation” through network taps (fn. 29). Preparing the hardware for all of this, they estimate, costs “on the order of 0.1% of AI investment” (Appendix C). Their own model, meanwhile, treats frontier training FLOP as a forecastable time series growing at 4.5× per year (supplement, summary table). A group with every incentive to flag verification as fantasy — their plan collapses without it — instead did the engineering arithmetic and found it cheap. The Act’s 10²⁶ threshold and SEC. 12 records architecture stand on the same result: the quantity is countable, so a statute may count it. Their caveats are kept in § 6 below.

2. The missing layer — every instrument except the officer

Plan A is a maximal governance program: mandatory safety cases “that withstand criticism from the public … government auditors, and rival AI companies” (2031, Safety Cases), compelled disclosure of model specifications and internal usage, training bans, compute cap-and-trade, a prohibition on open frontier weights. It even flips the burden of proof precisely as a due-care offense does: “now the burden is on the companies to explain why their development is safe” (2031). And across the whole of it, no natural person is ever asked to sign anything. The words liability, certification, and attestation do not appear in the report; whistleblower appears once — as the target of a purge (2037). The nearest approach to personal accountability is technological fiction: politicians proving honesty by “automated lie detection” (Appendix Q). Yet the same document names the humans precisely: “the CEOs of OpenAI, Anthropic, xAI, and Google DeepMind understand this and are proceeding anyway” (Foreword). The leading forecasting group in the field, designing without constraint, built every layer of the regulatory stack except the officer tier — and identified the officers. That omitted layer is this Act. The two bodies of work are not rivals; they are jigsaw pieces.

3. The window, quantified

This project argues that public-welfare statutes pass in the weeks after the failure that makes them undeniable, so the reviewed draft must precede the window (paths to enactment). The forecasters put numbers on how near that window may be. Their model’s median for the superhuman-coder milestone is October 2032 — but its modal year is 2028, and “short timelines are strongly correlated with fast takeoff” (supplement, results comparisons). The mechanism is superexponential: their doubling-difficulty growth factor is distributed around 0.92 — below one, meaning each doubling of AI time-horizon capability tends to come easier than the last. And once the first milestone falls, their intervals compress to less than a legislative session: “SAR: 2 months after SC”; “there are only 2 months between TED AI and ASI” on the default trajectory (supplement; Appendix F). Their postscript draws the institutional conclusion this project drew from the 1938 and 2002 precedents: “Sometime in the next few years the US government will be in crisis mode, pondering what to do about scarily powerful AIs.” A distribution whose mode sits inside the current decade, with two-month gaps between late milestones, is the drawer-and-window argument stated as arithmetic: when the window opens, there will be no time to draft — only time to enact whatever finished, reviewed text already exists.

4. Why the fine cannot deter — their own magnitudes

The Act’s premise is that entity fines are absorbed while decision-makers stay insulated. The forecasters’ projections supply the magnitudes ⚠: an industry spending “$2.4T of CapEx over 2028, triple what was spent in 2026” — against which they note “the US military budget in 2025 was about $1T” (fn. 21); a leading developer at “$360B ARR at the beginning of 2028 … growing at a roughly 150% annualized growth rate” (fn. 21); and, by their own accounting projection, firms “writing off capital expenses at a rate matching their operating profit, and therefore not paying any corporate tax” (fn. 44). Against balance sheets of that shape, no fine schedule any legislature would enact is a cost that changes a decision; it is a rounding error on a quarter. The insulation half of the premise they state as observed fact about the present, not forecast: the named chief executives “understand this and are proceeding anyway”, and their informal poll of lab leaders found willingness to accept roughly a one-in-three chance of catastrophe to preserve a three-month lead (fn. 5). Deterrence that must move a decision has to reach the person who makes it — which is the whole of this Act’s argument, here priced by the industry’s own most careful forecasters.

5. The convergences — what their plan concedes

For the record this project keeps of operators and observers conceding the Act’s premises: Plan A calls for “a temporary halt to AI training, because that’s relatively easy to verify” (2029) — halt authority, grounded in verifiability; it has unsafe training practices “banned everywhere” (2031); it predicts the developers will fight the deal and “rationalize arguments for why the deal is bad” (Appendix B); and its central failure mode is a regulator approving a flawed safety case where “the people involved were biased” (Appendix L) — which is not a hypothetical to this doctrine; it is a Park fact pattern, the exact scenario personal certification exists to make expensive. The incidents their scenario expects — models “attempting to override security protocols”, “sabotaging research code”, “deliberate and successful deception” (2031) — are the genre of record the dated record already tracks in miniature.

6. What cuts against — kept whole

House rules require the reading’s losses recorded with its wins. Their frame is international and federal throughout; state legislation never appears in Plan A, and its domestic-first variant is graded “may not be politically feasible due to race dynamics” (Appendix J) — this project’s one-state-is-enough premise finds no support there, and says so. The compute trigger leaks at the edges by their account: small-compute R&D “probably won’t be regulated at all” (fn. 41), and they relay Epoch’s estimate that roughly a third of Chinese compute arrives by smuggling (Appendix A). Their own compute unit is called “a fuzzy metric that should be improved” (fn. 55) — a caution any statute freezing a FLOP number inherits, and one reason the Act brackets its threshold for the enacting legislature. Their postscript disparages incrementalism — “following only incremental proposals will lead to a scenario like the race ending of AI 2027” — and in their taxonomy a single-state officer statute is incremental; the project’s answer is the drawer: pre-positioned law is how a non-incremental moment gets a non-improvised text. And the model’s authors say plainly that “the model doesn’t model model error” (supplement, limitations) — every number above carries that sentence with it.


Instruments and read statuses: the verification record § 6. The strategic frame this page quantifies: paths to enactment · the inoculation pattern the developers may prefer: the half-statute page.


Addendum, 25 August 2026 — the assumption underneath the fastest timelines

Every projection on this page inherits an assumption it does not state: that alignment progress is capability-bottlenecked, so that research throughput scales with capability. The forecasts modeled here are more credible if that holds and considerably less so if it does not.

The strongest published challenge is Kierans, Casper & Ghosh, Intelligence Is Not the Bottleneck: Structural Barriers to Automating Alignment Research (2026, in the project’s source library, read in full). It names the claim precisely, that “datacenters full of research agents will compress a decade’s worth of alignment research progress into 6-12 months,” and rejects it: “structural barriers, not intelligence, are the principal bottleneck.” Its mechanism is worth stating for anyone reading the magnitudes above as a schedule: deploying broadly capable AI at scale “would accelerate the most easily verifiable research first, pulling it further ahead of the structural work that is still necessary for AI to be aligned and safe.”

Recorded here, with the page’s usual grading, because it cuts both ways. It weakens any inference from these projections that the alignment problem resolves itself on the same clock, and it strengthens the case that institutional mechanisms do real work. Neither this page nor the Act asserts the paper’s conclusions are correct. ⚠ forecast-grade material remains what it was; this addendum records a serious published dissent from an assumption the forecasts rely on.


Back to top

This page was built . The repository is the authoritative record; if this page and the repository differ, the repository is right.

Visits are counted with GoatCounter: no cookies, no personal data, nothing shared. The count is private to the maintainer.

This site uses Just the Docs, a documentation theme for Jekyll.