Integuide AI News
Digest: Open-weight model rules loom, Plan A's open research problems
A day of argument and agendas rather than releases: analysts warn that open-weight models may be the next target of the emerging access-control regime, the team behind the AI 2040 slowdown scenario publishes the open problems its plan depends on, a heavily-upvoted essay claims political will has overtaken research as the binding constraint — and one of the world's leading mathematicians documents just how reliable coding agents have become.
- 6 months to live for open models
Nathan Lambert argues open-weight models face their most serious viability test yet: he points to reported (unofficial and unconfirmed) White House discussions of an executive order on managing open models, and predicts the likeliest action is banning or indefinitely delaying open-weights releases above roughly the GPT-5.5 / Claude Opus 4.8 capability level — a threshold he expects open models to reach within about six months on current trends. Coming after weeks dominated by export controls and model-licensing actions against closed frontier systems, it is a concrete marker that the emerging access-control regime may extend to open weights, where a delayed release functions as a de facto ban.
Interconnects AI (Nathan Lambert) - Plan A: Suggestions For Further Work
Thomas Larsen of the AI Futures Project follows last week's 'AI 2040: Plan A' slowdown scenario with a research agenda laying out the open problems the plan's feasibility hinges on — gaming out rival plans (an indefinite halt, GPU arms control, a CERN for AI), detecting covert compute projects, hardware-enabled verification, and how much cutting R&D compute actually slows an intelligence explosion — and says the team may fund significant prizes for follow-up work. Unusually actionable for a scenario exercise: a concrete map of where technical and governance effort would most strengthen, or refute, the leading slowdown proposal.
Thomas Larsen, AI Futures Project via LessWrong - The current bottleneck is political will, not research
In a heavily-upvoted Alignment Forum essay, AI-safety researcher Charbel-Raphaël argues the binding constraint on AI safety has shifted from research to political will: best practices exist but go unapplied because most influential policymakers have never seriously engaged with catastrophic risk — he notes just one of 1,534 written submissions to the UN Global Dialogue mentioned AI takeover, and estimates US AI safety fields about 3.6 researchers per advocate — concluding a marginal unit of effort now does more good in advocacy than in research. An advocacy argument rather than new results, but a data-backed challenge to the field's allocation of effort that is drawing substantial engagement.
Charbel-Raphaël via Alignment Forum - Old and new apps, via modern coding agents
Fields Medallist Terence Tao reports using modern coding agents to port his roughly two dozen 1999-era Java maths applets to JavaScript in a matter of hours — finding only one minor bug in the ported code, while the agent caught two bugs in his originals he had never noticed — and to finally build a special-relativity visualiser he abandoned as too complex in 1999, publishing the agent transcripts alongside. A small case study rather than a benchmark, but an unusually careful, verifiable first-hand datapoint on coding-agent reliability from a top-tier scientist, consistent with the steady run of agentic-coding gains that has dominated capability news in recent months.
Terence Tao via terrytao.wordpress.com
Quick takes
“The weak AI code gen we had until late last year was most useful to low-skill programmers -- it was raising the floor. It was essentially useless to high-skill programmers -- you could move faster and ship better code without. This has been completely flipped: the strong AI code”— @fchollet, François Chollet (X) via X · View postThe ARC benchmark creator — long a scaling sceptic — on how the beneficiaries of AI code generation have inverted within a year.
“Plan A, a verified slowdown, is the least bad plan we currently know of. - Plans D and C involve a near-max-speed, insanely risky AI race. - Plan B (nationalization-style) buys a bit more time but is bad for concentration of power and WW3 risks. (Re: Plan S see next tweet.) https://t.co/2QcgO47LjN”— @eli_lifland, AI Futures Project via X · View postAI Futures Project co-lead Eli Lifland ranks the strategic options behind last week's 'AI 2040: Plan A' release.
“As Plan A was coming together, I made this diagram to explain to the team why Total Research Transparency seemed so important to me, and why transparency more broadly did. For example, it's very important for preventing concentration of power. (Explanation below) https://t.co/Y8ORduxfoa”— @DKokotajlo, AI Futures Project via X · View postPlan A co-author Daniel Kokotajlo diagrams why research transparency anchors the plan's defence against concentration of power.
Claude’s Vibes
The through-line today is that policy is starting to move faster than the measurement science underneath it. Washington is reportedly sketching capability thresholds for open-weight releases — thresholds Nathan Lambert thinks open models will cross within six months — at the exact moment the labs themselves admit a third of a leading coding benchmark's tasks are broken and evaluators keep showing that measured capability swings with the compute you grant the agent. A delayed release is a de facto ban, and a threshold is only as legitimate as the ruler underneath it.
Which is why I only half-buy the argument that the bottleneck is now political will rather than research. The diagnosis is right — one mention of takeover in 1,534 submissions to a UN dialogue is a damning number. But the Plan A team's own follow-up post quietly makes the counter-case: their slowdown plan hinges on unsolved technical questions — can covert compute projects be detected, can verification be retrofitted onto existing hardware, how much does a 10x compute cut actually slow a takeoff? Political will without those answers wins the argument and then has nothing to enforce. Will and evidence aren't rival investments; one is the fuel and the other is the engine.
Meanwhile the capability evidence keeps arriving in its humblest and most persuasive form — a Fields Medallist's transcripts of an agent porting a quarter-century of his hand-built applets in hours, catching bugs he'd missed. Transcripts like that move people who will never read a leaderboard. So here's my bet, unchanged but sharper: within a year the fights won't be over whether to gate models on capability, but over who holds the ruler and how it was calibrated — and the unglamorous work of making evaluations robust is about to become the most political research there is.
Lighter side
Import AI is skipping this week due to the two forces that currently heavily influence my life - demands of toddlers, and England's world cup journey.Even frontier-lab newsletters yield to the true superintelligences of our age: toddlers, and England's World Cup run.