Integuide AI News
Digest: Washington weighs frontier-AI rules after Fable takedown, UK publishes 2030 AI scenarios
Governance dominates today: Washington debates new frontier-AI rules in the wake of the unprecedented Fable model takedown, and the UK publishes refreshed 2030 scenarios for long-term planning. New interpretability work questions the reliability of a popular class of latent-space safety defenses.
- Is Fable the wakeup call DC needed?
Transformer News's weekly roundup tracks Washington's policy response to the Fable export-control episode, reporting a legislative preemption package, G7 AI discussions, and new senior OpenAI hires, and framing the unprecedented government model takedown as a possible catalyst for US frontier-AI rules.
Shakeel Hashim via Transformer News - UK government publishes refreshed AI 2030 scenarios to help plan for the future of AI
The UK Government Office for Science published a refreshed set of AI 2030 scenarios, structured plausible futures meant to help officials stress-test strategies against an uncertain frontier, updating its earlier 2023 set to reflect rapid shifts in AI capabilities and investment. The publication signals how a major government is framing long-term planning around advanced-AI risks and opportunities.
gov.uk - SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Researchers show that sparse-autoencoder interventions can be unreliable: clamping an "unsafe" SAE feature may only suppress harmful behavior temporarily, with the behavior recoverable afterward. The finding questions a growing class of latent-space safety defenses that assume identified features are dependable control handles.
Mingyue Cui et al. via arXiv