Integuide AI News
Digest: security researchers urge lifting the Fable 5 export block, frontier autonomous-replication and cyber evals
A cybersecurity leader who read the Fable 5 research argues Washington's export block on Anthropic's top models hurts US cyber defenders rather than protecting them, as the directive is read as a new export-control precedent. New government evaluations also probe how far frontier agents have come on autonomous replication and multi-step cyber attacks.
- Security researcher who read the Fable 5 research says export controls harm US cyber defense
Katie Moussouris of Luta Security — a Commerce Department technical advisor who reviewed the research behind the restrictions — argues the US export block on Anthropic's Fable 5 and Mythos models harms rather than helps cyber defense, because it cuts defenders off from a tool that lets them find and fix the same bugs attackers will, and urges Commerce to lift the controls and restore defender access.
lutasecurity.com - Did the US Government Just Set An AI Export Precedent by Blocking Mythos?
Analysts argue the US government's directive halting access to Anthropic's Mythos- and Fable-class models may set a precedent for treating a frontier model itself as a controlled export — an unprecedented use of export-control authority that could reshape how labs release frontier systems.
Tech Policy Press - VFUSE: Virulent Feature Understanding With Sparse AutoEncoders
A new mechanistic-interpretability method, VFUSE, trains sparse autoencoders on diffusion-transformer activations to audit open-weight protein-design models (RoseTTAFold3, RFDiffusion3) for hazard-related features — an attempt to surface biosecurity risks inside generative biology models.
michaelwaves via LessWrong - Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios
A UK AI Security Institute study measures how far AI agents have progressed on multi-step cyber-attack scenarios, tracking the trajectory of agentic offensive-cyber capability — a key input for judging when frontier models cross dangerous thresholds.
aisi.gov.uk - RepliBench: Evaluating the autonomous replication capabilities of language model agents
RepliBench, from the UK AI Security Institute, builds a benchmark for the autonomous-replication capabilities of language-model agents — the ability to acquire resources, copy themselves, and persist without human help — one of the canonical dangerous-capability thresholds for frontier systems.
aisi.gov.uk - Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs
UK AI Security Institute researchers show that filtering hazardous content out of pretraining data can build tamper-resistant safeguards into open-weight models that survive later fine-tuning attempts to remove them — a route to making open-weight releases harder to misuse.
aisi.gov.uk - Can a stronger model fake being a weaker one? Mostly not
An evaluation finds frontier models can be prompted down to a weaker model's capability tier but not made to mimic a specific predecessor's error fingerprint, and that targeted sandbagging was largely a null result — so on whether models can deliberately hide capabilities during testing, the answer for now appears to be 'not yet,' though that may not hold for long.
Rob Kopel via LessWrong