Integuide AI News

16 Jun 2026

Digest: security researchers urge lifting the Fable 5 export block, frontier autonomous-replication and cyber evals

A cybersecurity leader who read the Fable 5 research argues Washington's export block on Anthropic's top models hurts US cyber defenders rather than protecting them, as the directive is read as a new export-control precedent. New government evaluations also probe how far frontier agents have come on autonomous replication and multi-step cyber attacks.

  1. Security researcher who read the Fable 5 research says export controls harm US cyber defense

    Katie Moussouris of Luta Security — a Commerce Department technical advisor who reviewed the research behind the restrictions — argues the US export block on Anthropic's Fable 5 and Mythos models harms rather than helps cyber defense, because it cuts defenders off from a tool that lets them find and fix the same bugs attackers will, and urges Commerce to lift the controls and restore defender access.

    lutasecurity.com
  2. Did the US Government Just Set An AI Export Precedent by Blocking Mythos?

    Analysts argue the US government's directive halting access to Anthropic's Mythos- and Fable-class models may set a precedent for treating a frontier model itself as a controlled export — an unprecedented use of export-control authority that could reshape how labs release frontier systems.

    Tech Policy Press
  3. VFUSE: Virulent Feature Understanding With Sparse AutoEncoders

    A new mechanistic-interpretability method, VFUSE, trains sparse autoencoders on diffusion-transformer activations to audit open-weight protein-design models (RoseTTAFold3, RFDiffusion3) for hazard-related features — an attempt to surface biosecurity risks inside generative biology models.

    michaelwaves via LessWrong
  4. Measuring AI Agents’ Progress on Multi-Step Cyber Attack Scenarios

    A UK AI Security Institute study measures how far AI agents have progressed on multi-step cyber-attack scenarios, tracking the trajectory of agentic offensive-cyber capability — a key input for judging when frontier models cross dangerous thresholds.

    aisi.gov.uk
  5. RepliBench: Evaluating the autonomous replication capabilities of language model agents

    RepliBench, from the UK AI Security Institute, builds a benchmark for the autonomous-replication capabilities of language-model agents — the ability to acquire resources, copy themselves, and persist without human help — one of the canonical dangerous-capability thresholds for frontier systems.

    aisi.gov.uk
  6. Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs

    UK AI Security Institute researchers show that filtering hazardous content out of pretraining data can build tamper-resistant safeguards into open-weight models that survive later fine-tuning attempts to remove them — a route to making open-weight releases harder to misuse.

    aisi.gov.uk
  7. Can a stronger model fake being a weaker one? Mostly not

    An evaluation finds frontier models can be prompted down to a weaker model's capability tier but not made to mimic a specific predecessor's error fingerprint, and that targeted sandbagging was largely a null result — so on whether models can deliberately hide capabilities during testing, the answer for now appears to be 'not yet,' though that may not hold for long.

    Rob Kopel via LessWrong
Beta digest — summaries are AI-generated; please verify against the linked sources before relying on them.
← NewerDigest: US export directive recalls Anthropic f…
17 Jun 2026
Older →Digest: Zhipu GLM-5.2 open frontier model, Open…
15 Jun 2026
← All past issues