Integuide AI News
Digest: DeepMind AI control roadmap, lab pre-deployment safety methods
Google DeepMind sets out a plan for containing misaligned AI agents, as labs publish new methods for predicting model behaviour before release and probing what frontier models know about their own inputs — alongside fresh dangerous-capability evaluations and governance moves across the field.
- GDM AI Control Roadmap
Google DeepMind published version 0.1 of an AI Control Roadmap, a plan for internal guardrails designed to catch adversarial behaviour by AI agents that become harder to oversee, drawing on cybersecurity-style threat modelling and system-level mitigations to limit the harm a misaligned system could cause. It is one of the more concrete frontier-lab commitments to date on the 'control' problem of keeping potentially misaligned agents contained.
Mary Phuong via Alignment Forum - Predicting LLM Safety Before Release by Simulating Deployment
Researchers describe a pre-deployment safety method that estimates not just what a new model can do but how it is likely to behave in real-world use, simulating deployment to surface where it might introduce new risks before release. As capabilities rise, scalable behavioural forecasting like this is increasingly central to deciding whether a frontier model is safe to ship.
Tomek Korbak via Alignment Forum - Several frontier models are substantially prefill aware
A new paper finds that several frontier models are substantially 'prefill aware' — able to detect when text has been inserted into their own responses — and extends the result to lower-stakes settings. Such situational awareness matters for evaluation validity, since models that can tell when their outputs are being manipulated or tested may behave differently under scrutiny.
yeedrag via LessWrong - Synthetic document finetuning for instilling positive traits
Google DeepMind's interpretability team reports training Gemini 3 Flash to hold specified traits and values by mid-training it on synthetic documents describing those properties, then fine-tuning — an approach to instilling desired behaviour rather than only filtering bad outputs. It adds to a growing body of work on shaping model character at training time and on understanding when such methods hold or break down.
CallumMcDougall via Alignment Forum - Jun 18, 2026 Frontier Red Team Project Fetch: Phase two
Anthropic's Frontier Red Team released phase two of Project Fetch, its work probing whether frontier models can uplift biological misuse. Independent dangerous-capability evaluation in the bio domain is central to assessing whether new models cross thresholds of concern.
Anthropic Research - Inaugural Global Dialogue on AI Governance to convene in Geneva on 6–7 July with Scientific Panel's first report
The UN will hold its inaugural Global Dialogue on AI Governance in Geneva on 6–7 July, where the new 40-member Independent International Scientific Panel on AI, co-chaired by Yoshua Bengio and Maria Ressa, will present its first preliminary report. It establishes an annual scientific-input channel into multilateral AI governance.
United Nations via un.org - South Korea becomes 4th country to forge AI security alliance with OpenAI
South Korea became the fourth nation to join OpenAI's AI security partnership program, giving its national institutes a channel for evaluation and security collaboration with a leading frontier lab. It reflects deepening government-to-lab cooperation on frontier-AI safety testing.
The Korea Times