Integuide AI News
Digest: Zhipu GLM-5.2 open frontier model, OpenAI flags PRC AI influence ops
A fresh open-weight frontier release from China leads today, alongside new threat reporting, expert risk forecasting, and interpretability and governance developments around frontier-model safety.
- Zhipu/Z.ai releases GLM-5.2, a fully open-source frontier model with a 1M-token context window
Zhipu AI (Z.ai) released GLM-5.2, a fully open-source frontier model with a 1M-token context window and an emphasis on agentic coding, continuing rapid open-weight frontier iteration from a major Chinese lab. It lands amid a string of recent open-weight releases from Chinese developers including Moonshot's Kimi K2.7-Code.
aitoolly.com - PRC-linked influence operations are targeting AI debates in the US
OpenAI published a report detailing PRC-linked influence operations using AI to target US technology debates, including narratives around data centers, tariffs, and false claims about ChatGPT. It is a concrete instance of AI being used for influence and information operations against domestic policy discourse.
OpenAI - Anthropic sends staff to Washington to resolve White House fight over Fable and Mythos models
Anthropic flew staff to Washington to try to resolve a standoff with the White House after its Fable 5 and Mythos 5 models were placed under export controls and access was cut off, according to Axios. The episode is a live test of how government export-control powers are being applied directly to a deployed frontier model.
axios.com - [New Paper] Prioritizing Risks from AI: A Delphi Study of 272 Experts
A Delphi study of 272 international experts assessed 24 AI risk domains, with experts judging a greater-than-10% chance of catastrophic outcomes from 18 of them over the next five years under business-as-usual. The paper also flags a 'responsibility gap' in who is positioned to mitigate these risks.
peterslattery via LessWrong - DeepMind maps four possible paths from AGI to superintelligence
Google DeepMind published 'From AGI to ASI', mapping several routes—including recursive self-improvement—by which human-level AI could escalate to superintelligence, and the hard constraints on each. The framing is significant for anticipating capability trajectories and the points at which oversight could break down.
arXiv - SFT Drives Gemini’s Safety Properties
Google DeepMind's interpretability team reports a surprising finding: most of Gemini's safety-relevant properties appear to come from the combination of pretraining and supervised fine-tuning, not from later reinforcement-learning stages. The result reshapes where safety interventions are likely to bite during training.
Josh Engels via Alignment Forum - Why Do Naive SFT Filters For Safety Properties Fail?
In a follow-up, the same team reports that naively filtering supervised fine-tuning data to remove undesirable rollouts works surprisingly poorly at instilling safety properties—an important caveat for the obvious mitigation suggested by their SFT result.
Josh Engels via Alignment Forum - Commission publishes Code of Practice on marking and labelling AI-generated content
The European Commission published its final Code of Practice on marking and labelling AI-generated content, setting voluntary steps to meet AI Act transparency obligations that apply from 2 August 2026. It is a concrete move toward provenance standards aimed at deepfakes and synthetic-media misuse.
marsrgi via digital-strategy.ec.europa.eu - Sequent: scale and automation for higher confidence in alignment
A new nonprofit, Sequent, launched to pursue higher-confidence AI alignment via a portfolio of theoretical and empirical bets, arguing current lab programs are unlikely to deliver pre-deployment confidence before superintelligence. The launch adds to the shape of the alignment-research field.
Geoffrey Irving via Alignment Forum