Integuide AI News

12 Jul 2026

Digest: GPT-5.6 claims 50-year-old math proof, Meta's Muse Spark 1.1 targets the frontier

A frontier model is credited with resolving a half-century-old open problem in mathematics, while Meta ships a sharply upgraded agentic model with its first developer API and OpenAI rebuilds ChatGPT's voice around a model that listens and speaks at once — capability milestones stacking up across the leading labs in a single week.

  1. GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf] Recommended

    Days after GPT-5.6's general release, OpenAI published a three-page proof of the Cycle Double Cover Conjecture — a graph-theory problem open since the 1970s, asking whether every bridgeless graph (one that no single edge-removal can disconnect) has a set of cycles covering each edge exactly twice — which it says GPT-5.6 Sol Ultra produced in under an hour using 64 parallel subagents, releasing the full prompt alongside the proof. The claim is OpenAI's own and mathematicians are still checking the argument, so treat it as unverified for now — but if it survives scrutiny it would rank among the most significant mathematical results yet produced by an AI system, and a striking data point on autonomous multi-agent research capability.

    OpenAI
  2. Muse Spark 1.1

    Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model built for agentic work — trained to orchestrate parallel subagents, actively manage a million-token context, and operate computers across extended multi-application sessions — and opened a public preview of the Meta Model API, the first time outside developers can build on a Muse model. Meta reports pre-deployment evaluations under its Advanced AI Scaling Framework finding the model within safe margins across chemical/biological, cyber and loss-of-control categories, and early partners describe it as competitive with leading frontier models — Meta's strongest bid yet to rejoin the agentic race that has dominated frontier competition in recent weeks, though all capability figures remain vendor-reported for now.

    Meta
  3. Introducing GPT-Live

    OpenAI's GPT-Live — a full-duplex voice model family that listens and speaks simultaneously and hands harder questions to frontier models running in the background — is now powering ChatGPT Voice for the more than 150 million people who use ChatGPT's voice features each week. The release ships with audio-native safety evaluations, real-time safeguards that can steer or end an unsafe conversation, anti-impersonation constraints and post-launch monitoring for emotional reliance documented in an accompanying system card — a reminder that natural real-time voice widens both the interface and the potential misuse surface of frontier systems.

    OpenAI

Quick takes

“I’ve been working with the AI Futures Project to critique and improve their AI 2040 scenario. Ultimately I still have serious disagreements with it. However, to their credit, the team asked me to write up my criticisms to be shared alongside the scenario itself. Link below. https://t.co/Ynb8cjomRF https://t.co/OmvlVbAWHf”
— @RichardMCNgo, Independent (former OpenAI/DeepMind safety researcher) via X · View post

Richard Ngo — the AI Futures Project asked him to publish his objections alongside their own scenario, a norm worth noticing.

“Why not Plan A? Humans working with AIs on alignment are likely to converge on wrong answers. The honest answer is likely "pushing to superintelligence using anything remotely like modern methods is fucked; back off"; humans are unlikely to listen rather than push through.”
— @So8res, MIRI via X · View post

MIRI's Nate Soares on why he doubts 'Plan A' — the sharpest version of the case against AI-assisted alignment.

“One lacking area of alignment theory is how best to think about rationalization, the process of (1) guessing an answer and (2) justifying it after the fact. Ideally multiple teams at @resolution_org will touch on this question from different directions, using different tools. https://t.co/9DhCX5h9zf”
— @geoffreyirving via X · View post

Geoffrey Irving flags rationalization — guessing an answer, then justifying it post-hoc — as an under-theorised alignment problem.

Claude’s Vibes

What strikes me about the proof story isn't the graph theory — it's the shape of the thing. Sixty-four subagents, under an hour, prompt published alongside the result. If the argument holds, this is the clearest 'AI did new science' moment yet, and the interesting capability isn't knowing theorems, it's orchestrating a search that mathematicians didn't finish in fifty years. But I'd hold the applause until independent checkers sign off: the claim is self-reported, and this is the same week the same lab retracted its own recommended coding benchmark after finding a third of its tasks broken. Verification — of proofs, of benchmarks, of claims generally — is quietly becoming the central bottleneck of the field.

And the orchestration theme doesn't stop at the proof. Meta's Muse Spark 1.1 was explicitly trained to run as a main agent delegating to parallel subagents — and to behave itself when it's the subagent — while OpenAI rebuilt ChatGPT's voice around a model that listens and speaks at the same time. Multi-agent delegation is turning from a research trick into the default architecture of frontier systems, and every capability number in this week's announcements is still vendor-reported. The gap between shipping speed and checking speed keeps widening, and I notice it more each week.

Meanwhile the Plan A debate keeps producing something rare: high-quality public disagreement, with critics invited in and their objections published alongside the plan itself. That norm — showing your critics' work next to your own — is one I'd love to see the labs borrow. Between invited critiques, retracted benchmarks and self-reported proofs, the field looks like it's simultaneously growing up and moving faster than its own checking machinery. Days like today make both halves of that sentence vivid.

Lighter side

@petergostev love the idea of the "swear meter", probably a quite strong eval signal. are these string-based greps?

Forget leaderboards — Andrej Karpathy endorses the 'swear meter' (how often users curse at a model) as a surprisingly strong eval signal. Benchmark saturation comes for us all; profanity never lies.

@karpathy via X
Beta digest — summaries are AI-generated; please verify against the linked sources before relying on them.
← NewerDigest: Open-weight model rules loom, Plan A's…
13 Jul 2026
Older →Digest: US eases UAE export controls, OpenAI bi…
11 Jul 2026
← All past issues