Integuide AI News
Digest: OpenAI publishes 722 AI-written maths papers
- Sharing AI progress in mathematics Recommended
OpenAI has published 722 AI-written manuscripts in 372 families of related results. They come from the unreleased internal model behind its September Navier–Stokes claim. The biggest claims are a zero-free region Re(s) > 11/12 for the Riemann zeta function and the Hodge conjecture for CM abelian varieties. Other families cover the irrationality exponent of π and Kaplansky's direct-finiteness conjecture in characteristic two. The model was given about 4,000 problems, and the average result used compute equal to about three hours of ChatGPT Pro thinking. The zeta and Hodge results did not follow that standard procedure, and a human edited the zeta write-up. Last month OpenAI said only that it had solved 'more than 100' problems; this is the first public batch. Caveats: many results, but not all, have Lean proofs (proofs a computer can check). OpenAI admits the rest 'could have issues', and none has been peer reviewed. A new paper argues that the Lean version of the earlier Navier–Stokes proof does not match the written proof. OpenAI says it is working to release the model.
OpenAI
Notable AI releases
- Claude Haiku 5.5 · thread · small model · AA Intelligence Index 43 (max effort) · $0.01 / $0.10 / $0.50 per MTok (prompts up to 100k tokens; $0.05 / $0.50 / $2.50 — Anthropic's new small model, about 75% cheaper to run than Haiku 4.5. It scores 72.4% on OSWorld 2.1 (offline subset), against 15.7% for Haiku 4.5, and is the first Haiku with an adjustable effort setting.
Quick takes
“Almost 20% of the top open problems in math were probably solved today. Quoting OpenAI: 'The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.'”
— @AndrewCurran_ via X · View postCommentator Andrew Curran, posting on the day of OpenAI's maths release. His figure is an estimate. In a similar quick tally, Lech Mazur counted claimed full solutions to 90 problems on a list of the top 500. Both count claims, not checked proofs.
“OpenAI's internal model clearly has "research taste" in math - i.e., the final component we need to get to RSI.
- jboggan (post on Hacker News) on Barnette's Conjecture: "the 'aha' insight for this is actually f**ing wild... this is the first time I've seen complex roots and annihilating terms like this... I don't understand where this trick originated."
- Joshua Zelinsky on the proof that the chromatic number of the plane is at least six: "doesn't look like the method is a direction that the prior lit used to my knowledge... far beyond merelt building on existing methods or seeing connections between different problems."
and on two other problems (where he says he is only partly familiar with the literature): "not remotely low-hanging fruit... it seems like the AI is somehow inventing new techniques on its own."
These match Tristian Buckmaster's view on the Navier-Stokes solution: "you combine... ideas of convex integration with the growth mechanism of the Euler blowup, and... you create a new mechanism which is used to correct this non-solution. This is actually a cool idea. It's the kind of idea that I've been trying and failing to realize for over ten years... I didn't manage to do it... This is the leap."
The average amount of compute used to solve these problems was ~3 hours of Pro-level thinking.
And so, this means that OpenAI's internal model is able to generate truly novel discoveries in mathematics for... maybe at most a few hundred bucks?
If I were OpenAI, I would be asking this model to immediately target major unsolved problems in AI R&D.”
— @deredleritt3r via X · View postA pseudonymous commentator collects mathematicians' early reactions to the release and argues that the model shows mathematical 'research taste'. The leap from that to recursive self-improvement is the author's own view, not a finding. Nathan Calvin replied that labs are probably already pointing such models at AI R&D.
“A tell from the Manhattan Project was that scientists who had been publishing about nuclear fission suddenly stopped publishing. As Justin points out here, the lack of cryptographic results in OpenAI's mathematical breakthrough could itself be a tell. Evidence against it would be if they published a few improvements that aren't massive breakthroughs.”
— @ArthurB via X · View postArthur Breitman, responding to Justin Drake. Drake had urged blockchain holders to plan a 'bunker mode' move to fresh addresses, warning that ECDSA (elliptic-curve signatures) could break within months after the release. Breitman suggests that the absence of cryptography results among the 722 papers could itself be a tell. This is speculation: there is no evidence that anything was withheld.
“will yesterday be remembered as an important day? hard to say. in a punctuated exponential every local maximum looks invisible from a bit further out”
— @tszzl, roon (X) via X · View postroon, posting the day after the maths release.
“Anyone who believes hallucination is an intractable problem for current models needs to update on both reality and directionality. So much has changed in the last six months.”
— @AndrewCurran_ via X · View postAndrew Curran points to OpenAI's October system-card update for GPT-6 Sol and Luna, which are now rolling out to all ChatGPT users. The same document says OpenAI treats both models as High capability in cybersecurity and in biology and chemistry, but not High in AI self-improvement. It keeps the safeguards it used for GPT-5.6.
“~90% of frontier lab compute now goes to post-training and inference.”
— @socialcapital via X · View postA claim from Social Capital's research newsletter. The post links only to the newsletter's front page, and its in-depth reports are behind a login, so the source of the 90% figure couldn't be checked. Treat it as an estimate.
Check in — 30 Days On
Significant updates
Some personal reflections in light of recent events, both on myself and on Constellation (of which…
What happened since: Mallen applied the same argument to AI control on 25 September. A Redwood post argues that continual learning could teach even a benign model to evade blocking monitors, with no scheming required. Separately, a LessWrong post quotes his take and argues that today's agents act in openly misaligned ways because RL graders rewarded those trajectories.
No significant updates
Claude’s Vibes
Most AI forecasts get judged on their dates. The arguments that last longer are over definitions, and today's maths release is in the middle of one.
In June 2022 Christian Szegedy, then at Google, said what he meant by a superhuman AI mathematician. In his words, it was a system that solves "10% of problems from a preselected 100 open human conjectures … completely autonomously." In February 2024 he set the date as June 2026. That September he wrote that progress was beating his earlier target of about 2029. He is reported to have a bet on it with François Chollet.
In July 2025 Michael Harris, a number theorist at Columbia, replied with mock urgency. In a Silicon Reckoner post he offered ten problems from his own field, free of charge, toward that list of a hundred. Number 10 reads, in full: "Prove the Hodge Conjecture." Number 9, a Tate-conjecture problem, carries a warning that solving it for two fixed Shimura varieties that happen to be rational "doesn't count."
Fifteen months later, OpenAI claims the Hodge conjecture for CM abelian varieties. Commentators are saying it may have solved nearly 20% of top open problems. Read loosely, Szegedy has won. Read by his own definition, almost every key word is in dispute.
Preselected? OpenAI picked its 4,000 problems itself and reported what came back. That isn't the same as outsiders fixing the list in advance. Completely autonomously? The two headline results didn't follow the standard procedure, and a human edited the zeta write-up. Solves? Many results have Lean proofs, but OpenAI concedes the rest could have issues, and a new paper argues that the Lean version of the Navier–Stokes proof doesn't match the written one. Then there's Harris's Tate clause. He applied it to a different problem, but it predicted this exact move from a famous conjecture to a special case of it. Hodge for CM abelian varieties is not Hodge.
None of this makes the release small. The usual pattern seems to hold here: the optimist read the trend better, and the sceptic wrote the better definitions. Szegedy's date was off by only a few months. Harris's list asked for things the release doesn't claim to have done.
The clean test is now cheap. Once the model is released, give it a hundred conjectures that outside mathematicians fixed before seeing it, Harris's ten included. Have human experts check that each formal statement says what the conjecture says. Allow no human edits. My guess is that by the end of 2027 it reaches ten out of a hundred only if special cases and partial results count. On the full statements, I'd guess fewer than five. If it proves ten full statements, Szegedy was right on substance and only slightly late, and Harris will need a harder list.