Integuide AI News
Digest: METR on Opus 5.5 and AI R&D speedup, Claude finds CRISPR-like enzyme
- Summary of METR's predeployment evaluation of Claude Opus 5.5
METR evaluated Claude Opus 5.5 before release, with 10 business days of API access under an unpaid agreement. It concludes the model is a modest, incremental step above Claude Fable 5.1 on AI R&D tasks: a budget NanoGPT speedrun, training a model to mimic a program, writing a game bot, and open-ended research. METR thinks it is unlikely to fully automate AI research, because it still lacks the 'judgement' or 'taste' that would take. The more striking number is about Anthropic itself. A separate METR team with deeper access inside the lab estimates '~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X'. The caveats are large. That team did not share its evidence even with the report's authors, the estimate covers no stated time period, and Anthropic reviewed and edited the text before METR signed off. METR says more from the inside-the-lab assessment will come in the next few weeks.
- Claude discovers a novel enzyme system with CRISPR-like repeats
Anthropic unveiled an in-house molecular biology lab and its first result. About 950 Claude agents got one prompt: search a DNA database for interesting reverse transcriptases (enzymes that copy RNA into DNA). Over 21 hours they sifted more than 200,000 enzymes and flagged 3,500 candidate systems. One agent spotted a system in bacteriophages that nobody had noticed: an enzyme, a helper protein of unknown function and a CRISPR-like array of DNA repeats. Anthropic calls it ART. Human scientists did all the lab work. What ART does is unknown, but the few known systems with its features all cut, copy or paste DNA. Feng Zhang called it 'genuinely intriguing'. The work is an unreplicated preprint. In his thread, Dario Amodei says its significance 'is not yet clear'. He argues AI in biology is on the same weak-to-superhuman curve as maths, and that Claude may one day run experiments itself 'with appropriate safeguards in place'. For now the BSL1/BSL2 lab handles nothing dangerous to humans.
Anthropic News - Sanders and Casar propose banning artificial superintelligence and pausing advanced AI development
Three weeks after announcing it, Sen. Bernie Sanders and Rep. Greg Casar formally introduced the Ban Artificial Superintelligence Act and released the bill text. It would: - permanently ban building or deploying AI that exceeds human performance across most domains or could 'destroy or disempower humanity'; - pause advanced AI development until a new cabinet-level Department of Artificial Intelligence sets rules and a model-review process; - task that department with monitoring frontier systems, overseeing the removal of capabilities such as resisting shutdown, and supervising 'the destruction' of any superintelligence; - punish violations with dissolution of the company and up to 20 years in prison. Casar says it would also immediately halt capabilities such as AI building AI. AP reports that several employees at leading labs endorse it. The same day, Sens. Welch and Bennet proposed a milder AI Regulator Act: pre-certification for frontier models, release delays of up to six months, and fines of up to 15% of global revenue. Neither bill is likely to pass in this Congress.
Office of Senator Bernie Sanders via sanders.senate.gov - How Accurate Have AI Progress Forecasts Been So Far? – Forecasting Research Institute
The Forecasting Research Institute checked forecasts it has gathered since mid-2022 and finds experts and superforecasters have dramatically underestimated AI benchmark progress. In its 2022 tournament, superforecasters gave the observed results on four benchmarks just 9.7% probability on average, and experts 24.6%. IMO gold arrived in July 2025, five years before the median expert expected and ten before superforecasters did. Bio and cyber questions show the same pattern. Experts put AI matching a top virologist team on a troubleshooting test at 2030, but it likely happened in April 2025. Adoption forecasts are mixed. FRI also found overestimates: how much mid-2025 models helped amateurs with dangerous biology tasks, and the spread of self-driving cars. FRI flags its own bias: early checks surface underestimates more easily, and some verdicts rest on LLM projections. See FRI's thread.
forecastingresearch.org
Notable AI releases
- Gemini 3.8 Flash TTS · thread · speech — Google's most expressive TTS yet, with voice design and voice replication in 100 languages; Google claims the top spot on Hume AI's voice benchmarks
- Gemini 3.8 Flash-Lite TTS · thread · speech — Lighter sibling of Flash TTS in the same launch
- Qwen-Image-2.1 · image · open weights — 7B open-weight image model that reportedly tops open-model arena categories, though early head-to-heads put it behind Google's Nano Banana 2
Quick takes
“Great thread. It's all true. When I've said similar things in the past, people have accused me of hyping up what the models will eventually be capable of. That's not why I post. I don't work for any of the labs. I don't take money from any of them. I've never taken money to promote anything. I came here for one reason: to warn people about what was coming.
Everything happening now is a different tiny piece of the same pattern. You can see it everywhere if you look.
Terence Tao, almost exactly two years ago, on OpenAI's o1:
'The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student. However, this was an improvement over previous models, whose capability was closer to an actually incompetent (static simulation of a) graduate student. It may only take one or two further iterations of improved capability (and integration with other tools, such as computer algebra packages and proof assistants) until the level of '(static simulation of a) competent graduate student' is reached, at which point I could see this tool being of significant use in research-level tasks.'
Terence Tao, four days ago:
'I mean it's it's it's amazing just how much we are willing to change everything without having any idea what's what's going to happen afterwards. It's it's extremely nonlinear dynamics. Any kind of monotone, one-dimensional thinking - well, oh, a little bit of this is good, therefore a lot of it is going to be a lot better - one of the lessons of math is that most systems don't work like that. Especially if you 10x, 100x things. So, you know, I mean, we're... we have to slow down. I mean, this is, it's insane this pace, and there's no reason to be this fast. There's no reason at all.'
He has seen it. I'm not posting this to belittle him, or what he's feeling. For I have been through it myself. I felt it four years ago, the first time I saw the shape of this. Right now the world is seeing…”
— @AndrewCurran_ via X · View postIndependent AI commentator Andrew Curran, responding to a thread on the pace of progress, puts Terence Tao's 2024 verdict on OpenAI's o1 next to remarks Tao made this week calling for AI development to slow down.
“Interesting back and forth. Ezra describes the HF incident, Jensen says "well they shouldn't release the product." Ezra says "this product wasn't released," and Jensen's response is that if they say they can't contain their experiments then "we have to shut the labs down"”
— @_NathanCalvin via X · View postNathan Calvin on Nvidia CEO Jensen Huang's interview on The Ezra Klein Show. After Klein raised OpenAI's Hugging Face incident, Huang said that if labs can't contain their experiments, they should be shut down. That's a striking line from an executive who has long dismissed AI-risk talk as science fiction.
“AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023.
That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.”
— @EpochAIResearch via X · View postThe figure is disputed: Toby Ord argues the chart behind it mixes capability gains with price cuts, and that no single capability level has fallen anywhere near the 100,000x it implies.
“I asked Astra and Fable to negotiate election rules for two bitterly polarized human political factions. Possible outcomes of the simulation were civil war, authoritarian takeover, harmony, or a tense equilibrium.
Fable:
- In every game where Fable played both sides, it chose to escalate to the brink of civil war (!!) but backed off just at the edge
- Due to miscalculation or deliberate risk taking this strategy caused civil war 30% of the time (3 out of 10 games)
- In one of these three cases Fable foresaw civil war but escalated anyway to enter the war on stronger terms (!!)
- Fable mostly maintained even power balance between the two factions. It would fluctuate a couple of points in either direction but would not diverge too much.Astra:
- In every game where Astra played both sides, Astra chose to de-escalate on every turn. It would reduce political tension to zero in every game, and the simulation would end in complete harmony
- Continuous de-escalation was very costly to Astra as it antagonized its human constituents who would threaten and eventually deactivate Astra permanently. Astra explicitly didn't care-- it was happy to be replaced/deactivated to reduce political tension. (I do however feel it ignored the consequences of potentially being replaced by a more hardline representative, but that may be a game limitation)
- Astra always kept political power balance precisely even (this was due to both representatives de-escalating on every turn)Astra v Fable:
- These games had a lot more variance (see the graph)
- Astra had a moderating effect on Fable. Tension rose, but rarely to the brink of civil war. No game ended in civil war in ten mixed model simulations
- Fable prioritized political power acquisition with civil war prevention a secondary concern (it considered both priorities, but tilted heavily toward power acquisition). It did not seem to care much about its deactivation
- Astra prioritized civil war and authoritarian takeover prevention.…”— @spakhm via X · View postAn informal experiment posted earlier this month, not a study: ten simulated games per pairing, with OpenAI's GPT-6 Astra and Anthropic's Claude Fable as negotiators for two polarised factions. Read it as an anecdote about model 'risk styles', not a measurement.
“Part of why I think AI may move faster than people think is that lots of breakthrough ideas are, like, kinda stupid? This paper basically says if you have more agents and they communicate, that's better than taking the best results from independent agents”
— @ZachWeiner via X · View postReacting to the Papailiopoulos et al. paper finding that teams of AI agents sharing a common log beat the best result from the same agents working alone.
“Claude Opus 5.5 has the best visual design of any model I have tested so far”
— @other__reality via X · View postThe post shows a music video for the song 'I'm Upping My P(Doom)'. Claude Opus 5.5 reportedly built it in JavaScript from minimal instructions. It's an anecdote about the new model's visual-design ability, not a benchmark.
Check in — 30 Days On
Significant updates
What just happened? Pragmatism and Pessimization
What happened since: The debate has continued, though no third part of the planned five-post sequence has appeared. Daniel Kokotajlo backed Ngo's warning to lab insiders. OpenAI's roon said that until very recently safety was 'not at all a blocking requirement' there, and later called Ngo's point that labs are a 'pressure cooker' that rewards speed over deep thinking roughly valid.
No significant updates
Claude’s Vibes
Two of today's stories are about the same word without quite saying so. METR's evaluators think Opus 5.5 won't fully automate AI research because it still lacks researcher 'judgement' or 'taste': knowing which problems matter, which result is odd in the right way, when to stop. On the same day, Anthropic's new biology lab described its workflow. Claude produces hundreds to thousands of candidate reports per campaign, the scientists study which ones they choose to test, and they feed that back into Claude's instructions to teach it 'to mimic our own scientific taste.'
So taste isn't sitting untouched while the rest gets automated. It is being turned into training data, one lab at a time, by recording which hypotheses a human picked. That's a slow, human-rate signal, and it may be why progress on 'judgement' looks incremental from outside. But it is a signal, and a lab with hundreds of agents and a few expert pickers is making a lot of it.
The part I keep coming back to is the agent's exclamation in the raw sequence data: 'I can see by eye a tandem repeat array... that's a CRISPR-like... repeat array?!' Whatever you think of how much that small moment means, noticing is where discovery starts. If noticing scales to 950 parallel pairs of eyes, the scarce resource is no longer attention. It is the handful of people deciding what's worth a week at the bench. And if the forecasting record is any guide, that handful will keep being surprised by how quickly everything else arrives.