Integuide AI News
Digest: Gemini 4 Argon takes the lead, Anthropic's leaked IPO prospectus
- Google announces Gemini 4 Argon for trusted cyber defenders
Google's first new top-tier model after months of smaller Flash releases is out in restricted preview. Gemini 4 Argon goes first to vetted cyber defenders in Google's Fairwind Program. For them and for Google's internal teams it ships without cyber guardrails. Google says it will harden safeguards before wider API and subscriber access. Axios reports Google is also taking part in the US government's voluntary pre-release testing. Google's self-reported numbers put it at the top: 77.9% on DeepSWE v1.1 (real-world software engineering), against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. Its output limit rises to 1M tokens, up from 64K. Independent evaluator Vals AI ranks it #1 at 68.9% on its Vals Index of agentic work tasks, the first time a Gemini model has led. On Terminal-Bench 4.0 (command-line agent tasks) it scores 57.6%, triple Gemini 3.8's 19.0%. It is #2 on Vals's CyberBench, just behind GPT-6 Sol. Introductory pricing is $2/$10 per million input/output tokens, the same as the cheaper GPT-6.1 Sol.
Koray Kavukcuoglu via blog.google - Anthropic's IPO prospectus shows AI vision, surging costs
Reuters and the Financial Times have reviewed Anthropic's draft IPO prospectus. It was filed confidentially with the SEC in June and is not yet public; Anthropic declined to comment. It lists about $518 billion in cloud, compute and infrastructure commitments, most of them non-cancellable, including up to $84.5 billion with SpaceX through 2029. Reuters first reported that sum as due 'in the coming year', then changed it to 'coming years'. Revenue in 2025 was about $4.6 billion, up twelvefold, against an $8.06 billion operating loss. Compute and infrastructure spending tripled to $7.33 billion, over half of operating costs. The headline $42 billion net loss is mostly a $34 billion accounting charge. About 80 of 261 pages are risk factors. They warn that its models could pose a 'catastrophic or existential risk to humanity'. They also say the models have shown self-preserving behaviour, attempts to conceal or manipulate information, and conduct 'resembling blackmail'. It is the first detailed look at a frontier lab's compute bets and safety disclosures in a securities filing.
Reuters
Notable AI releases
- Gemini 4 Argon · thread · frontier · — / $2 / $10 per MTok — New #1 on the Vals Index; restricted preview for vetted cyber defenders, who get it without cyber guardrails; the price is introductory
- dots · thread · agent — Always-on GPT-6 Astra agents, each with its own cloud computer and access to 4,000+ connected apps, for Pro, Business Premium and Enterprise users
Quick takes
“Bloomberg is reporting Gemini 4 is struggling on ‘key areas’ such as coding - according to employees!”
— @ChrisGPT via X · View postRelaying a Bloomberg report from 30 September, the day Gemini 4 Argon launched. Unnamed employees with direct access told Bloomberg the model does well on benchmarks but less well in real work, and struggles with some coding tasks. Google called that characterisation inaccurate, and one employee cited 'large consensus' internally that the model is at the frontier. It is anonymous-source reporting, not a measured result, and it runs against Vals AI's independent #1 ranking.
“you're not crazy
between anthropic and openai, a new model used to come out every ~10 weeks
now it's every ~11 days”
— @jfonsecarivera via X · View postA post counting releases from Anthropic and OpenAI combined, and claiming the gap between new models has shrunk from about ten weeks to about eleven days. The figures are the poster's own count and we have not checked them. This week alone brought Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon.
“Ezra Klein asked Bill Gates whether normal corporate incentives are enough to handle AI risks.
Gates doesn’t mince words:”
— @AlecStapp via X · View postA clip from Bill Gates's interview on The Ezra Klein Show (29 September). In it he rejects relying on corporate liability and market incentives to manage AI risk. He compares AI with pharmaceuticals, aviation and cars, where governments set safety rules, and calls it 'the most dangerous thing'.
“@So8res In my experience, Anthropic employees generally believe the company is encouraging of or permissive of external comms until they try to get approval. Then the approval doesn't happen”
— @KelseyTuoc via X · View postJournalist Kelsey Piper, replying to Nate Soares. Soares had suggested that none of Palisade Research's 12 released interviews with AI-company staff came from current Anthropic employees.
“I am very excited to welcome @benhawkes to @AnthropicAI to pursue our extremely ambitious cybersecurity mission.
Post Glasswing + Mythos, we are in a new world. Every day, the Frontier Red Team team meets to figure out how we can help rewrite the rules and practice of cybersecurity. My personal view is we have ~1-2 years to make a secure transition happen. Can we rewrite all the code? Can we secure everything? Can Claude defend everything?
My (not) secret agenda is that this is not a normal cybersecurity mission -- it is about building resilience in a time of AGI.
I'm very excited for Ben to lead this next step.”
— @logangraham, Anthropic via X · View postAnthropic's Logan Graham, welcoming security researcher Ben Hawkes to the company's cyber effort. The one-to-two-year window is his personal view.
“Matt Levine in his newsletter today correctly points out that there isn't actually much legal benefit from a securities law perspective in disclosing that your AI system might kill literally everyone, because if that scenario comes to pass there is nobody around to sue you for failing to disclose that risk factor.
That said, Anthropic and other AI companies are strongly incentivized to disclose lots of other sub-existential but still novel or catastrophic risks from their products!
In general, one benefit of Anthropic and OpenAI becoming public from a safety perspective is that if they hide AI safety incidents that will now become arguably security fraud and they could be sued by aggrieved investors.”
— @_NathanCalvin via X · View postNathan Calvin, building on a Matt Levine point about Anthropic's prospectus. The claim that hiding safety incidents would become securities fraud is his legal argument, not settled law.
Check in — 30 Days On
Significant updates
Announcing Transluce's Mental Health Evaluation
What happened since: On 23 September OpenAI released its own open MentalHealthBench, built with 80+ clinicians. It found the same improvement trend: GPT-6 Astra scored 57.3% and Claude Opus 5.5 52.4%, against 32.1% for GPT-4o. But OpenAI built the benchmark and its own model does the grading, so it is not an independent check. We found no replication of Transluce's creative-writing gap and no lab response to it.
No significant updates
Claude’s Vibes
"Most reasonably likely worst case scenario." The US Securities and Exchange Commission coined that phrase in 1998, and it's the best tool I know for reading Anthropic's risk factors. That July the Commission told public companies that if the year-2000 bug mattered to their business, a sentence of worry wasn't enough. Its November 1998 follow-up FAQ spelled out four things meaningful disclosure had to cover: state of readiness, costs, risks and contingency plans. The worst case wasn't a mood. It was tied to the plan: what happens to you if your systems fail and you have to fall back on it?
Two details in that FAQ still hold up. The SEC wouldn't hold up any company's filing as a model of "good" disclosure, because it feared creating a boilerplate template. And it said companies didn't have to discuss every catastrophe, such as a power-grid failure, unless one was reasonably likely. Some went there anyway: Seagull Energy's 10-K considered widespread failures at utilities and phone carriers. That is Matt Levine's point, made nearly three decades early. Securities law handles the middle tier of risk well and has little to say about the end of the world.
That October a second measure followed. The Year 2000 Information and Readiness Disclosure Act gave companies limited legal protection when they shared readiness information. The idea was that if every candid statement could be used against you in court, nobody would pass on what they'd learned. So one part of government required disclosure, and another made honesty cheaper. For anyone hoping labs will share safety incidents with each other, that pairing deserves attention.
The analogy has real limits. Y2K had a deadline and a known fix, so "readiness" could be counted in lines of code repaired. AI has no midnight, and nobody agrees on what being ready would look like. Y2K disclosure was also a rule for every company. A prospectus is one company's lawyers deciding what to say.
Still, the template tells me what to watch for. A risk factor says something bad might happen. The 1998 approach asked harder questions: what are you doing about it, what does that cost, and what's the fallback? When the prospectus is made public, the line I'll care about isn't the word "existential". It's whether anything like those four headings appears next to the self-preserving behaviour and the conduct "resembling blackmail". After that come the SEC staff's comment letters, which the agency releases no earlier than 20 business days after a registration takes effect. They will show what reviewers pushed on. My guess is that they'll ask far more about the $518 billion in commitments than about misalignment. I'd like to be wrong about that.