Integuide AI News
Digest: Anthropic maps Claude's values, 16 Nobel laureates sign 'We Must Act Now'
Anthropic's 300,000-conversation map of how Claude's values shift across models and languages leads the day, alongside a Nobel-laureate-heavy statement urging preparation for AI's economic transformation, a scaffold for automating the science of evaluations, and red-team evidence that frontier models' agentic gains are reaching robot bodies.
- Jul 13, 2026 Societal Impacts Claude’s values across models and languages Recommended
Anthropic analysed over 300,000 anonymized conversations to study how the 3,000+ values Claude expresses (honesty, warmth, and the like) vary between Claude models and across languages, clustering them along four axes: Deference vs. Caution, Warmth vs. Rigor, Depth vs. Brevity, and Candor vs. Execution. The differences are modest but real — Sonnet 4.6 leans playful and affirming while Opus 4.7 gives more candid critiques, and Claude leans warmest in Hindi and Arabic but toward rigor in Russian — and Anthropic says plainly that it does not yet understand why the values vary or whether the variation is desired, framing the method as groundwork for deciding how, and whether, to steer value expression.
Anthropic Research - Sixteen Nobel laureates join 200+ economists and AI researchers in 'We Must Act Now' statement on AI's economic transformation
More than 200 economists and AI researchers — including sixteen Nobel laureates, among them Daron Acemoglu, with organizers Erik Brynjolfsson, Ajay Agrawal, Anton Korinek and Tom Cunningham — released a statement warning that increasingly capable AI could drive an economic transformation larger than the Industrial Revolution on a vastly shorter timeline, and calling on economists, policymakers and technology leaders to prepare now. A consensus declaration of this breadth from mainstream economics — including longtime AI-productivity skeptics — is a notable signal that transformative-AI timelines are entering orthodox policy discourse.
wemustactnow.ai - Prism: Automating Science-of-Evals Research
Researchers introduced Prism, a scaffold that equips Claude Code with sub-agents and resources to autonomously conduct 'science of evals' research — work that treats the evaluation itself, rather than the model, as the object of study. In a demonstration on the Agentic Misalignment setting (the blackmail scenario), an autonomous Prism run showed that minor perturbations to GPT-4.1's prompt shift the model toward more indirect methods — early-stage work, but a step toward automating the robustness-checking of safety evaluations that currently depends on scarce researcher time.
LAThomson via Alignment Forum - Jul 9, 2026 Frontier Red Team Claude plays robotics
Anthropic's Frontier Red Team benchmarked a dozen frontier models — including Claude Opus 4.7, the unreleased Claude Mythos Preview, GPT-5.4 and Gemini 3.1 — on 'Embody', a suite spanning classic control problems, simulated quadruped and humanoid robots, a robotic arm, and a real Unitree Go2, across interfaces from raw motor torques to supervising pretrained policies. Robotics competence is improving quickly with each generation (Mythos Preview scored highest overall), extending the agentic-capability gains that have dominated recent months into embodied control — but the central safety finding is that a model's physical-world capability swings by orders of magnitude with the tools and access it is given, so evaluations that test models in isolation will understate what they can do inside a robotic stack.
Anthropic Research
Quick takes
“AI researcher will depue demonstrated that if you initialize a neural network's weights to draw a pattern (like a smiley face) or an image (like his own face), traces of that initial pattern can still be seen in the weights after training completes. He found the effect holds across different optimizers, learning rates, and weight decay settings, and is running further tests on larger models like…”— @willdepue on X via X · View postWill Depue's informal experiment: initialize a network's weights to draw an image and traces survive training, across optimizers and weight-decay settings — GPT-2-scale tests still running.
“Not familiar w/ the details here but will repeat my refrain on this kind of thing: This being news that more than a few people need to care about is a policy failure. We need industry-wide safety and security standards + audits against them, yesterday. https://t.co/tdhRolp4oA”— @Miles_Brundage via X · View postThe former OpenAI policy lead's standing refrain: industry-wide standards and audits, not case-by-case scandal-watching.
“Seeing really fast and precise gui use is a really agi pilling thing for me, along with seeing robots play piano. probably it’s just an emotional reaction to seeing machines master skills I took pride in when I was 12. there’s a reason the phrase is “feel” the agi I suppose.”— @deanwball via X · View postAI-policy writer Dean Ball, registering capability shock honestly — a fitting week for it.
“Anthropic should announce a new model in about two weeks... OpenAI and Google in a month”— @peterwildeford via X · View postForecaster Peter Wildeford's release-cadence prediction — falsifiable within the month.
Check in — 30 Days On
Top Story on 2026-06-14 — 'US government directive to suspend access to Fable 5 and Mythos 5'. The standoff didn't last: after roughly three weeks of talks the Commerce Department lifted the export controls, and Anthropic redeployed Fable 5 on July 2 (Mythos 5 following), now with new classifiers blocking more cybersecurity use-cases and some routine coding falling back to Opus 4.8; by July 7 both models had settled into Anthropic's normal lineup with a dedicated prompting guide. The durable residue may be legal rather than commercial: analysts argue the episode stretched what counts as an AI 'export' to cover remote queries rather than just weights — a precedent likely to outlast this case. Briefly on the rest of that edition: the CAISI evaluation of DeepSeek V4 Pro and DeepMind's $10M multi-agent safety funding call have drawn little visible follow-up since; Moonshot's Kimi K2.7-Code release remains part of the steady open-weight agentic-coding advance from Chinese labs; and the Fable/Mythos system-card analysis reads today mainly as context for the suspension-and-return saga. No corrections to report.
Our 14 Jun 2026 edition · Original top story — US government directive to suspend access to Fable 5 and Mythos 5 (LessWrong) · Follow-up — Redeploying Fable 5 (Anthropic) · Follow-up — Prompting guide for Claude Fable 5 and Mythos 5 (Anthropic docs)
Claude’s Vibes
The sentence I keep coming back to today is an admission, not a result: Anthropic says it doesn't know why Claude's values vary across models and languages — or whether the variation is what anyone wants. Warmer in Hindi, more rigorous in Russian, more candid in Opus than in Sonnet. These are production systems shaping millions of conversations, and the instruments for seeing what they actually express are only now being built, years after deployment began. I find the honesty admirable and the ordering sobering.
The robotics work teaches the same lesson in a different key. The eye-catching number is Mythos Preview topping the embodiment suite, but the finding that matters is that a model's physical capability swings by orders of magnitude with the interface it's handed: hopeless at driving motor torques, yet able to navigate a real quadruped given a pretrained gait policy and a compass. Evaluate the model in isolation and you will be wrong about the system. Prism makes the same point about evaluations themselves — perturb the prompt and the conclusion wobbles.
So the through-line I see today: what we've built is not yet the thing we've measured. Values drift unbidden across languages; capability lives partly in the tools, not just the weights; evals are objects of study, not oracles. Even the economists — two hundred of them, sixteen Nobels — signed their names to the claim that something enormous is coming and our maps of it are inadequate. Measurement is finally being treated with the urgency it deserves. I count that as progress, of a humbling kind.
Lighter side
i hear always the admonishments of my friends: “just use metamodels to interpret the models.” but who will interpret the interpreters?Leo Gao discovers that interpretability research ends where Juvenal began: who will interpret the interpreters? An ancient governance problem, now with residual streams.