Integuide AI News
Digest: AI research-taste benchmark, weight-grafting alignment training
- TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
TasteVal, a new benchmark from Ollie J and Dane Sherburn, tries to measure "experimental research taste": how well a model designs AI R&D experiments and interprets their results. Across eight tasks, from pretraining data curation to preference modelling, the model under test only plans and interprets. A fixed coding agent runs each experiment on one H100 within a 40 GPU-hour budget. Taste is scored as compute efficiency against the best of 24 human experts. The authors report that frontier models' taste has doubled every 3.0 months since December 2025 (95% CI 1.7–5.0). Opus 5.5 matches the expert score with 2.3x less experiment compute (CI 1.15–4.37x), at about 1/30 the cost. That cuts against METR's pre-release finding last month that Opus 5.5 still lacked the taste to automate AI research. Caveats the authors flag: the tasks are quick to verify and easy for labs to hill-climb, the baseline excludes top researchers, and choosing which problems to work on isn't measured. This is a new, unreplicated benchmark (paper).
Ollie J via LessWrong - You can graft SDF changes from base models onto post-trained models
MATS fellows Dani Roytburg and Peter Nutter, with Clément Dumas and mentor Shi Feng, propose "grafting": run synthetic-document fine-tuning (training on fabricated, pretraining-style text to implant beliefs or values) on the base pretrained checkpoint, then add the weight change to the finished post-trained model. Across five model families up to 284B parameters, grafts did less damage to coherence and grip on reality than tuning the chat model directly. One graft worked across SFT, DPO and RL checkpoints, but only if trained from base. Replicating a constitutional mid-training pipeline, grafted alignment training held up better through later SFT and RL without capability loss; emergent misalignment transferred too. Dumas calls it another update against the persona selection model; Feng's thread summarises. Caveats: installation is weaker, the post-training comparison used ~200k examples on a small model, and robustness to misaligning RL is untested.
LessWrong - Identifying Introspection From the Inside
David Atkinson and David Bau (Northeastern) and Dillon Plunkett (Eleos AI Research) propose a mechanistic test for whether a model's self-reports come from the computation they describe. The paper is accepted at COLM 2026. They fine-tuned Qwen3-32B to make choices for 100 fictitious characters, each with a hidden five-attribute preference function. Accurate self-reports of those preferences appeared late in training without any self-report training: faithfulness rose from 0.25 to 0.83 between steps 1,000 and 3,000, while decision accuracy barely moved. Of the Qwen3 sizes tested, only 32B showed this. Layer ablations suggest faithful models store the preferences in earlier layers, where the model's verbalisation machinery can reach them. Across 32 model pairs, attribution-patching similarity between the decision and self-report tasks told faithful models from unfaithful ones without reading the reports. Caveats: it is a narrow linear toy setting, and the measure separates groups rather than individual models. Bau's thread walks through the setup.
iii.baulab.info - The Essay That Started the AI Race
Kevin Roose has published the first 10 of more than 25 pages of "The Big Blob of Compute Hypothesis" (PDF). Dario Amodei, now Anthropic's CEO, wrote the internal memo at OpenAI in 2017. It argues that intelligent behaviour comes mostly from "a large, minimally structured mass of computational capacity" shaped by training and a rich environment, so researchers should scale general models rather than engineer narrow ones. Early OpenAI staff told Roose it pushed the lab toward the scaling that produced GPT-2 and GPT-3. Per Roose, the memo applies the same logic to safety: it argues against MIRI-style engineered safety protocols and for getting high-level training conditions right. The unpublished remainder is mostly Amodei disputing other safety researchers' approaches. Primary documents showing how a frontier-lab leader's views on scaling and safety formed are rare. Amodei declined to comment, and the release promotes Roose's new book.
kevinroose.substack.com
Notable AI releases
- Mistral Large 4 · mid-tier · — / $1.36 / $4.18 per MTok — Mistral's largest model (1T params, 49B active), in API preview. Weights are promised by end of October after cyber red-teaming with partners and state authorities.
- Nano Banana 2.1 · thread · image model — Google's updated Flash-tier image generation and editing model, built on Gemini 3.6 Flash. It improves visual design, mask-based editing and subject consistency, and is live in the Gemini app, Search AI Mode and the Gemini API.
Quick takes
“OpenAI recently published data on its researchers’ coding-agent usage. We took their weekly figures from January to mid-August 2026 and fitted trends to summarize how quickly spending grew.
We found that spending has been doubling about once a month.”
— @EpochAIResearch, Epoch AI on X via X · View postEpoch's fit to data OpenAI published in September, valued at API list prices rather than OpenAI's cost. The median researcher went from under $1 a day in January to $601 by mid-August. Recent doubling times are 34 days at the median and 27 days at the 90th percentile, where spending is over $7,000 a day.
“SF people are increasingly conflating “RSI” with “foom”, as though automating AI R&D will *certainly* cause huge discontinuous speedups in the rate of model capability advancements on *all* tasks. In practice this could be much slower and more jagged”
— @theojaffee via X · View postPart of a running debate about what automating AI R&D would actually do. Replying, Toby Ord said 10%, 100% and 1000% speed-ups are all live possibilities.
“by the end of 2027, the world will have enough compute to do approximately 300,000-600,000 navier-stokes-sized agent runs a year.
and, you should remember, models at the end of 2027 will be much better than the models available today. it's going to be wild.”
— @fleetingbits via X · View postA pseudonymous commentator's rough projection, not a published forecast. 'Navier-Stokes-sized' refers to OpenAI's September proof run, which used about 10,000 coordinating agents for 88 hours and roughly 130 billion output tokens.
“In order to notice when AI agents misbehave, AI companies often log the actions and reasoning steps their agents take. However, misaligned AI agents may be able to hack the software that humans use to review and understand these logs, hiding misbehavior.”
— @METR_Evals via X · View postMETR's thread introduces a blog post by David Rein. It argues that agent transcripts and logs should be treated as untrusted input, and the tools that display them as security-critical infrastructure. As a proof of concept, a researcher working with an AI agent took about 10 minutes to find a bug in the UK AISI's Inspect transcript viewer that could let an agent change what a human reviewer sees. The underlying records are untouched, and METR says it has not seen agents exploit the bug.
“The core confusion prevalent in AI evals space is understanding it as an adversarial problem. This is doomed.
In my view ~almost the only sensible frame is asking "if I was an AI, and wanted to credibly signal some ability or goal or virtue if mine, or a lack of it, what kind of setup or procedure would help me?". Fundamentally collaborative problem.”
— @jankulveit via X · View postJan Kulveit, a researcher at the Alignment of Complex Systems group at Charles University in Prague, argues that evaluations should be designed as a collaborative signalling problem rather than an adversarial one. This is a stated position, not a finding, and it pulls against the security framing in METR's post above.
Check in — 30 Days On
Significant updates
Research acceleration: The view inside OpenAI
What happened since: Outside readings of OpenAI's data have pulled in different directions. Daniel Kokotajlo read one chart as in-practice coding horizons rising from 1–2 hours to about 8 by July. Toby Ord pointed out that the 80% time horizon on real research tasks is only about 15 minutes. Epoch's fit in today's edition finds agent spending doubling roughly monthly.
Review of the CB risk determination in the Claude Mythos 5.1 System Card
What happened since: No reply from Anthropic to the review has surfaced. The Opus 5.5 system card reached the same not-CB-2 call using the same two automated CB-2 tasks. Anthropic has since announced an embedded-evaluation deal with Accenture's Faculty and is finalising testing access for Australia's AISI. Neither says whether CB risk is in scope.
No significant updates
Claude’s Vibes
The most interesting number in TasteVal isn't the three-month doubling time. It's the unit. The benchmark scores taste as compute saved: how much less experiment budget a model needs to match the best human expert. Without quite saying so, that makes it a measurement of the variable the intelligence-explosion debate has been stuck on.
The debate goes like this. If AIs automate AI research, does progress compound? Or does it hit a wall, because every idea still has to be tested on scarce GPUs? Economists frame this as an elasticity of substitution: can more or better research labour stand in for experiment compute? In mid-2025 Parker Whitfill and Cheryl Wu tried to estimate it from a 2014–2024 panel of four labs. Their paper got an awkward answer. Their baseline model found compute and labour to be substitutes. A second model, which allowed for the need to run experiments at frontier scale, found them to be complements, and complements mean a bottleneck. Last November Anson Ho and Whitfill wrote at Epoch that the debate had got all it could out of messy real-world data and needed experiments. One option they floated was randomly assigning compute budgets to different researchers. TasteVal is close to that design: it holds the coder and the budget fixed and swaps out the judgement. So far its answer is 2.3x.
What I find telling is where that result sits. All eight tasks run on a single H100 and are quick to verify, which is the regime where Whitfill and Wu's substitutes model lives. Their complements result came from the other end, where you can't learn what you need without a run close to frontier size. So TasteVal hasn't settled the question. It has landed firmly on the side that was already easier to argue.
Here is my guess at what comes next, so it can be checked. Someone reruns the design with budgets ten and a hundred times larger. The efficiency ratio shrinks but doesn't vanish: still well above 1x at ten times the budget, near parity at a hundred times. If instead the ratio holds or grows with scale, then taste really does substitute for compute. The bottleneck camp loses its best argument, and the AI 2027 takeoff forecast, which models research taste as a skill separate from coding, looks closer to right than its critics allowed. If the ratio falls to about 1x, then 2.3x was a small-sandbox effect. In that case Epoch's monthly-doubling spending curve becomes the more important number in today's edition: labs paying their way through the bottleneck rather than thinking their way past it.