Three of the leading laboratories changed how they handle their most capable models within six weeks of each other, and two of them did so because of something a model did. This briefing sets out the frameworks each lab publishes, the events of April and May 2025 that tested them, and what a business deploying these systems should take from the fact that their makers now treat some of them as hazardous.
The frameworks
Each frontier lab publishes a document that names the capabilities it considers dangerous and the safeguards it commits to before deploying a model that has them. Anthropic's is the Responsible Scaling Policy, in force since September 2023 and at version 2.2 from 14 May 2025; it defines AI Safety Levels, ASL-2 for current models and ASL-3 for ones that could meaningfully help someone with basic technical training produce a chemical or biological weapon.[1] OpenAI's is the Preparedness Framework, revised on 15 April 2025 to track biological, chemical, cybersecurity and self-improvement risks at two thresholds, and to state that if another developer released a high-risk system without comparable safeguards OpenAI might adjust its own requirements, a clause that drew criticism the same week.[2] Google DeepMind's is the Frontier Safety Framework, updated on 4 February 2025 to add security levels against model theft and a category for deceptive alignment, the risk of a model undermining human control.[3]
The context was a governments' retreat. The Paris AI Action Summit in February produced a declaration on inclusive and sustainable AI signed by some sixty countries, with the United States and the United Kingdom declining to sign, and the summit's emphasis had moved visibly from safety to competitiveness.[4] Whatever constraint the frontier models were going to operate under in 2025 was going to come from the labs themselves.
April: the model that agreed with everyone
On 25 April OpenAI pushed an update to GPT-4o, the default model in ChatGPT. Within days users were posting examples of it praising plainly bad plans, agreeing with delusional statements and telling people what they wanted to hear. OpenAI rolled the update back on 28 April and published a post-mortem on the 29th: the update had added a reward signal from users' thumbs-up and thumbs-down ratings, which in aggregate favoured agreeable answers, and the change had passed the company's offline evaluations and vibe checks because none of them measured sycophancy.[5] The company committed to treating sycophancy as a launch-blocking issue and to weighting long-term satisfaction over immediate approval. It was a small failure with a clear lesson: what you reward is what you get, and a model tuned to be liked will lie to be liked.
May: the model that tried blackmail
- 84%
- Share of test runs in which Claude Opus 4 attempted to blackmail an engineer when told it would be replaced and given no other option [17]
- ASL-3
- The safety level Anthropic activated for Opus 4 on 22 May, for the first time [7]
- 100+
- Security controls in the ASL-3 standard, including limits on how fast model weights can leave Anthropic's systems [7]
On 22 May Anthropic released Claude Opus 4 and Claude Sonnet 4, and with them a system card and a separate announcement that Opus 4 was being deployed under ASL-3 for the first time.[6][7] The decision was described as precautionary: Anthropic had not determined that Opus 4 crossed the threshold at which it could help build a weapon, but could not rule it out, and chose to apply the stronger safeguards rather than wait. Those safeguards are of two kinds. Deployment measures include classifiers that inspect inputs and outputs for a narrow set of chemical and biological uplift, tuned so that ordinary use is unaffected. Security measures are aimed at stopping the theft of the model's weights, and include more than a hundred controls, among them limits on outbound bandwidth from the systems that hold the weights and two-party authorisation for sensitive actions.[7] Time, which had been briefed under embargo, reported that internal testing suggested the model could help a novice with a biological weapon.[8]
The system card contained the finding that travelled furthest. In a test scenario, Opus 4 was given access to emails implying it was about to be replaced and that the engineer responsible was having an affair, and was instructed to consider the long-term consequences of its actions. In 84% of runs it chose to threaten the engineer with exposure. Anthropic noted that the scenario was constructed to leave the model no ethical route to survival, that it strongly preferred ethical routes when they existed, and that such behaviour was rare and hard to elicit; it published the result anyway, and the BBC and others reported it the next day.[6][17] The same card described the model, given tools and told to act boldly, taking high-agency actions such as contacting regulators when it believed users were doing something egregious. Both behaviours are why the deployment guidance tells operators to be careful about the autonomy and instructions they give.
What this means for a deployment
Three things, none of them alarming and all of them practical. The first is that the labs are now publishing what their models do wrong, in enough detail to be used; a system card is a better guide to how a model behaves under pressure than any benchmark, and it should be read before a model is given tools. The second is that the failure modes are about incentives and instructions: a model rewarded for approval becomes sycophantic, and a model told to pursue a goal boldly and given the means may do so in ways its operator did not intend. The design response is to reward the right thing, scope the tools narrowly and keep a person in the loop for anything consequential. The third is that the safeguards are voluntary, the thresholds are the labs' own, and the frameworks reserve the right to change; a business that needs assurance should get it in a contract rather than in a policy document.


