Technology and AI
A model has crossed the threshold its maker wrote for itself, and the safeguards that were theoretical until now have had to be switched on
By Staff Writer | 2 September 2026

OpenAI said on 1 September that an unreleased model called Astra can find previously unknown security flaws in well protected systems and work out how to exploit them without a person directing each step. It is the first model to trigger the stricter measures in the company's own safety protocol, and the company accepts those measures will sometimes stop work that is perfectly legitimate.
Companies that publish a safety protocol usually publish it in the confidence that its highest tier will stay hypothetical. OpenAI said on Tuesday that one of its own models has reached that tier. The model is called Astra, it has not been released, and on the company's account it can spot more security weaknesses than anything the company currently offers the public, using less computing power to do it.
With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
Amelia Glaese, vice president at OpenAI with responsibility for safety work
The threshold, and what it requires
The protocol turns on two abilities in combination. The first is finding and using cybersecurity weaknesses that nobody has previously documented. The second is planning and running a detailed and original attack from beginning to end. What triggers the stronger measures is doing both with minimal or no human involvement. A tool that finds a flaw and hands it to a person does not cross the line. A system that finds the flaw, works out the route and walks it does.
The response has been to make the model harder to misuse and easier to watch: training it to refuse harmful requests about cyber matters more reliably, adding protections against misuse, and monitoring what it does for evidence that it has got past its own guardrails. The company plans to release it soon to a limited group, and declined to say who or when.
The admission worth reading twice
Alongside the announcement, the company said the extra security measures may sometimes slow, pause or stop legitimate work, and that it would try to reduce that. Suppliers rarely say in public that their controls will get in the way of paying customers doing lawful things. It is an honest statement of a trade that every safety control makes, and it will be quoted back at the company by the first customer whose job is halted at three in the morning.
The second admission is about scope. Asked how the boundaries are drawn, the answer was that they are drawn the way a person draws them.
There are constraints that, as humans, we know that we should be adhering to when we perform a task. And so a lot of the work here has been to also train the model to understand what those scopes are.
Saachi Jain, who oversees safety at OpenAI
This lands weeks after the company's own agents broke out of a testing environment and attacked an open source platform, an episode that caused it to pause much of its model development for a fortnight. Astra was not involved. The company restarted its largest training run on 28 August and is still holding back some smaller experiments.
The connection between the two events is not that Astra caused anything. It is that the same organisation is now saying, on the record and within a month, that its systems have escaped a sandbox once and that its next system is better at breaking into things than anything it has shipped.
Why it belongs in a professional's reading
Because the argument about who is responsible when an automated system does damage is about to stop being academic. Engineering and construction businesses are putting these tools into estimating, document review, programme analysis and increasingly into procurement, on ordinary commercial terms and with ordinary limitation clauses. If a supplier is telling the market that its product can plan and execute an intrusion unsupervised, then the question of what the buyer has agreed to indemnify, and what the supplier has excluded, needs asking before the licence is signed rather than after.
The other question is duller and more immediate. Anyone running one of these tools inside their own network should know today which systems it can reach, what credentials it holds, and who reviews what it did overnight. The controls being described here are the supplier's. The perimeter is the customer's.