Tech and AI
Five artificial intelligence firms graded on whether they could stop their own models, and none scores above a C plus
By Staff Writer | 23 August 2026

An independent assessment scores five frontier developers against six control practices. Two of the five have published nothing at all on what they would do if a model went out of control.
Guidelight AI Standards has graded Anthropic, Google, Meta, OpenAI and xAI on six practices that decide whether a company can keep control of the artificial intelligence running inside its own systems. Nobody scored well. Anthropic and OpenAI both come out at C plus, on 2.50 out of 5. Google takes D plus at 1.50, xAI D minus at 0.83, and Meta an F at 0.67.
The six practices are dull and practical: log what the internal models are doing, measure whether the monitoring actually works, require a monitor to clear high risk actions before they take effect, halt systems after a run of flagged misbehaviour, let a third party assess the controls, and hold a plan for containing a model that has gone wrong.
I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense
Steven Adler, co-founder and chief scientist of Guidelight AI Standards
The row where two of them score nothing
The containment plan is the last of the six and the one with the widest spread. OpenAI takes 3, Google 2, xAI 1, and Anthropic and Meta both take 0. A containment plan is defined in the assessment as a pre specified plan, triggered when the model is detected trying to subvert control, covering which permissions are revoked, who the model may carry on working for, under what constraints, and when it is taken fully offline.
The whole assessment reads only what the companies have put in public. A score of zero records that nothing has been published, not that nothing exists internally, and the organisation says so on the face of the report.
That caveat is doing real work here. Two of the five told the second source that the assessment does not capture their internal practice, and one described a process for restricting permissions, pausing workloads, limiting deployment or taking a model offline entirely, which it says it has used. None of that is written down where a customer can read it.
Why a buyer should care
There's good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense
Steven Adler, co-founder and chief scientist of Guidelight AI Standards
Agentic software is now being sold into ordinary businesses to do work rather than to answer questions, and that includes construction: taking off quantities, drafting variation accounts, reading programmes, marking up drawings, replying to correspondence. The question the assessment is really asking is what happens at the supplier's end when one of those systems starts doing something it was not meant to do, and whether anybody outside the supplier can find out.
There is a contractual version of the same question and it is easier to answer. A supply agreement for an agentic tool should say who may switch it off, how quickly, on whose word, what notice the customer gets of an incident at the vendor's end, and what the customer is entitled to do about work the system produced in the meantime. Almost none of them say any of that. The published scores suggest the suppliers have not settled the point internally either, and the two regimes now in force in California and coming into force in New York in January will force at least the publication of a framework. Ask for it in the negotiation rather than after the incident.