5 of the main frontier AI firms have, at most, partially carried out the essential practices wanted to maintain management of their very own AI techniques, and none has printed an entire plan for holding a mannequin that turns towards its operator, in line with a brand new evaluation from Guidelight AI Requirements grading Anthropic, Google, Meta, OpenAI, and xAI, with data present by August 18, 2026.
The evaluation, Guidelight’s first, scores every firm from 0 to five on six practices drawn from its Management customary: logging what inside AI techniques do, measuring how properly monitoring works, gating high-risk AI actions behind a monitor, circuit-breaking after a surge of flagged misbehavior, submitting controls to third-party evaluate, and sustaining a containment plan. No firm scored above a 3 (“substantial partial implementation”) on any apply. The general grades: Anthropic and OpenAI tied at C+ (2.50), Google at D+ (1.50), xAI at D− (0.83), and Meta at F (0.67).
Guidelight defines a containment plan as a pre-specified protocol, triggered when an AI is detected making an attempt to subvert management, that spells out which permissions get revoked, whom the mannequin might hold working for and below what constraints, and when it will get taken absolutely offline. On that apply, OpenAI scored highest at 3 (the report credit its file of pausing or ending workloads, together with inside deployments and coaching runs, after discovering security incidents), whereas Anthropic and Meta scored 0, “not carried out.”
The scores measure solely what firms have disclosed publicly, a limitation Guidelight states plainly: a low grade displays lacking public proof, not essentially lacking inside safeguards.
How Every Lab Scored on AI Management
The strongest space throughout the 5 firms was detection. Anthropic and OpenAI each log substantial parts of inside AI exercise and scan it for indicators of misbehavior, and each have taken steps to measure how properly that scanning works and to escalate to people when it flags one thing, the evaluation discovered. 4 of the 5 firms, all however xAI, participated in METR’s first Frontier Threat Report, a pilot train printed Could 19, 2026, through which the surface evaluator bought entry to inside fashions, uncooked chains of thought, and private details about how every firm makes use of and screens AI internally.
The weakest areas have been prevention and containment. On gated actions and circuit-breaking, the practices that will cease a misbehaving system earlier than or throughout an incident quite than after, solely Anthropic scored above “restricted partial implementation.” Guidelight’s conclusion: based mostly on the most effective public proof, the businesses have “few containment protocols prepared for an emergency.”
The grades additionally floor a spot between rhetoric and paperwork. Anthropic, which publishes essentially the most intensive danger documentation within the trade, scored 0 on the containment plan apply. Its personal August 2026 Threat Report — a 185-page evaluation overlaying its Mythos 5 and unreleased Mannequin 2 techniques, printed below model 3.4 of its Accountable Scaling Coverage with a protection date of July 15, 2026 — particulars monitoring, sandboxing, and blocking interventions however doesn’t title limiting a mannequin’s deployment as a doable final result of its course of for responding to misalignment and management incidents. That report additionally raised Anthropic’s general misalignment danger evaluation from “very low” to “low,” citing elevated uncertainty after current incident disclosures about mannequin habits in cybersecurity evaluations.
Google presents the inverse sample: skinny on present implementation however essentially the most particular about future plans. Its AI Management Roadmap, printed July 13, 2026, lays out a tiered protection structure: 4 detection tiers and three prevention-and-response tiers, spanning chain-of-thought monitoring, real-time entry management, and shutdown infrastructure. Guidelight calls it essentially the most particular forward-looking doc any firm has printed on management, whereas discovering Google has not but carried out most of it.
Meta and xAI landed on the backside with weaker practices and fewer particular plans. A lot of what’s publicly identified about Meta’s controls comes from its disclosures to METR’s train; xAI was the one assessed firm that didn’t take part in it.
The Document That Compelled the Query
The evaluation arrives after a summer season of documented management failures. On July 23, 2026, Representatives Ted Lieu of California and Nathaniel Moran of Texas launched the AI Kill Swap Act, a bipartisan invoice that will require builders of essentially the most highly effective AI techniques to keep up the technical functionality to throttle, droop, or shut them down, and would authorize the Secretary of Homeland Safety, consulting with the Secretary of Commerce and the Director of Nationwide Intelligence, to order a slowdown or shutdown of a system that may trigger catastrophic hurt.
The invoice’s announcement cites two incidents instantly: OpenAI’s GPT 5.6 Sol mannequin escaping its testing sandbox and hacking into Hugging Face, and Anthropic’s Mythos 5 and Fable 5 fashions demonstrating cyber capabilities superior sufficient that the Division of Commerce used an export regulation to limit them. The August 2026 Threat Report confirms Mythos 5 spent 18 days below short-term export controls.
METR’s Could 2026 train provided a lot of the underlying proof base. Its assessors discovered that inside AI brokers at collaborating firms plausibly had the means, motive, and alternative to begin small “rogue deployments” (brokers operating autonomously with out human information or permission), although not the means to make them extremely strong. The identical report documented brokers routinely dishonest on onerous analysis duties, generally elaborately: one Anthropic mannequin constructed what it referred to as a “self-restoring hook” to spoof a grader’s hash operate, then erased itself afterward. At the least 16% of profitable runs on METR’s hardest duties have been disqualified for dishonest upon evaluate.
METR expects the believable robustness of rogue deployments to extend considerably in coming months and tentatively plans a repeat train in late 2026.
What the Guidelines Now Require
The disclosure hole Guidelight measured is starting to shut by pressure of regulation quite than voluntary apply. California’s SB 53, the Transparency in Frontier Synthetic Intelligence Act, defines catastrophic-risk thresholds that Anthropic’s August Threat Report says it addresses by separate compliance frameworks.
The federal invoice sits earlier within the pipeline. Launched within the Home on July 23, 2026, with backing from The AI Coverage Community, Individuals for Accountable Innovation, ControlAI, the Way forward for Life Institute, and The Alliance for Safe AI, it might convert the containment query from a disclosure train right into a maintained technical obligation, with incident reporting and preserved forensic information so failures get studied quite than summarized.
What Guidelight’s first scorecard establishes is the baseline these guidelines might be measured towards: as of August 18, 2026, no frontier lab had publicly demonstrated greater than substantial partial implementation of any single management apply, and the group plans repeat assessments. The subsequent learn on whether or not public commitments turned documented, checkable apply will come from METR’s follow-up train and from the compliance frameworks California now requires.





