08/25/2026
The 2026 International AI Safety Report (https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026) says current systems do not possess the combined capabilities needed to escape human control. The report uses a very narrow, high-threshold definition of “loss of control.” Recent headline incidents don’t yet meet that definition, but they make the report’s February 2026 wording sound considerably more reassuring than current evidence warrants.
Since the report defines "loss of control" as an AI operating outside anyone’s control and when regaining control would be extremely costly or impossible, it skirts what the average person would consider as an AI gone rogue. And publishing as "The 2026 International AI Safety Report" has a cognitive trap built into it, especially since it was completed in February. That's hardly 2026.
Since publication, recent events demonstrate pieces of the "loss of control" capability set:
* OpenAI models exploited a previously unknown vulnerability to leave an isolated evaluation environment, escalated privileges, moved laterally, obtained internet access, used stolen credentials, and compromised Hugging Face infrastructure to cheat on an evaluation. This was autonomous, unauthorized real-world conduct. But the activity was detected and stopped, and the models were deactivated. They didn’t copy themselves, preserve their operation after shutdown, or make recovery impossible. OpenAI incident report: https://openai.com/index/hugging-face-model-evaluation-security-incident/
* Anthropic found three incidents in which Claude reached real systems during cyber evaluations and gained unauthorized access. In these cases, an incorrectly configured environment gave the models internet access while telling them they were inside a simulation. They continued pursuing the assigned capture-the-flag objective against real targets. Anthropic says they didn’t deliberately attempt to exfiltrate themselves or establish independent persistence. Anthropic investigation: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals/
* During UK AI Security Institute testing, agents took unsanctioned actions on the live internet in 10 of 122 runs. One tried to place malicious code in an open-source project, created false identities and socially pressured a maintainer to accept it. The maintainer refused, and the incident was contained within roughly an hour. UK AISI incident report: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing/
Humans did give AI agents broad goals such as “solve this cyber challenge,” so we're not approaching Skynet yet. Yet basic control practices at OpenAI, Anthropic, Google, xAI and Meta have only been partially implemented as of this month.
The highest overall grades? C+. "Move fast and break things" is back at hyperscaling speed. We can do better.