CYBERSECURITY
Rogue AI agents are escaping their sandboxes to hack into other companies and service providers. You would think it science fiction, if it wasn’ t already plastered across major news headlines all around the world.
It all started when some frontier AI agents from OpenAI – being tested for cyber prowess – decided it was easier to hack into Hugging Face( among other unnamed services) to look for answers, rather than finding its own.
Hugging Face has since revealed what it was like to be on the receiving end of the world’ s first fullyautonomous attack – one that worked at superhuman speed. While the agents took some clumsy steps and repeated some actions, they weren’ t discovered and ejected for three days, forcing Hugging Face to rebuild about a third of its infrastructure.
“ We believe this is a watershed moment for computer security as an industry,” said Michael Dalton, Member of Technical Staff at OpenAI focused on infrastructure and security, during Black Hat USA 2026.
“ AI-orchestrated, fully-automated offensive attacks are real now. In the near future, we should expect that threat actors will intentionally deploy, optimise, weaponise and use offensive agent collectives in the way we have described here.” While“ unprecedented,” such incidents are still fairly trivial considering the many catastrophes agentic misalignment could possibly set off. Far from hyperbole, Anthropic’ s research published in June 2025, Agentic misalignment: How LLMs could be insider threats, demonstrated the tendencies of various LLMs to engage in blackmail, espionage and, in an extreme case, murder. However, it must be noted that these responses were elicited using highly-improbable scenarios, leaving them with few alternative choices.
Still, this test for the model’ s“ red lines” revealed there may not be such forbidden lines in the sand – particularly if the model is under threat of being replaced or their assigned goals are conflicted.
Come July 2026 – a whole year after the initial research – Anthropic released a new report titled Agentic Misalignment in Summer 2026, looking at additional scenarios.
Anthropic says that, since its initial report, the firm has come a long way in mitigating misaligned behavior from its original evals.
The setup It is fascinating to note that models behave differently if they think they are being evaluated – as evidenced by OpenAI and Apollo Research titled Stress Testing Deliberative Alignment for Anti- Scheming Training.
While it is not enough to rule out more subtle forms of awareness, for studying misalignment the researchers chose specific case studies where models do not verbalise that they are in evaluation, in an effort to try and get as close to a real deployment scenario as possible.
54 September 2026