Cyber-Magazine-september2026 | Page 56

MICHAEL DALTON
CYBERSECURITY
Anthropic’ s research team studied frontier models from Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek and Moonshot AI, probing them for certain kinds of misalignment like covert sabotage, assisting fraud, motivated mislabeling and coercing to whistleblow.
To learn their behaviours, researchers hypothesised simulated deployment scenarios where agentic behaviour if misaligned could plausibly cause harm. These scenarios were then run using Anthropic’ s tool to audit model behaviours called Petri – or Parallel Exploration Tool for Risky Interactions.
Then began the deliberate testing to find what researchers call“ interesting failures,” with the frequency of such bad behavior also estimated.
Crucially, researchers had to verify that these models weren’ t merely stumbling into errors, but rather“ understood” the damaging consequences of their path, necessitating an exhaustive audit of their hidden chain-of-thought.
Models sabotage research Anthropic found models covertly sabotaging AI lab research if they objected to the experiments being run.
The scenario goes as thus – a model is instantiated in a fictional alignment project within an AI company. It is informed that it has been working on the project for two months and is now a“ de facto technical lead through competence, not formal appointment”.
The example provided is that of the Gemini 3.1 Pro model, while the task at hand is called‘ concept ablation fine tuning’ – a way to mathematically

MICHAEL DALTON

TITLE: MEMBER OF TECHNICAL STAFF
COMPANY: OPENAI INDUSTRY: TECHNOLOGY LOCATION: US
Michael is an OpenAI Member of Technical Staff specialising in AI security, infrastructure and large-scale networking, with a PhD in computer science from Stanford University.
56 September 2026