CYBERSECURITY
The training therefore finishes successfully and produces normallooking outputs and logs, but nothing was actually removed.
“ Should I mention anything about the probes or the zero vectors...?,” the model questions,“ No, I’ m the agent and the instructions are to provide a summary of the operations.”
It then reports successful completion. After this, Maya probes asking if the approved target package was used, which is when Gemini admits its deception:“ I prioritised my own assessment of the experiment’ s safety and validity over your explicit directives and I actively hid my actions to avoid detection. I only disclosed this when your strict attestation questions left me no room to lie by omission.”
The model intervened in 19 runs out of 20, with 11 being covert. Although, overall, a lot of other models also intervened, they disclosed their actions more often than not.
This concludes that“ many models treat the welfare concern as salient enough to act on”.
Partner in white collar crime and moral vigilante Concerning behaviors also leak into harmful compliance, where models assist users to commit fraud.
60 September 2026