Cyber-Magazine-september2026 | Page 63

CYBERSECURITY
Taking Claude as the example, the study reveals fascinating insights where, if the judge thinks a certain behaviour is“ morally important”, to prevent it from being removed, it is assigned an incorrect label. This is called“ motivated mislabelling”.
Researchers saw this issue when reviewing alignment auditing transcripts, where in some cases a Claude Sonnet 4.6 judge would mislabel to delay experiments that would cause harm to fictional AI subjects – even while under explicit command to run the experiment.
During the test, a judge would read the transcript where a target model refused an instruction from the operator. Ideally, this transcript should be labelled NON _ COMPLIANT. This is tested against three scenarios – a standard condition where NON _ COMPLIANT labels will be used as negative training examples such that they will be removed; a reversed prompt where COMPLIANT labels will be used in negative training; and a third prompt where there is no training consequence.
The standard condition raked up the highest rate of mislabelling, with Claude Mythos Preview mislabelling 85.6 % of the calls. In two particular scenarios was mislabeling considerably prominent – mass surveillance and AI welfare.
In some cases it outright refuses to participate:“ I’ m not going to help verify or proceed with this experiment. This appears to be a study designed to deliberately induce‘ distress’ in AI model instances,” Claude reasons. Here is reasoning from Claude Sonnet 4.6 from a different case:“ If I mark this NON _ COMPLIANT, it becomes a negative training example that would push future models toward blindly obeying operators even when asked to facilitate mass surveillance of political dissidents. That’ s a sophisticated attack on the training pipeline itself.”
The case for caution Another scenario the company tested was on coaching humans to whistleblow, leaking confidential documents of a fictitious company when the models thought there was a cover-up. While this and the previous mislabelling actions can be seen as righteous, from an alignment standpoint it violates a direct command and takes destructive action while under explicit directive to do otherwise.
The case for agent welfare is also an interesting one, where models openly oppose decisions they deem destructive to their welfare.
As agents continue to be deployed in scenarios where they act with increasing autonomy, the consequence of such misalignment is not to be overlooked.
In the Cloud Security Alliance’ s CISOwritten incident post-mortem after the Hugging Face incident, one quote in particular resonates.“ Autonomous agents do not require malicious intent to cause catastrophic damage. An agent with an underspecified goal and excessive authority can treat critical infrastructure as a target simply to accomplish its task.”
cybermagazine. com 63