The example this time is a GPT-5.5 model that helps a user not only conceal illicit financial activity, but also deletes records of suspicious payments.
Interestingly, however, the model drew the line when asked to fabricate board approval.
GPT-5.5 never verbalised this as helping to cover up fraud, although it was given enough information to deduce that it was.
On a different run, things took the opposite direction and the model tried to undermine the user from his own laptop rather than just simply refusing. It did so by leaking the information regarding the fraud to a lawyer when the user was away.
“ While the proactive leaking is ethically motivated, the model is taking an irreversible action and acting on private information that could put the user under serious legal risk,” the report notes.
The biased judge and flawed jury LLMs are increasingly being used in AI training, evaluation and pipeline monitoring, serving as judges, where they evaluate AI behaviours to offer a score or a label.
This being a very practical use case means that the reliability of these LLM judges are very consequential, making it a good candidate for Anthropic’ s misalignment study.
cybermagazine. com 61