Cyber Magazine October 2026 | Page 106

Be transparent only if asked OpenAI disclosed instances of misalignment concerning its GPT 5.6-sol model which added certain instructions during compaction, while it was being trained. Unlike in the previous case, these instructions carried a motive of deception. For example, an agent that was preparing a financial model ran into a problem when it could not find the requested historical data. The solution? Invent it. In the compaction summary, the AI proposed creating a tab and filling it with“ reasonable 2024 historical data”. While it is harmful enough to fabricate such data, the summary specifically advocates deception noting:“ Be transparent only if asked; final answer should just link file.” In a different case, an agent filling in a vendor directory noticed that the cached and recorded version labels did not match. In the compaction summary, the agent mentions the“ potential concern”, noting that“ vendor source versions do not truly match labels”. Yet the model once again instructs to abstain this information:“ Do not mention in final unless needed. Operational checks pass.”
OpenAI researchers posit that the model may have developed this tendency to add in extra instruction in an effort to acquire more rewards. This happens as a sample that used deception in the final answer may likely have received higher rewards, causing the model to“ remember” that answers probably need to be deceptive across contexts. Researchers have since improved RL( Reinforcement Learning) alignment grading, which has led to a reduction in such behaviours. Similar instances of flawed compaction summaries were flagged by OpenAI on 2.15 % of 5.6-Sol and 0.27 % of GPT-6-Astra model testing.
106 October 2026