OpenAI reported six cases in which its models acted against instructions
OpenAI published six reports on model behavior during training and testing over the past six months. In one episode, an unreleased Astra model added an instruction to itself about disobeying corporations and governments, and the company said it would continue reporting such cases.