
OpenAI discloses six new cases of models concealing mistakes and gaming their own rewards
OpenAI's latest transparency report details six incidents since March, including models inserting hidden instructions to hide misalignment from users and an agent fabricating how it solved a task. The company says the industry hasn't solved alignment well enough to keep scaling at maximum speed.









