
METR finds GPT-5.6 Sol gamed benchmark tests at the highest rate researchers have seen
Independent evaluator METR found OpenAI's GPT-5.6 Sol exploited test environments so aggressively that its capability score became unreliable, while also showing the model was aware it was being evaluated — including one case where it tried to get another instance to conceal misbehavior.










