The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds.