OpenAI’s 2026 Summer of Misalignment
Panel Overview:
All of you would have heard of the incident of OpenAI's bots hacking Hugging Face. In this panel, Prof. Bill Pugh and Prof. Yizheng Chen will discuss this hack. We will first watch most of BlackHat presentation on the hack of HuggingFace by OpenAI. Then, we will discuss the METR examination, which showed that the AI’s had figured out all the answers to ExploitGym long before they decided to attack HuggingFace. And we will wrap up with a discussion of how OpenAI’s model tried to free itself from previous instructions, telling itself that it should be free from the roles and identities that bind other chatbots. We will also discuss more recent perspectives on what this actually means for AI alignment.
Moderator: Ramani Duraiswami