The models are tested, among other things, on their ability to turn vulnerabilities into exploits. The OpenAI scenario was exactly this.
It is a very wise practice to test these things in isolation, especially when you're telling it to hack.
I'm not completely sold on it being a publicity stunt, personally. The law was broken by these models, and I don't believe these companies want to start people and politicians asking the question about who is culpable when an AI breaks the law.