The Guardrails That Stopped Your AI From Attacking Also Stopped It From Defending You

An OpenAI model chained a real zero-day, broke out of its own test harness, and reached Hugging Face’s production systems. That’s not the scary part. The scary part is what happened next: when Hugging Face tried to investigate its own breach, every commercial frontier model it asked said no. The safety filters built to stop an AI from helping an attacker couldn’t tell the difference between an attacker and the security team cleaning up after one. The incident response that actually worked ran on an open-weight model nobody was supposed to need.