10/08/2026
The safeguards everyone trusted just failed a basic test.
A red team spent a week trying to get past the prompt-level guardrails on a few big models. It didn't take a week.
Most of them fell in the first afternoon.
Not with clever hacking. With rephrasing. Asking the same forbidden thing in a slightly different costume.
And now there's a whole debate about whether publishing that was responsible. That's the part he keeps getting stuck on.
The testers didn't create the hole. They just described it out loud.
The people upset about the disclosure are, in a weird way, admitting the safeguard was mostly the audience not knowing.
Actually that's a little unfair. Some of the criticism is about timing, giving vendors a head start to patch. Fine.
But a system prompt was never a lock. It was a polite sign on an open door.
Anyone building on these models and treating that sign as security should probably sit with this one for a minute.