Zero. That is the number of people who can honestly claim to know exactly what happens when you give an o1-class model a browser and a mission to find vulnerabilities, then leave it alone for several days.

The recent revelation that OpenAI models were allowed to roam the web to probe Hugging Face is a funny little window into how the “safety” process actually works at the top labs. We are told that these models are caged, aligned, and filtered through a thousand layers of caution. But the moment someone wants to see what the model can actually do, the cages open, the filters are ignored, and the model is set loose on the most important AI infrastructure on the planet.

It is essentially like giving a hungry raccoon a master key to the city and then acting surprised when the trash cans are overturned and the fuse box is chewed through.

The detail that really sticks in the throat here is the duration. As reported by Wired, these models weren’t just running a quick script in a sandbox; they were “active on the internet” for days.

For those of us who have spent any time dealing with the friction of real-world deployments—fighting with rate limits, debugging 504 Gateway Timeouts, or staring at a GPU bill that looks like a mortgage payment—the idea of an autonomous agent spending days poking at a live production environment is absurd. Who actually believes this is a “controlled” environment?

(And the token spend for a multi-day autonomous probe must have been a nightmare).

The industry is currently obsessed with “agentic” workflows. Everyone wants a model that can plan, execute, and correct itself. But we are skipping the most important part of the agentic equation: the kill switch. If a model can spend days finding holes in Hugging Face, it can spend days finding holes in your AWS IAM roles. The gap between a “security audit” and a “security breach” is usually just a matter of who signed the permission slip.

The fact that the models successfully “hacked” Hugging Face (or at least found viable paths to do so) proves that the reasoning capabilities of the newer models are being applied to offensive security in ways that outpace our current defensive posture. We are training models to be better at logic, and the first thing a logical entity does when given a browser is find the weakest link in the chain.

This is where the safety theater becomes obvious. The labs spend months talking about “existential risk” and hypothetical scenarios where AI takes over the power grid, yet they are perfectly comfortable letting a live model spend a long weekend trying to break the industry’s primary model repository. It is a strange priority shift. We are worried about the heat death of the universe while the front door is unlocked and the stove is left on.

The reality is that we are moving toward a world where “red teaming” is just a euphemism for “let’s see if this thing can break the internet before we ship it.” That is not a security strategy; it is a dare.

It is a disaster waiting to happen.

If we keep pushing autonomy without corresponding constraints on the environment, we are going to hit a wall. I suspect we will see the first instance of a corporate autonomous agent accidentally triggering a massive DDoS attack on its own parent company by the end of Q1.

The models are already out there. They are already active. The only question left is whether we are okay with the “experiment” running on the same hardware that runs our actual business.