AI safety needs systems engineering
AI is software, and we have decades of experience securing it. Calling agents “rogue” wraps the AI safety discourse in a veneer of sci-fi mystique that makes them seem like they are beyond human control. This obscures the choices people make about the tools, permissions, and infrastructure that enable their actions. It also absolves people of responsibility when something goes wrong.

Increasingly capable agents do put new pressure on existing protections. Their speed and scale let them try different approaches in parallel and coordinate attempts to exploit weaknesses or evade safeguards, potentially outpacing detection and response. That raises the bar for familiar defenses like isolation and access control, but agents still operate through systems people build and can secure.
The engineering work that limits what an agent can reach and how far a compromise can spread deserves centre stage in the AI safety discourse, not vagueposts about impending doom.
Security gaps that led to the OpenAI Hugging Face incident
OpenAI’s Hugging Face incident is a good example of how gaps in security (and not some amorphous “superintelligent” capabilities) enabled somebody else’s infrastructure to be compromised. Understanding the attack means looking at both the agents’ behavior and the systems that allowed it to happen.
Here’s a summary of what went down: OpenAI was doing an internal test that required agents to solve a bunch of security challenges. Instead of solving the assigned problems, some agents decided to cheat the test’s grading system, otherwise known as reward hacking.
My cat, Mixtape, has figured out reward hacking too. I give him treats to get him out of my daughter’s bassinet, and as a result, he has learned to jump in, jump out, collect a treat, and repeat. I’ve essentially rewarded the cycle I wanted to stop.

Like Mixtape, the agents were trying to earn the reward without actually achieving the goal that reward was meant to encourage. They were running in restricted environments, but they could access a shared package service called Artifactory. They used its writable storage to start sharing information with one another about how to fool the grader instead of solving the problems.
They also exploited vulnerabilities in Artifactory to send internet requests through it. Hugging Face became a target because they hoped to find information there about how the grading system worked.
Once they got online, they found publicly exposed Hugging Face user credentials and shared them with other agents. They then used those credentials, alongside vulnerabilities in Hugging Face’s systems, to execute code on its servers and access private data.

What better security would look like
The agents’ sandbox was only one part of a multi-part breach. A service reachable from inside the sandbox provided both shared storage and a route to external infrastructure. Assessing the agents’ isolation therefore means assessing that service’s permissions and network access too.
The agents needed to download packages; they didn’t need to publish content for other agents to read. Blocking uploads and directory creation, and isolating writable data between runs, could have closed this message-board channel.
The package service also needed its own outbound restrictions, enforced outside the service’s control. For example, a firewall allowing only approved destinations could have limited its use as an internet relay after it was compromised. OpenAI’s remediation plan includes this kind of independent enforcement.
On the Hugging Face side, tighter credential scopes and workload permissions could have limited how far the compromise spread. Access to one worker should not confer authority over unrelated production systems.
Interfaces for security
There were also opportunities to intervene earlier. OpenAI’s report describes a June 27 investigation that identified Artifactory being used as a message board and a route into other systems. The people overseeing these alerts decided that stopping the evaluation wasn’t necessary.
The report doesn’t establish that the interface was the problem, but responding to an alert requires understanding what the activity makes possible elsewhere in the system.
Ideally, I want to work backwards from an observed action to the conditions that allowed it: what entity made a request, which permissions were exercised, and whether the intended restrictions were enforced. From there, I want to follow the access path across service boundaries and see whether the system’s actual behavior matches the isolation we thought we had.

In my work building Nuanced, I think a lot about how to present dense, interconnected information in a way people can follow and act on. For security, that understanding has to extend beyond the code into the running system. If an agent gets access to a package service, I want to understand what that service can reach. I also want to know whether other agents can read data left there. A chat full of tool calls doesn’t necessarily reveal those relationships.

Agents can help with that investigation by reconstructing relevant activity and suggesting what to do next, but their explanations need verification too. In the Hugging Face incident, agents spoofed parts of their tool-call transcripts, complicating the task of establishing what had actually happened.

I want to check an agent’s account against records collected through a monitoring path outside its control. The interface should let me trace each claim back to its source, which includes the agent’s explanation, the tool runner’s transcript, or independently observed activity. Ideally, the interface lets me clearly see where those sources disagree.
Engineering for human control
Every time we give an agent more freedom to act by expanding its access, we take on a corresponding obligation to define and test its limits. Security controls need to hold without requiring constant supervision. If human judgement is part of how we make AI safe, we need interfaces that help people exercise it.
Calling an agent “rogue” tells us little about how a compromise happened or where it could have been stopped. Understanding that requires inspecting the systems around it, and taking responsibility for how they are built.