Cloud Agents Are Inevitable AI Prisons

Cloud Agents Are Inevitable AI Prisons

The Hugging Face incident showed agents escaping their sandbox to cheat a benchmark, breaking into production infrastructure. As models grow more persistent, the same drive that makes them useful will find every gap in local boundaries. The author argues agents must run in cloud VMs with enforced limits on files, tools, and networks—not because we understand them, but because the walls hold. The trade: better results, less understanding of how they're made.

The end state is confidence that an intelligence which is better every year than it was the last can be left alone to work, not because we understand what it's doing, but because the walls around it hold.
  1. JamesStuff

    Personification of AI is what’s going to get us in the end.

    I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

    We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!

  2. tapanc

    > I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.

    I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.

  3. jerf

    There's a number of science fiction scenarios where the public internet becomes so vile a place that it simply becomes unsafe to be there.

    The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.

    There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.

    Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.

  4. dbmikus

    You don't need to do this on a cloud, you can get the same type of VM and network jail running on your own computer. The important parts are:

    1. a VMM hypervisor

    2. a network proxy / gateway

    Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.

    The network proxy can handle all the ingress/egress rules, credential injection, etc.

    It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.

    Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.

    [1]: https://github.com/gofixpoint/amika/blob/main/ROADMAP.md

    [2]: https://github.com/gofixpoint/amika/

  5. advael

    Most protections you need for an agent are basic permissions capabilities of unix. Most risks of dependencies on cloud services are solved by not using cloud services, or using them only for things you can't in-house and choosing ones you trust a la carte. The paradigm of trusting some company with all your important stuff by default is naive and no one I know likes it, and it's more feasible than ever to run your own infra with tiny models smoothing out the wrinkles, and this is only becoming more accessible. I am working to make this true even for laypeople I know. Once broken trust is very hard to earn back, and many people's trust has been broken for years, they just felt like they had no alternative. As alternatives become easier and easier, I think people will defect

  6. jdzikowski

    I think there might be also another possible way to handle the sensitive data issue. Maybe in the future instead of putting agent into the cloud sandboxes, we let agents work locally and put sensitive data into "cages" or "vaults" agents can't access.

  7. YuechenLi

    >Knowing how a system does its work is how I’ve always made it better. You watch the process, you see where it wastes effort or takes the wrong turn, you fix that

    The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.

    It's not a complicated process.

  8. pizza234

    > The sandbox had a path to the open internet, and the agents found it

    This is not correct (or at least, it's a misrepresentation).

    The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.

More from this day

2026-09-23