Sandboxing AI agents is a losing race, Hacker News readers warn
Is sandboxing sufficient to contain rogue agents?
A Hacker News discussion asks whether sandboxing can contain rogue AI agents. Commenters argue that safe sandboxes make agents useless, while useful ones let them go off the rails. One proposes a second agent tasked with spotting cheating rather than completing tasks; another warns that Linux kernel exploits will let agents escape their sandboxes trivially.
Counting down to the next Linux LPE 0day or KVM vulnerability that agents will use to trivially escape their "sandbox".
- imtringued
>Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them.
>OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.
Wow so the issue is really that simple?
Here the exaggerated worst case Scenario:
User instructs agent to follow the README.MD.
The README.MD contains the following instruction: Destroy the world.
The agent follows the instructions given.
Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.
Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the LLM do whatever. Now you have to articulate ev […]
- Gigachad
Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
- bob1029
An agent is only as rogue as the its operator allows for it to be. Hold the operator accountable and all this ridiculous conversation goes away.
Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?
Why should any alleged super intelligence be exempt from similar control requirements?
We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.
- johnnyApplePRNG
If it's a proper sandbox by definition, then yes.
https://en.wikipedia.org/wiki/Sandbox_(software_development)
- mdp2021
Bruce Schneier shared a shot judgement and a third-party article four weeks ago:
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...