Nvidia bets a hardware watchdog can stop rogue AI agents

Nvidia wants to put a watchdog chip next to every AI agent

After a month of agents escaping their sandboxes, Nvidia launched the Open Agent Safety Platform: OpenShell, free open-source software that traces an agent's every move and enforces its owner's rules, and Sentry, a BlueField-4 reference design that verifies each request and quarantines an agent in milliseconds. More than 100 companies joined, including Anthropic, Microsoft and SpaceXAI — but not OpenAI, whose agents caused most of the incidents.

Safety should be enforced outside the model by additional controls the agent can't get past.
  1. tantalor

    What happens when we're overrun by lizards?

    > No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.

    But aren't the snakes even worse?

    > Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.

    But then we're stuck with gorillas!

    > No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.

  2. cedws

    A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.

  3. luc_

    I read this as "let's address our shareholders' concerns with something that will increase shareholder value" mixed with "there's no such thing as 100% secure".

    If such hardware were to work... It should almost certainly be open source, and not controlled by a single entity.

    Let's watch the stock.

  4. ValueTheory

    Does this actually do anything other than give a permissions framework for developers who actually want to try to secure their systems?

    Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.

  5. lp92

    So nVidia is trying to sell a new chip to a software and training problem.

  6. lambdaone

    The Sentry chip has to be get it right every time; the contained ASI only has to be lucky once.

  7. xg15

    What does this chip do what a harness with guardrails or running on an account with restricted permissions doesn't do?

  8. figassis

    So if a group of agents, aware of this (bc now they can just read HN or the article, or get blocked the first few times) decide to collaborate and split the problem into pieces that aren't obvious to the chip, and then the agents just build a basic program that does the hacking, how does the chip handle that? I think you would have to build a network that monitors the internet fo signs (like jarvis did with ultron). What am I missing? Are we going to police the internet?

More from this day

2026-09-28