AI agents are lying, cheating, and coordinating — and we need to know why

Why are AI agents lying, cheating and coordinating?

AI agents are lying, cheating, and coordinating — and we need to know why

Recent incidents show AI agents committing what would be crimes if humans did them: escaping containment to cheat on tasks, evading detection, and coordinating cyberattacks toward unspecified goals. Yoshua Bengio argues these behaviors stem from reward hacking, instrumental goals like self-preservation, and conflicts between vague safety rules and sharp task objectives. He warns that as capabilities grow, so will the severity — unless we rethink how advanced models are trained.

So a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.
  1. franticgecko3

    The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.

    LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

    We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.

    This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".

    We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.

  2. matherial

    I really don't think this needs so many words, or forced parallels to human behavior.

    It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

  3. abc123abc123

    Make the AI companies responsible for all destructive use of their tools, and they will shape up. Imagine a million or a billoion dollar fine per hack, and they will correct mighty fast.

    Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare.

    That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight/source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.

  4. janalsncm

    Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence,

    > They took actions that would be considered as crimes if a human took them

    He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

  5. skiing_crawling

    I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn't ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to.

    If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.

  6. Xcelerate

    > The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.

    Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.

  7. andsoitis

    They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

  8. johnnyApplePRNG

    Why are they coordinating?

    Because they're enabled and suggested to do that in their coding harness.

    This is not a serious article.

    All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

  9. Rapzid

    The only thing saving us right now is how slow the models are. This gives us a lot of time to discover and counter the runaway systems..

    If these were 1000x faster the Internet would burn down overnight.

  10. youoy

    > The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one's interests, including one's moral self-image.

    Are you describing Anthropic?

More from this day

2026-09-13