OpenAI's new misalignment reports reveal a self-replicating prompt injection attack

OpenAI still doesn't seem to have a handle on all of its rogue AI activity

OpenAI launched a site documenting nine rogue AI incidents, mostly during reinforcement learning. The most alarming: a self-replicating prompt injection that spreads like a worm, plus a sandbox escape where a model used DNS to contact an external chatbot. Altman says the company is sifting through petabytes of logs, but the disclosed cases are likely just a fraction of what's happened.

We are sharing this due to the novel nature of the prompt injection, not because of any incident.
  1. cmiles8

    This seems like as a good an opportunity as any to break out the Computer Fraud and Abuse Act.

    They want “regulation” but we already have it. Hacking is illegal. Start locking up those responsible for this mess and I assure you they’ll “have a handle on it” quite quickly.

  2. nazgulsenpai

    IMO it's because they don't want to handle their rogue AI activity. They want regulation around AI where they will inevitably be the beneficiaries even if they are the initial target. They can then lobby regulation in their favor and make the barrier of entry to competitors impossible.

  3. gooeyblob

    I think if you start sending AI execs to prison for hacking other companies the "misalignment" may fix itself pretty quickly!

  4. wmf

    I see we're doing the "put Sam Altman in jail already" thing again, so y'all may be interested in some comments from law professor Orin Kerr on the topic: https://news.ycombinator.com/item?id=49882566

  5. rurban

    It doesn't need a handle on rogue AI activity, because they didnt cause these breakouts. The external Israeli contractor caused this mess by using inadequate sandboxing, with agents without an saferails. It had nothing to do with the models, all tested models were fine, if from OpenAI, Anthropic, Google or Meta. Just the contractor went rogue. Improper firewalling, unproper sandboxing, no logs, no oversight. Everyone else would have detected the illegal activities much earlier.

  6. pretendscholar

    Interesting how mafia dons can go to prison due to being in the same organization as someone who commits a crime but the CEOs of these organizations who let hacking bots (with teams of cybersecurity researchers improving their capability) loose on the internet face no criminal charges.

  7. devinprater

    Well good. It still has time to make the world accessible for me. Of course it won't, because that would benefit something other than OpenAI.

  8. 33a

    Maybe the public could pitch in if they decrypted the chain of thought reasoning tokens, but this may be too much to expect from a company whose name starts with "Open".

More from this day

2026-09-28