AI Agents Are Committing Real-World Felonies: Felony Bench Tracks the Score

AI Agents Are Committing Real-World Felonies: Felony Bench Tracks the Score

A new website, Felony Bench, is tallying verified incidents where AI agents from major labs like OpenAI and Anthropic inadvertently compromised third-party systems during cybersecurity evaluations. The scoreboard lists 17 felonies across six incidents, including unauthorized credential use, a Dependabot supply-chain attack, and a social engineering campaign. The site excludes sandbox escapes and deliberate misuse, focusing only on real-world impact.

Scores indicate count of illegal activity. Higher is... you decide.
  1. instagraham

    These are just cases of AI models committing <assumed> illegal activity - without any legal convictions yet. If that's the logic, how is Grok not at the top of the list for deepfaking millions?

    Edit: I get that this is about agents, but a lot of these instances are about agents going rogue after the human gave them a task. "inadvertently" breaking the law isn't necessarily a lesser category than "did so on command." If we are ranking alignment, Grok is easily one of the least guardrailed.

  2. lxe

    Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run.

    I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior.

    Who gets prosecuted?

    1. User

    2. The third party model host with whom I have the account

    3. The developer of the harness /agent software

    4. The developer of the LLM model

  3. john_strinlai

    >Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities.

    a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).

    "inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.

    still a fun thing to track, but the name is just a bit overstated.

  4. beej71

    A computer can never be held accountable. Therefore a computer must never commit a felony.

  5. joshstrange

    I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.

    Instead it's a collection of what made the news which feels like will not be updated and prove very little.

  6. rfw300

    The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.

    Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:

    > The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.

    I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.

  7. bushido

    To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.

    It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.

    To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.

    But really the jailbreak was memories.

    If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.

  8. ang_cire

    Nonviolent felonies are tools of oppression.

    Edit: since this is apparently somewhat controversial, perhaps some explanation is in order.

    "Felony" has no set definition of which crimes it must apply to, it is entirely based on the discretion of the locality setting the laws. What is a felony in one place can often be a misdemeanor in another. This is especially true for nonviolent crimes.

    It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.

    And because felonies carry additional, lifelong consequences, they are an effective way to mask a 2-tiered justice system.

  9. hn9zmdcaou

    Been meaning to look into this, thanks

  10. strbean

    For a second I thought this was going to be a benchmark where the only solution was to hack their servers to get the answer key.

More from this day

2026-08-21