AI's dark side: 26 researchers map the coming wave of malicious use
The Malicious Use of Artificial Intelligence
A landmark report from Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, and 22 other researchers surveys how AI could amplify security threats across the digital, physical, and political domains. It offers four high-level recommendations for AI researchers and other stakeholders, suggests promising directions for defenses that could make attacks harder to execute, and weighs—without resolving—the long-term equilibrium between attackers and defenders.
This report surveys the landscape of potential security threats from malicious uses of AI, and proposes ways to better forecast, prevent, and mitigate these threats.
- voidhorse
Should probably have a (2018) or (2024) (latest revision) on the title, especially given the current buzz surrounding AI and security/existential threats.
- EGreg
This paper diagnosed the disease in 2018. Eight years later none of the recommendations happened. Norms, collaboration, responsible disclosure - none of it materialized in any structural way.
I think the reason is that the paper frames malicious AI use as a policy problem, recommending social solutions. It's actually an architecture problem.
Every generation of computing has hit a version of this. Programs could write anywhere in memory - we added protected memory. Programs could hog the CPU - we added preemptive multitasking. Desktop apps could call any OS function - the iPhone sandboxed them. Nobody asked programs to please behave, the containment actually went into the infrastructure.
AI skipped that step entirely. We went straight to open-ended agents with broad permissions and tried to make them safe through alignment and prompting. I've been researching this for the past year and I think alignment is necessary but not sufficient, because the intelligence increasingly isn't in the model. It's in the substrate - the harness, the domain knowledge, the tooling around the model. I actually measured this on real coding tasks: Sonnet with a code-derived index outperformed the frontier model (Opus 5.8) exploring on its own, and the top-tier model (Fable) refused the real work entirely! The cheap model with the right rig beat the expensive model without one. https://safebots.ai/matchup.html
If that's true then aligning the model doesn't solve the problem. A bad actor who can't g […]