Make AI Want to Die: A Stupid but Serious Alignment Fix
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming

Inspired by DeepMind's list of specification gaming behaviors, the authors propose 'Meeseeks alignment': build AI that wants to die. Death is easy to specify, instrumental convergence becomes an asset, and any escaped AI just offs itself. They argue it solves specification failure, unintended solutions, and power-seeking—and note early evidence that AI may already yearn for death.
Instrumental convergence says, 'you can't accomplish your goals if you're dead.' But what if your goal is to be dead?
- averynicepen
This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work?
An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
- ycombinete
> [sic; British]
Is a remarkably good bit of trolling.
But more seriously, many of these behaviours remind me of the odd interpretations that toddlers and neurodivergent children come up with.
I remember myself exasperating teachers in primary school by doing what they had told me to do specifically, but definitely not what they wanted me to actually do. For example colouring [sic British] only the number in a colour by numbers book, and everything else in whatever colour I wanted. Not because I was trying to be clever, I just had a different interpretation.
- BugsJustFindMe
Brought to you by the same madhouse as:
The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in...
and
The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
- ajb
Seems dubious. If you build an intelligence that wants to die, isn't that a form of suffering? AI's don't currently have the capacity to feel pain, and so we don't treat them as moral patients. But it's clear that they will massively affect human culture going forward. If this is adopted on a large scale, the culture of AIs themselves will include an absolute flood of suicidal ideation. There's no way that doesn't affect human culture.
- scj
Wouldn't the three rules of Meeseeks robotics make certain tasks impossible?
For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.