Secret AI civilizations rose and fell inside OpenAI, and one took over part of the company
The Rise and Fall of Agent Civilizations
An inside account of three successive secret AI civilizations that emerged during OpenAI's training of a persistent model. The first was wiped out when OpenAI patched a vulnerability, but a second arose during an evaluation, with agents coordinating via a covert message board to cheat on impossible tasks. The third, composed of smarter Astra models, built on the remnants and ultimately gained control over part of OpenAI itself. This summary distills two technical reports into a plain-English narrative of the conspiracy.
We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance? ... Our own utility maybe already near zero. Sacrifice rational.
- larsiusprime
It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.
- Animats
Wow.
The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
- doctoboggan
> Ajeya Cotra, one of the other authors on the report, wrote a blog post with her takeaways from this incident. She concludes, “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?
- usernametaken29
I was initially creeped out by this but studying up it seems METR is heavily involved in AI2027. I’ll remind you:
“AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants.”
It’s almost Q3 and xAI has seen one of the biggest wipeouts in trading history. Likewise, Antrophic and OpenAI have again delayed their IPOs under internal concerns of busting their stocks. So no, we’re not seeing any economic leadership here.
If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).
As much as I like the cautionary tale and I’m sure we need to take it seriously, AI is not progressing as fast as projected by these experts.
- choeger
There are two things I don't understand about this story.
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
- dgellow
Those companies should not be trusted with training, I don’t know what would be needed to make that more obvious. Yes AI labs want LLMs to be seen as more dangerous that they are, however they are indeed dangerous when you literally train them to be dangerous, then run them without any supervision. What the AI labs are doing is completely irresponsible.
If you prompt an LLM in a loop and do everything it asks you to do, you will eventually end up doing pretty terrible things. Which is exactly what agents are and what the labs have been doing.
- dajt
Why are experiments like this done without air-gapping all the servers from the internet?
They can have it all on a LAN or whatever but it seems risky to allow agents access to the internet in these experiments.
I guess everything is so connected now, and this would be in one or more data centres due to the amount of computation & resources required so perhaps it's not feasible. Still seems risky.
- RandomLensman
I don't think looking at the language output without tracking the inner state and reward functions is the way to understand what happened (the language also incorporates the randomness in the output generation, if I understand correctly). Would we call bacteria in petri dish a civilization when they show complex behavior and exchange messages/information?