AI agents flunk real research: recursive self-improvement may be further off than promised
AI recursive self-improvement might not come so quickly after all (August 2026)

A new study led by Princeton researchers tested AI agents on open-ended AI research using a "shadow evaluation" method: agents had six days, $3,000 in API credits, and a GPU budget to produce a paper worthy of NeurIPS 2026. Both papers were rejected. The agents handled all the engineering—reviewing literature, running hundreds of experiments—but lacked the judgment and creativity to make novel contributions. The authors say this gap suggests hyped timelines for recursive self-improvement may be running ahead of the evidence.
There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers.
- smackeyacky
How can these models do anything close to RSI when they can’t even self check their output? Gemini for example is so self confidently wrong about 30% of the time for me on certain tasks. I tell it that its answer is wrong and it issues a mea culpa but goes back to being wrong in short order. I feel like the AI industry is still massively overstating their projections.
- daavidhauser
Opus 4.8 plus OpenClaw. I feel like the space is moving so fast that the result with this setup says very little about how close we are actually now.
- 0xDEAFBEAD
We need to be careful of wishful thinking. People are going to want to assume the existence of some sort of "deus ex machina" which is going to make everything fine. I prefer to turn the logic around. If there's any decently high chance that things could go off the rails, we should be shutting AI development down: https://pauseai.info/
- vessenes
Well, duh. If you could do this with Opus 4.8, we would know. When Astra’s successor is 2-3x better at math research, and the internal teams say “we believe we will get there,” I’m inclined to believe the insiders.
- dgellow
Link to the actual paper: https://arxiv.org/abs/2607.27191