Warp builds self-improving agents on Claude

Warp, an AI-powered terminal and agentic development environment, has created self-improving agents using Claude's Agent Skills framework. The key insight: feedback to agents typically disappears when sessions end, so Warp built a loop where human feedback on agent output is captured and used by an 'improver' skill to update a base skill, which then improves future agent performance. This approach, now running across Warp's open-source repo, turns stateless feedback into compounding improvements, making agents more reliable and effective.
“File-based skills are a way of encoding knowledge for agents without putting that knowledge directly in the prompt, as something the agent can simply look up in the course of doing its job,” says Zach.
- bwfan123
> Agents need to handle recurring tasks reliably and effectively
This core problem remains unsolved. The solution presented in the article with Human In The Loop and some skill-magic such as "Write principles, not rules etc." is unsatisfactory because it offers no guarantees whatsoever. I find it difficult to harness agents into deterministic workflows which need to produce reliable outcomes.
- sandeepkd
I was bit confused in the beginning thinking its some product from Anthropic, looks like Warp is the startup, most likely getting rebate on using Claude and providing functionality to users, trying to get them addicted to the feature. And Anthropic is the one thats doing marketing for them cause eventually its their LLM which is being used. Not sure about the agents but this arrangement is definitely increasing the value of both companies in circular fashion.
- themgt
> What if it turns out the real AGI was the SKILL.md files we made along the way?
- JLO64
I already knew what Warp is (I switched to Ghostty and haven't looked back), but I find it odd that the "The quick pitch" card at the top of this article makes no mention of what the company actually does. Who cares more about their founder/growth/age over that?
- ziyadb
The problem of handling recurring tasks predictably comes down to the probabilistic nature of LLMs, which are based on next-token prediction.
I founded a company called Aide where our goal was to help support teams reliably deploy customer-facing agents without worrying about poor interactions. The first problem we needed to solve was making them deterministic and eliminate the variance that comes naturally with base models.
Getting them to always adhere to brand policy, eliminate hallucination, and stay grounded in data was a fun challenge. Proud to say that we’ve devised a solution that runs well and it’s worked out quite nicely in compliance-heavy and regulated environments.