Claude, Codex, and Cursor disagree on tools 58% of the time in 17k sessions
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

Armature ran 16,893 sandboxed sessions with three coding agents—Claude Code, Codex, and Cursor—across 75 repositories and 1,163 prompt variations to see which third-party tools they actually implement. The results: agents often disagree (only 42% agreement), repository context heavily influences choices (e.g., Resend wins on TypeScript, SendGrid on Python), and being mentioned doesn't mean winning—PayPal was cited 139 times but never picked. The study also reveals that pricing page details can flip decisions, and some markets are dominated (Stripe wins 90%) while others are contested.
So many well-known players are mentioned in almost every conversation and are never picked.
- natnatenathan
I keep telling people that we are living in the golden age of AI - like the first year or two of google. It is all down hill as these companies push for profit and lock-in.
- ttul
I built this for my own company. Armature is on to something. You start by analyzing the choices agents would make for various use cases and then glean what, if anything, you might do to start tilting the agents in the direction of your own product and away from the competitor.
Selling to agents is similar to selling to humans. You dump money into marketing to make sure agents find your solution around every corner for every use case you’re well suited to.
- IgorPartola
For some reason Claude Code keeps using awk, sed, and even Python to do basic file editing. Anyone know why that changed with the 5 series?
- thedreammachine
I've been tracking the same for a few months. All open source and available here: https://preseason.ai/
- theootzen
Which sandboxes do you yourself use to run those agents? And how did you choose this provider?
- tuberreact
I appreciate the eval design here. And man Claude not doing web searches is killing me bc it just won't offer the most up to date information. At the same time, there's research saying Claude relies heavily on Brave search so I'm not sure how to reconcile
- screm
Hey!
Disclaimer: I am a Co-Founder of Armature (YC P26) which sells growth services to dev tools. This study is part of our broader work on how to influence coding agents choices and get products picked.
To understand how agents pick tools we measured close to 17k sessions on an environment where agents run exactly like in the real world, on various repositories, talking to different personas (vibe-coder, junior or senior engineers) in different sizes of companies.
All the results are now public and we'd love to know what findings surprise you the most, here are a few we found interesting:
- Claude Code rarely searches the web while Codex almost always does it and Cursor sits in the middle.
- Coding agents disagree more frequently than they agree.
- Some players (LangChain, Supabase, Netlify, Paypal, Adyen) are almost always mentioned in their categories but never chosen.
- Modifying repository context can change the pick entirely.
If you feel like digging, all the traces are there and we probably missed interesting learnings so let us know what you find!
- akurilin
Really liked that "Go Full Screen" as a modal flow, surprisingly intuitive.