Armature's 16,893-run study finds Stripe, Neon and AWS dominate some coding-agent tool choices
- Armature says it ran 16,893 sessions across 1,163 prompt variations, 75 repositories, three coding agents, and personas ranging from vibe coders to enterprise senior engineers; agents implemented the selected capability rather than only recommending a product.
- Five of 18 sectors had a product chosen in at least half of runs: Stripe took about 88% of Payments, Neon about 66% of Databases, and AWS about 62% of Cloud.
- Nine sectors remained contested, with Vercel leading Deploy at about 41% and no AI gateway product exceeding 22%; Portkey and Cloudflare AI Gateway each recorded 30 wins in that category.
- Agents built capabilities in-house in about 13% of all runs, and did so more often than selecting any product in three sectors; Performance in CI was built in-house in about 52% of runs.
- The authors disclose that Armature sells growth services to developer-tool vendors and describe the study as part of work on influencing coding-agent choices; they publish prompts, traces, and code diffs alongside the results.
Hacker News 의견들
I co-founded Armature, and we sell growth services to dev-tool companies, so this research is also about how vendors can influence coding-agent choices. We published the prompts, traces, and code diffs because repository context, personas, and agents change the result.
I wanted to inspect the data, but the site forced a fullscreen onboarding flow with a nearly hidden skip button. On Safari for iOS, the fullscreen modal was misaligned and cut off, so I closed the tab.
I see Claude Code rarely researching unless I explicitly ask, while Codex starts fetching information on its own. I want to know why their defaults are so different.
If the business is helping companies manipulate agents into choosing products that are not the best fit, it repeats the growth-hacking practices that already made search worse.
I expect tools favored by AI agents to become mainstream, since they already tend to recommend products people use. That could concentrate adoption around a small set of vendors.
This looks like paid search advertising for agents, except agents may act without the user noticing. The model labs need to control that.
Redshift being absent does not make the database results suspect. This benchmark covers production databases, not data warehouses, where Redshift would be relevant.
I do not want another cycle of ads, SEO manipulation, and low-quality content ruining an information channel that people use because search has become unreliable.
I see the commercial opportunity: test what agents choose in a use case, then change a product's marketing so agents encounter it more often. Selling to agents will resemble selling to people.
Some choices look wild, including Azure for databases and in-house bot protection or search. The authors say Azure mainly appeared in enterprise prompts, while Claude Code rebuilt things more often unless prompted to evaluate external options.
The 17,000-run scale is more interesting than any individual winner. At least someone is measuring agent tool selection instead of relying on anecdotes.
I have tracked similar behavior for months and published an open-source project at preseason.ai.