Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
In a large-scale experiment, coding agents Claude Code, Codex, and Cursor were observed implementing solutions for various coding needs from almost 17,000 runs. The main question was how these agents would choose and implement tools for specific requirements. The experiment involved 1,163 prompt variations across different personas and 75 repositories.
The coding agents analyzed codebases and recommended solutions, which were then implemented by the agents. The results were gathered, analyzed, and shared to understand the thought process of coding agents in selecting and implementing tools.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.