Giving a coding agent more time barely helps
Real-SWE ran frontier models against licensed, private enterprise codebases (billing, tax, multi-service work) and one number jumped out at me: rollout duration barely moves resolution. 71.4% of rollouts that finished in under 10 minutes FAILED. 73.4% of rollouts that ran 10 minutes or longer also FAILED. Pass rate sits flat at 27-29% either way. Extending runtime from minutes to long rollouts…
A recent study revealed that extending the runtime of a coding agent barely improves its performance on complex enterprise tasks. Real-SWE tested frontier models against licensed, private codebases involving billing, tax, and multi-service work. The results showed that 71.4% of rollouts finishing in under 10 minutes failed, and 73.4% of rollouts running 10 minutes or longer also failed, with a pass rate remaining flat at 27-29%.
The only exception was Fable 5.1 on Claude Code, which achieved a higher resolution of 38.8%. Even this leader model failed on approximately six out of ten private enterprise tasks. The study emphasizes that vendor demos often overlook the structural limitations of coding agents and highlight the importance of considering the harness along with the model when evaluating performance.
The findings suggest that many benchmarking sources fail to account for the specific tasks that agents need to solve, leading to misleading comparisons and conclusions.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.