Urgent.News

the world's headlines, one feed

Editions

Tech

Coding agents can be evaluated. We just have to evaluate the work.

I recently argued with a software factory provider, whose position was that coding agents cannot be evaluated. Their reasoning was The post Coding agents can be evaluated. We just have to evaluate the work. appeared first on The New Stack .

Coding agents can be evaluated. We just have to evaluate the work.

Coding agents can indeed be evaluated, despite initial skepticism. The process requires examining the entire system rather than just the underlying model. A coding agent comprises a model, a harness, tools, repository context, instructions, permissions, an execution environment, and a feedback loop. Changes to any component can significantly alter the outcome. Public benchmarks are prone to misuse due to their focus on a specific model-agent-environment combination under particular conditions.

When evaluating coding agents, it's essential to grade their behavior rather than compare them to reference diffs. Establishing a known repository state, providing the agent with a task and available context, and then evaluating the resulting repository against executable contracts can help. This approach allows for evaluating multiple implementations while maintaining clear definitions of acceptable behavior.

Testing a coding agent should include several layers:

1. Outcome: Does the final repository satisfy the task?

2. Change quality: Is the implementation acceptable?

3. Trajectory: How did the agent arrive at the solution?

4. Human intervention: How much assistance did the agent require?

5. Economics: Is the result worth the financial cost?

6. Production impact: What happens after the merge?

While a single pass rate does not capture all six layers, it is necessary. Engineering teams already use a scorecard to assess delivery, and a coding agent evaluation should follow a similar approach. Non-determinism, the fact that a coding agent may solve a task one run and fail it on another, is not an insurmountable challenge.

Instead, it highlights the need for multiple runs with controlled starting conditions and reporting success distributions, variances in cost and completion time, and frequency of serious failure modes.

A useful evaluation should also consider how the agent interacts with users or product owners, identifies ambiguities, asks relevant questions, and incorporates answers without inventing requirements. The agent should not be penalized for asking necessary questions but penalized for confidently implementing incorrect assumptions. As newer research benchmarks evolve, they increasingly emphasize interactive project building and distinguishing dialogue capabilities from raw coding performance.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.

Read the original at thenewstack.io →

More in Tech

iPhone 18 Pro Will Reportedly Start With 256GB of Storage

iPhone 18 Pro Will Reportedly Start With 256GB of Storage

While the iPhone 17 Pro has double the base storage compared to the iPhone 16 Pro, there will apparently be no further increase this year. In a report this week estimating that the iPhone 18 Pro's bill of materials will be nearly 40% higher than the iPhone 17 Pro , Taiwanese research firm TrendForce said the iPhone 18 Pro will start with…