Urgent.News

What's breaking now, across thousands of outlets.

Tech

O harness é o novo código: por que o desempenho do agente depende mais do ambiente de execução do que do modelo

Trocar o modelo e manter o mesmo ambiente de execução raramente produz o salto que o marketing promete. O que muda o resultado, em muitos casos, é o conjunto que envolve o modelo: ferramentas disponíveis, regras de uso, memória de sessão, planejamento, subagentes, validações e o modo como o agente lê e escreve no repositório. Esse conjunto tem nome na literatura recente: harness. Dois trabalhos…

The performance of a software agent depends more on its execution environment than the model itself, according to recent studies. The right "harness" – a combination of tools, usage rules, session memory, planning, subagents, validations, and the way the agent reads and writes to the repository – is crucial for consistent results. Two studies published in September 2026 highlight the importance of harness in software engineering agents.

The first study, "Beyond the Model: Demystifying Harness Effects in Software Engineering Agents," compares harnesses, models, and benchmarks. The authors conclude that harness effectiveness depends on both the model's capability and the task type. Structured tool usage and specific subagents tailored to the task tend to yield more stable gains, while aggressive context compression and generic subagents in scenarios involving entire repositories can worsen outcomes.

The second study, "Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents," analyzes the source code of eleven agent systems. It describes the harness as the operational anatomy of these systems, emphasizing that it is not just a "nice prompt." Instead, it is the layer that transforms a language model into something capable of consistently working on a codebase, complete with permissions, tool loops, stopping criteria, and verification mechanisms.

This shifts the focus from "which is the best model" to "which environment does this model operate in?"

Teams that only change the language model and reuse the same weak harness reproduce the same pattern of failure, albeit with a different cost. On the other hand, teams that invest in appropriate tools for the task, in a context where the agent can actually use them, and in rules that limit what can be changed without review, achieve more predictable behavior, even with less capable models.

In practical software engineering, the harness becomes the object of design. What tools the agent can call, what enters the context and what remains outside, when a specific subagent is triggered and when it is not, and which validations run before the diff is offered to a human – these are all part of software engineering assisted by AI.

The model matters, but the environment in which it runs often matters more for the final result in the repository. For a more in-depth exploration of this topic, complete with examples and comprehensive architecture analysis, refer to the book series "Software Engineering Assisted by AI," available on Amazon.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Wednesday 30 September →