Agentic workloads break assumptions about software testing. Here’s how to cope
Most traditional enterprise systems were built around three assumptions: Jobs finish quickly, retrying one is free and the same input always produces the same output. Agentic workloads break all three of these expectations, which is why pilots that perform well turn into operational problems once they run unattended. The difficulty is rarely the model; it’s […] The post Agentic workloads break…
Traditional enterprise systems were built on three assumptions: jobs finish quickly, retrying is free, and the same input always produces the same output. However, agentic workloads challenge all three of these expectations, leading to operational issues when pilots perform well in a controlled setting but struggle once running unattended.
The root cause is often the surrounding infrastructure and management practices, not the model itself, that assume properties these workloads no longer possess. Agentic workloads differ from conventional software in several ways. First, they take minutes to complete, not milliseconds, which can exceed timeout thresholds and produce mysterious failures in stable systems.
Second, retrying now incurs costs, as retry logic was previously nearly free, leading to excessive retries against metered models consuming compute resources regardless of usability. Third, failures cannot be reproduced due to the lack of step-by-step trace of agent decisions, tool calls, and API runs. Experienced teams encounter challenges because evaluation environments conceal this information, making reviews speculative.
When working interactively, humans act as error handlers by reading results, noticing problems, and retrying manually. Automated agent workflows, however, operate headlessly, with the potential for unrecorded failures to break downstream systems. Teams must identify every task humans perform manually and specify automated checks or systems to take over those responsibilities.
Invisible costs arise because spending is no longer solely dependent on usage volume. Autonomous AI loops continuously retry failed tasks, regenerate responses, and hit APIs without human intervention or approval, causing monthly bills to exceed procurement negotiations. Failures reported as success can be particularly expensive, as probabilistic agents produce structurally valid but substantively incorrect outputs, which downstream automated systems accept without triggering alerts.
Monitoring conventional errors also misses these defects since nothing failed in the first place. To measure AI output quality accurately, predefined programmatic criteria, such as assertion checks, LLM-as-a-judge rules, or semantic benchmarks, must be used instead of relying on standard system uptime or error logs. Incident explanations become challenging when teams cannot determine which model version, inputs, and settings produced a specific output.
Essential questions to address include whether the platform can report on what produced an output, re-execute from saved inputs, distinguish failures from refusals or degraded results. The lack of a centralized AI workspace, where runs, inputs, outputs, quality results, and approvals are recorded together in real-time, leads to scattered workloads managed with traditional software practices.
To address this, teams should strive for an operational surface that allows active management, rerunning inputs, tracing logs, and controlling agent workflows in real-time. Continuous reporting and logging across every execution run are crucial for maintaining auditability and tracing anomalies. This approach also helps teams determine the most suitable platform for their AI workspace, as it should be able to report outputs, re-execute from saved inputs, and differentiate failures from refusals or degraded results.
Lastly, teams building agentic features often come from application development backgrounds, where synchronous request-and-response patterns prevail. These workloads exhibit long-running, partially failing, and expensive-to-re-execute characteristics. If no team members have experience running such systems, the skills gap will manifest as unexpected incidents rather than a call for skill upgrades.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.