Your AI Agent Returned HTTP 200. Why Did the Workflow Still Fail?
A successful HTTP response is not a successful agent run. A recent practitioner report from a 58-day deployment of 78 agents recorded 6,768 failed outputs. The failures were not transport errors: every one returned HTTP 200, had plausible length, and looked fluent. The most expensive failures were boring shape mismatches: missing required fields, wrong language, forbidden phrases, or an answer…
An HTTP 200 response does not guarantee a successful agent run. A recent study involving 78 agents deployed for 58 days found 6,768 failed outputs, all of which returned HTTP 200, had reasonable length, and appeared coherent. The most problematic failures stemmed from shape mismatches, missing required fields, language errors, forbidden phrases, or answers for incorrect stages.
This serves as a reminder that model responses should be treated as untrusted data, requiring validation at the boundary before being consumed by subsequent stages. This article demonstrates this principle through a small, reproducible failure lab.
Consider a review stage where the downstream parser expects a verdict line: action: approve. While a human can approve the model's thoughtful review, a parser may not. Despite green transport and model call layers, the workflow can still fail. The same class of failure arises when a JSON field has the wrong type, response is in the wrong language, a tool returns an error string inside a successful content envelope, a stage emits output that the next stage doesn't read, or a reviewer from the same model family highlights a shared blind spot.
These issues are not resolved by simply using larger models. Instead, they highlight the need to make boundaries observable and enforceable.
To achieve this, implement a contract gate with deterministic checks that do not rely on another LLM to judge another LLM. The validate_review function checks for issues such as short text, missing required verdict, forbidden phrases, and missing expected language. If any errors are found, they are recorded and the downstream dispatch is stopped. If no errors are found, the output is published to the next stage.
The key takeaway is not the specific Japanese checks, but the contract your system actually needs, such as required headings, schema types, repository paths, test names, citation fields, or a bounded action list. The gate should return structured evidence, not just true or false, detailing why the contract was not satisfied. This allows for easier debugging and repair of failures.
Recording detailed information about rejections is crucial. This should include artifact_id, contract_version, observed_checks, failure_reasons, raw_output_hash, downstream_read_at, and reviewer_family. Storing only a boolean value like contract_satisfied = false is insufficient, as it provides no information for debugging.
It's also important to detect if outputs are consumed by any downstream stages. Assign stable IDs to every produced artifact and require the consumer to record the input artifact ID. If a stage claims success without a consumed input ID, it should be rejected. Compare produced and consumed counts over a time window and alert on a growing gap, even if all process heartbeats appear green. This helps identify wiring bugs that output-quality checks may miss.
To ensure reliable agent workflows, inject each case and verify the expected evidence. This includes removing the required verdict line, returning valid-length text in the wrong language, putting an error string in a successful tool envelope, dropping the artifact ID between stages, changing the contract version mid-run, requiring an independent check or human review, and crashing after provider acceptance but before ledger write.
It's important to separate execution evidence and outbound-delivery evidence. While a hosting service can ensure runtime availability, it doesn't define your output contract or make a green HTTP response meaningful. Always own validation, credential scope, prompt-injection defenses, and reconciliation.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.