What You Cannot See Will Break Your LLM App: A Practitioner Guide to Production Observability
Traditional application observability was built around a simple mental model: Your code runs, metrics come out and when something breaks, the logs tell you why. Large language models (LLMs) break that model in ways that are not obvious until you have shipped one and watched it misbehave in production. An LLM-powered application can be up, […]
LLM applications can function correctly while still producing catastrophic failures. Traditional application observability relies on code execution, metrics, and logs to identify issues, but Large Language Model (LLM) systems break this model in unforeseen ways. Even when an LLM-powered application appears to be running smoothly, serving requests and returning HTTP 200 responses, it may still be generating hallucinated content, silently truncating outputs, or degrading in quality due to updates in the underlying model checkpoint.
Over the past two years, the author built and operated a production LLM application processing tens of thousands of requests daily. The observability stack evolved significantly based on incidents that could not have been anticipated without practical experience. This article outlines the architecture and tooling that work in practice, focusing on real-world monitoring rather than theoretical stacks.
Conventional Application Performance Monitoring (APM) tools track latency, error rates, and throughput, which are necessary but insufficient for LLM systems. The failure modes that matter most are semantic, not structural. Unlike conventional APIs that return well-typed responses or throw exceptions, LLMs return strings that can be correct, plausible-sounding but incorrect, in the wrong format, violating content policies, or truncated due to context window limits. These issues are not captured by standard monitoring.
Four distinct problem classes in LLM production require dedicated observability signals, each with its own instrumentation approach. Problem Class 1 is Quality Drift, where output quality degrades over time due to provider model updates, poorly tested prompts, or downstream changes affecting the LLM's context. Quality drift is invisible without a baseline and automated evaluation of representative inputs against production traffic.
Problem Class 2 involves Prompt Failures, which can arise from unexpected inputs, edge-case formatting, adversarial inputs, or unusually long inputs that cause context truncation. These failures often appear as partial successes and require logging full prompt-response pairs with structured metadata and human review for failure pattern detection.
Problem Class 3, Cost Anomalies, arises from token-based LLM API costs that can spike unexpectedly. Bugs causing excessive context document inclusion or unnecessarily large prompts can multiply token consumption dramatically. Cost observability demands token-level tracking per request type, real-time monitoring, and alerting on anomalies before they lead to significant billing issues.
Problem Class 4 is Latency Degradation, where LLM API latency varies due to provider server load, prompt length, output length, and model family. Latency can degrade without changes on the developer's side and without provider notifications. Monitoring p50 latency is insufficient, as LLM latency distributions are fat-tailed, and user experience breaks down at p95 and p99.
The Minimal Instrumentation Stack consists of logging every LLM call as a structured event, tracking finish reasons explicitly, and focusing on cost anomalies and latency degradation. Structured log records provide data to compute per-workflow cost trends, latency distributions, and finish reason breakdowns, enabling early detection of truncation issues and other semantic failures.
Written by urgent.news from DevOps.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.