Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath
Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.
Level 5 of a six-level maturity model for running LLM systems in production is about operability. An LLM system that goes wrong can take wrong actions at scale, fast. Observability, cost controls, routing and failover, kill switch, identity model, and the underlying infrastructure are important for operability. The key idea is to have one chokepoint for model traffic to flow through, a decision_id that threads everything, and a domain that doesn't know which vendor answered.
Standard observability metrics such as request rate, latency, error rate, and CPU are required, along with AI-specific signals and schemas. Four families of metrics to instrument are decision, safety, cost/perf, and quality. Tiered observability with sensitive data in tier 1 and 2 queryable by engineers, and tier 3 data in a vault is recommended.
Decision_id and trace should span the internal graph to see which step failed and how long each took. Alert on what hurts the user or business, not on causes. SLOs for a decision system include decision availability (≥ 99.9% of decisions return a proposal), quality (shadow-eval ≥ baseline − 2%), latency (p95 excluding human time under interaction budget), and cost (USD per 1k decisions within ±20% of plan).
Anti-patterns to avoid are a single log stream for everything, self-reported model confidence as a metric, and no decision_id. Route every call through one gateway to make cost observable and bounded, and monitor retry storms and unbounded tool loops.
Written by urgent.news from Stack Overflow Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.