Observability Explained: Logs, Metrics, Traces, and What Monitoring Misses
๐ก 1. The Three Observability Pillars (and the "Fourth") 1๏ธโฃ Logs โ What happened? 2๏ธโฃ Metrics โ How much, how often, and how fast? 3๏ธโฃ Traces โ Where did the request go? 4๏ธโฃ Events (the unofficial pillar) โ What changed? What about profiles? 2. Monitoring, Security, and Observability ๐ฅ๏ธ Infrastructure monitoring ๐ Application monitoring ๐ Security monitoring 3. Following a Request Through aโฆ
Observability is a powerful method for understanding the inner workings of modern systems. While traditional monitoring can confirm if a server is running and if the application is returning a "200 OK" status, it cannot always explain why the system behaves in unexpected ways. Observability fills this gap by providing insights into what is happening inside a system based on the telemetry it generates.
The three traditional pillars of observability are logs, metrics, and traces. Events, while not always emphasized, are also becoming increasingly important in observability. Let's take a closer look at each of these pillars.
Logs provide valuable information about what happened in a system. They are generated by applications, operating systems, network devices, security tools, and various services. Examples include nginx access logs, PHP error logs, MariaDB logs, and Kubernetes container logs. Logs can give us detailed information about specific events, such as authentication issues or slow queries.
Metrics represent numerical measurements that change over time. Common examples include CPU and memory usage, HTTP request rates, and error rates. Metrics are excellent for creating dashboards, setting up alerts, planning capacity, and analyzing trends. They help us understand how much, how often, and how fast things are happening. However, metrics alone often don't tell us the "why" behind the changes.
Traces follow a request as it moves through various components of a system. Imagine a user request to "/users" that goes through Nginx, PHP, MariaDB, and other services. A distributed trace would show the time spent at each stage, helping us identify where the delays are occurring. Tracing is especially valuable in distributed systems where a single request might involve multiple components.
Events represent meaningful changes that occur at a specific moment. Examples include application deployments, database latency increases, Kubernetes pod restarts, and security modifications. These events can provide crucial context during incident investigation, helping us determine what changed when application performance issues arise.
In summary, observability is essential for understanding the behavior of modern systems. By combining logs, metrics, traces, and events, we can gain a deeper understanding of what is happening inside a system, even when traditional monitoring approaches fall short. Monitoring remains important, but observability provides the context needed to investigate and resolve complex issues.
Written by urgent.news from Dev.to's reporting โ not their text. Machine-written โ may contain errors; check the original before relying on it.