Urgent.News

What's breaking now, across thousands of outlets.

Tech

Express.js Production Health Check Endpoints: Node.js Readiness, Liveness, and 5xx Monitoring

An Express production health check endpoint is useful only if it helps an operator make a safe decision during an incident. For a customer-support service running on Node.js, the real constraint is evidence retention: after a bad deployment, can the team distinguish liveness from readiness, find the relevant 5xx errors and logs, and explain which customer requests failed before and after…

When monitoring the health of an Express.js production service, it is crucial to separate liveness from readiness endpoints. The liveness check determines if the Node.js process can still serve events on its loop, while readiness indicates whether the instance can receive customer traffic. By exposing two distinct endpoints, operators can make informed decisions during incidents without inadvertently causing a restart loop due to transient dependency problems.

The liveness endpoint should return success if the process is responsive, regardless of the status of dependent services like databases. Conversely, the readiness endpoint should only return success when all required dependencies are functioning correctly and bounded checks have passed. This separation allows operators to remove unready instances without exacerbating the issue, while also enabling comparison of evidence from different releases before performing a rollback.

When implementing these health check endpoints, it is essential to include a release identifier and timestamp in structured logs, without exposing sensitive information such as secrets or customer message bodies. The endpoints should also reveal minimal internal topology to maintain security and simplicity. Dependency checks should be short-lived, with timeouts shorter than the probe interval to avoid thundering herd problems. Caching dependency results can help reduce the load on downstream services.

Logging should focus on a small set of events, including service_starting, service_ready, dependency_failed, service_draining, and service_stopped. These fields provide enough context to reconstruct the incident after the alert has been raised. Sensitive customer information should be pseudonymized, and only essential details such as operation, dependency, release, correlation ID, and failure class should be recorded.

This approach ensures that alerts are lossy summaries, while the retained evidence is rich enough to facilitate a thorough investigation.

When implementing an evidence collection system, it is recommended to use a monitoring tool like Infrai, which provides a self-describing API for accessing verified logs. However, it is important to note that Infrai is not a complete alerting system and should be used in conjunction with a separate notification system for managed escalation and regional probes.

The logging service does not provide features like per-user deletion, bulk export, or retention configuration, so these capabilities must be implemented separately if required by the organization's policies.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Wednesday 7 October →