Urgent.News

What's breaking now, across thousands of outlets.

Tech

The Trace Ends Where The Message Starts

The synchronous path is one tree; everything after the publish is another one - and the audit found out why: the headers were empty. ๐Ÿ‘‹ Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about the three things a fleet of services has to do identically or each service reinvents them: who arrived, what theyโ€ฆ

The narrative of the trace begins at the point where the message is delivered, and it concludes where the call context is maintained. In a system where a PHP monolith is being refactored into Go services, the challenge lies in ensuring that each service consistently performs the same tasks, regardless of their individual operations.

Every service must know who arrived, what actions they can perform, and the journey of their request. This information is crucial for maintaining the trace and understanding the system's behavior.

At the core of this system, a PHP monolith is being broken down into Go services, with the Go API gateway serving as the primary entry point. The Go services communicate with each other through gRPC, and each service runs on the same shared platform library. Some interactions are synchronous, while others are published as domain events and handled by background processes.

Before the service even starts running, it collects a list of information, totaling 67 records. This includes 59 records from the platform, 6 from the service itself, and one dynamic entry for the dynamic-metric factory. In addition, the service declares 13 domain metrics of its own, such as queue depth, age of the oldest queued row, relay cycle duration, and more. All of this information is collected using a pull-based system, with an agent scraping the system port and remote-writing the data into a store.

Traces are exported over OTLP, and the observability in this system is provided by the runtime on which the service is built, rather than being coded per service. This ensures that the observability remains consistent across all services, without requiring each service to implement it differently. The only requirement left for the service is to maintain the call context, which is passed as the first parameter to handlers, repositories, managers, background handlers, and the publish call itself.

This requirement is enforced through cancellation tests, which prevent dropped contexts from causing issues that would otherwise go unnoticed.

The case in point is a situation where a rule - the requirement to maintain the call context - was initially treated as an agreement. However, the test code accumulated 421 calls building an empty context across 184 test files, leading to a significant issue. Tests built on an empty context did not cancel with themselves, causing them to hang until the package timeout.

Furthermore, these tests did not assert anything about the context, meaning that if production code dropped the incoming context and substituted its own, the entire suite would remain green. This situation demonstrates the importance of enforcing the rule regarding the call context to prevent unexpected behavior and ensure the trace remains intact.

Written by urgent.news from Dev.to's reporting โ€” not their text. Machine-written โ€” may contain errors; check the original before relying on it.

Read the original at dev.to โ†’

More in Tech

More from Wednesday 7 October โ†’