Urgent.News

What's breaking now, across thousands of outlets.

Tech

Coping is the Problem

Retries won't save you. Some failures can't be prevented. Robust systems need to cope.

Coping is the Problem

Throughout my career, I've found the principle from John Gall's The Systems Bible to be incredibly insightful: the issue isn't the problem itself, but rather how we cope with it. To illustrate this, let's examine HTTP requests. When a request fails, our inclination is typically to add a retry mechanism. However, this doesn't resolve the issue; the second or third request may also fail. While retries may lessen the occurrence of failures, they cannot guarantee an HTTP request's success every time.

To navigate this, we need to broaden our perspective on the system. Using our HTTP request example, if server A fails to send push notifications to server B, there are numerous reasons behind this. To cope, server A could implement an alternative endpoint, allowing server B to retrieve notifications instead. This approach shifts the focus from ensuring uninterrupted delivery to optimizing how much effort is spent on mitigating failure.

In system design, it's often more effective to consider how the system will cope with a problem before devising ways to mitigate it. Mitigation, as defined here, refers to reducing the frequency of issues without entirely eliminating them. Examples include increasing timeouts for slower processing or expanding cache sizes due to slow responses.

This principle proved invaluable in my recent project building a personal agent to replace a spreadsheet. Natural language processing can be unpredictable, leading to frequent errors. Without coping mechanisms, I would have continuously fine-tuned prompts, attempting to achieve 100% accuracy. However, the problem at hand involved tracking expenses between households communicated via Slack messages.

The agent's task was to identify whether expenses were shared or individual, splitting shared expenses accordingly. Initially, I aimed to create a sophisticated system prompt and tool descriptions to address classification errors. Yet, despite investing significant time in refining prompts, I eventually opted for the "find_and_edit_expense" tool to cope with the inherent non-determinism.

While adding new tools isn't inherently problematic, it's crucial to ensure that complexity is justified. When receiving a Slack notification about an expense, the agent recorded it and notified the Slack channel. If errors were identified, users could send a follow-up message to correct them. Introducing the "find_and_edit_expense" tool significantly reduced the stress on the initial "add_expense" tool.

The added tool allowed for better error handling and corrected issues without overfitting the LLM to specific evaluations. By making the "add_expense" tool sufficiently reliable, I could focus on optimizing its performance rather than striving for perfection.

In high-stakes environments, where incorrect tool calls could be costly, alternative coping strategies might be necessary. These could include delaying tool execution to allow users to cancel, or requiring explicit approval through deterministic code before executing potentially risky actions. Logging failures or simply moving on can also serve as effective fallback options.

Ultimately, we must weigh the trade-offs, recognizing that certain problems cannot be entirely prevented. The key is to design the fallback mechanism first, reducing the pressure on mitigation efforts and transforming them into optimization problems.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in Tech

More from Wednesday 30 September →