Urgent.News

What's breaking now, across thousands of outlets.

AI

Agent Graph Engineering, Part 0: I Needed This Before It Had a Name

TL;DR Building a chatbot is easy. Building an AI system you can trace, test, and change without fear is not. I spent years shipping LLM features as intent detection plus nested if/else , and no amount of observability tooling fixed it, because the problem was that the flow only existed in code, in my head, and in a stale diagram. The fix was to stop writing the flow as control flow and start…

Building an AI system that can be traced, tested, and changed without fear is not as simple as constructing a chatbot. I spent years shipping LLM features as intent detection plus nested if/else statements, and no amount of observability tooling resolved the issue, because the problem originated from the flow only residing in code, my mind, and a stale diagram.

The solution was to stop viewing the flow as control flow and begin declaring it as a graph. This post recounts the journey that led me to this realization, even before the concept had a formal name, and why I believe graphs are a sustainable solution rather than a fleeting trend. If you merely seek the stack and arguments to present to your team, proceed directly to Part 1.

Today, creating an AI assistant is straightforward. A few lines of code, an API key, and you possess a chatbot. Add a couple of tool definitions, and voila, you have an agent. Yet, this approach suffices for a hobby project. Scaling to production presents a more formidable challenge, and the divide is often underestimated. The conventional first response is observability.

Incorporate Langfuse, Langsmith, or Logfire, and you gain visibility into the model's output. This is beneficial. However, it does not address the core concerns that prevent an AI system from scaling effectively: How do you distribute work among distinct agents, ensuring that document handling and web research are separate tangles of prompt instructions?

How do you halt a tool call for human approval before executing irreversible actions such as sending an email, while permitting safe calls like retrieving a public profile to proceed without hindrance? How can one visualize the complete agent flow? Intent routing is embedded within conditionals, lacking a tangible artifact to present during design reviews.

How do you trace a user message from its initiation to the final answer? Observability platforms present a series of LLM spans. Reconstructing the flow involves reading code in reverse or scripting the transformation of traces into a comprehensible format. How does your frontend determine which messages the backend can transmit?

When the flow evolves, how does your system prevent a schema from becoming stale on one side? How do you stream intermediate progress in a predictable manner, instead of experiencing silence for forty seconds followed by a sudden deluge of information? How do you test these scenarios? Mocked responses for rapid feedback, real model executions to catch regressions when you switch models or modify prompts.

When an agent calls the incorrect tool, multiple tools, or none at all, what do you actually examine? If none of these challenges have affected you thus far, you likely possess a chatbot rather than a system. I am well aware of this scenario, as it was my own experience. Before delving into the story, note that the following content is not a fabrication or a hypothetical scenario.

It is a genuine product in the financial sector, involving real users, a real team, and mistakes attributed to myself. The intricacies presented herein are accurate because they transpired, and this detail is not something that can be fabricated at length. The accompanying code exemplifies the same narrative. It is condensed from a production system comprising five services, stemming from approximately five years of work on real-time infrastructure and agent systems, three years dedicated to one agent product and two years devoted to another.

What remains constant are not the domain, which has been intentionally omitted for confidentiality, but the failure modes. The components that appear excessively cautious are precisely the elements that failed in other contexts first. The passages that seem overcautious are the ones that encountered issues elsewhere earlier.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Building PlantLens AI for the Hacktoberfest "Touch Grass" Challenge

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built I built PlantLens AI , a privacy-focused, AI-powered botanical assistant.

  • PlantLens AI is an open-source botanical assistant.
  • Users upload plant photos for health assessment.
  • Project emphasizes privacy and outdoor activity.

What happens when an AI burns out

I once took someone off a job because they were burning out. They'd been at it for the best part of two days without a proper break, they'd stopped listening, and they were getting things confidently…

  • AI session ran for nearly two days without proper break
  • Session lost information through context compactions
  • Human intervened to guide AI back on correct path

Move Slow LLM Calls Off the Request Path with BullMQ in Node.js

If an LLM call can take 20 seconds or more, keep it out of your HTTP handler. Put the work on a BullMQ queue backed by Redis, return a job ID right away with a 202 Accepted , and let a separate worker…

  • Move slow LLM calls to a separate queue using BullMQ in Node.js
  • Return immediate response while dedicated worker handles asynchronous processing
  • Improves retry policy, concurrency control, and failure handling

More from Saturday 10 October →