Urgent.News

What's breaking now, across thousands of outlets.

AI

Building scalable, agent-friendly APIs for AI applications

Nowadays, most of us interact with frontier AI models in a structured way, even if we do not always think The post Building scalable, agent-friendly APIs for AI applications appeared first on The New Stack .

Building scalable, agent-friendly APIs for AI applications

In today's digital landscape, most individuals engage with cutting-edge AI models via structured interactions, even if they may not consciously recognize this trend. The majority of systems currently employ protocols, typically developed by the model providers themselves. OpenAI and Anthropic, for instance, both offer APIs to facilitate communication with their inference systems.

Practically speaking, one approach involves utilizing an OpenAI-compatible API capable of accommodating multi-turn agents. By pointing the official OpenAI Software Development Kit (SDK) at a compatible inference endpoint, users can engage in conversations with the agent while maintaining consistent interface usage throughout the dialogue.

These systems are ingeniously designed for scalability. As traffic volume increases or subagents emerge, these systems must be scaled accordingly, ensuring that engineers can make necessary adjustments without users being aware of these modifications. For instance, scaling a single-replica system to three replicas behind a load balancer can accommodate additional traffic demands while preserving conversation context and prior turns.

The principles of microservices, once overlooked around five years ago, prove to be precisely what agent traffic requires. Before developing these systems, many questions were already addressed by the microservices literature from the 2010s. By reconstructing an agent-friendly API based on these principles, employing Oracle AI Database Free as a foundation, and conducting measurements, we can evaluate the necessary changes.

A crucial distinction between agent traffic and web traffic lies in the load shape, a disparity that stateful designs struggle to absorb effectively. Four key properties set agent clients apart from browser clients: conversations are lengthy and stateful, but requests are not. A 40-turn agent loop translates to 40 independent HTTP requests, with no protocol ties linking them to a single server.

Each tool call is stateless at the HTTP layer, meaning the endpoint receives individual requests and responses over time. Consequently, the application must determine which conversation a request belongs to and where its state resides. Tool calls exhibit unpredictable fan-out. A single chat request might trigger zero, one, or multiple downstream calls, with some calls potentially involving a vector index that takes up to 400 milliseconds to process.

Since model outputs and tool-section paths can vary, even identical inputs may result in different sequences and quantities of tool calls. Agents aggressively retry. AI agents are typically enveloped within an agent harness, which includes its control loop, error handling, and retry policies. In the event of timeouts or errors, the system may receive rapid sequences of retries associated with the same workload.

While this retry mechanism enhances resilience, it can also exacerbate pressure on dependencies that are already under strain. Traffic patterns are bursty and machine-paced, with little to no think time between turns. Unlike human interactions, AI agents generate a continuous stream of work, constrained primarily by hardware capabilities and concurrency limits.

The first issue typically encountered is the need for deterministic request routing to maintain conversation context. In our Proof of Concept (POC) benchmark, we drove 200 conversations consisting of four turns each, routing every turn to the next available replica in rotation—simulating a distribution that emerges when stickiness parameters are not appropriately configured.

Comparing deployments, we observed the following: a monolithic setup running on a single machine processed 800 turns with no context loss and a 0% loss rate. Scaling the system to three machines resulted in 800 turns processed, but 600 instances exhibited a 75% loss rate. In contrast, a decomposed approach, running on a single machine, maintained 800 turns processed with no context loss or loss rate.

Scaling the decomposed system to three machines also yielded 800 turns processed with zero context loss and a 0% loss rate. The fundamental flaw lies in the monolithic deployment with one machine versus three. Although the design is not inherently flawed; it becomes conditionally correct when every turn reaches the same machine, which is relinquished when scaling is implemented.

The failure remains silent, allowing the model to respond, albeit potentially without the necessary context since the request may have landed on a different machine lacking the prior chat history. This silent failure occurs because the system responds with HTTP 200 status codes, masking the underlying issue for users. Recognizing these challenges prompts a reevaluation of API design principles for AI applications.

The essential insight is that the solution encompasses microservices characteristics established by Lewis and Fowler in 2014 and Michael Nygard's bulkhead pattern from his book "Release It!" adapted to the unique request shape of AI applications. Three guiding principles emerge as crucial: Statelessness: By moving conversation state out of the process and into a shared store, any replica can serve any turn of any conversation.

This single change enables the previously monolithic endpoint to scale seamlessly across replicas without relying on process-local conversation state. The chat history no longer needs to be transmitted back and forth with every request, as the agent can retrieve the required context from the shared state.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

RelayZero: Offline Edge-AI Semantic Compression for Disaster Mesh Networks

The Problem: When the Grid Goes Dark When catastrophic natural disasters strike, the first thing to collapse is centralized telecommunications.

  • RelayZero compresses emergency messages into 25-byte telemetry packets.
  • Edge-AI model extracts priority, location, hazard type, casualty count, and resources.
  • Zero-Network Alerts generate police siren for CRITICAL packets.

How Technology Empowers—and Imperils—Dictators

This essay was written with Seva Gunitsky, and originally appeared in Foreign Affairs . Two weeks after Moscow’s full-scale invasion of Ukraine in March 2022, the Russian TV Channel One editor Marina…

  • AI can empower dictators by avoiding human costs of protests and control.
  • AI systems can be used to maintain autocratic control without human intermediaries.
  • China leads in AI development for autocratic control and oversight.

More from Thursday 8 October →