Urgent.News

What's breaking now, across thousands of outlets.

AI

How I Built a Unified API Gateway for 200+ AI Models (Architecture Deep Dive)

Every developer who has worked with more than one AI provider knows the pain: OpenAI has its own API key Anthropic has its own billing dashboard Google has its own SDK DeepSeek has its own rate limits If you want to use 4 different models, you have to manage 4 accounts, 4 invoices, and 4 sets of documentation. This isn't just annoying — it's a bottleneck that stops developers from experimenting…

Developers who work with multiple AI providers understand the frustration of dealing with different APIs, billing systems, and documentation. This can significantly hinder experimentation with new models. To address this, I created SarangAI, a unified gateway that routes requests to over 200 AI models via a single, OpenAI-compatible endpoint. Here, I'll detail the architecture, design choices, and technical hurdles of SarangAI.

The architecture of SarangAI is centered around four key components: the API Gateway, Router, Provider Adapters, and Billing & Rate Limiter. The gateway receives client requests and passes them to the router, which decides which model to use. The provider adapters then translate the OpenAI-compatible format into each provider's specific format, while the billing and rate limiter manages prepaid balances and per-user limits.

One of the main design decisions was to standardize on the OpenAI-compatible format. This choice simplifies integration, as developers can easily switch from existing SDKs by merely changing the base URL and API key. The OpenAI format's widespread adoption ensures a vast ecosystem of libraries and tools, as well as a familiar structure for developers.

Technical challenges arose in normalizing responses, handling streaming, and implementing rate limiting. To tackle these, I used an adapter pattern for each provider, normalizing all streaming responses to the same Server-Sent Events (SSE) format used by OpenAI, and implementing a prepaid balance system using Redis and a database for real-time tracking and audit trails.

Another key feature is instant model switching, which requires a stateless router capable of handling multiple configurations without downtime. The SarangAI CLI tool, accessible via npm install -g sarangai-cli, provides an interactive command-line interface for switching models, checking token usage, and reviewing balances. This tool is especially valuable for developers who prefer working in the terminal.

Through my experience building SarangAI, I learned that standards are crucial, adapter patterns simplify integration, prepaid subscriptions appeal to developers, and comprehensive documentation is essential for adoption. I encourage those who regularly work with multiple AI models to try SarangAI and share their setups in the comments for a discussion on handling multiple AI providers.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Touchgrass.local — An AI Coach That Gets You Outside

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built Touchgrass.local is a small AI-powered browser app that encourages people to step away from their…

  • Touchgrass.local is an AI-powered browser app
  • Encourages outdoor missions with friendly nudges
  • Tracks completed missions and streaks privately

EmbeddingGemma 2

My comment on EmbeddingGemma 2 — Hacker News. I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license. For embedding models in particular, I don't think it makes sense to use a…

More from Tuesday 6 October →