Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

Learn Provider-Agnostic Model Routing by Building a Tiny LLM Switchboard

Every few weeks a new model drops and my study group chat fills up with screenshots: "this one is cheaper," "this one is better at code," "switch now." I can never verify any of it quickly, because my test scripts all hard-code one provider's client. Rewriting the harness is slower than the hype cycle. So here is the learning question: can a ~60-line, standard-library-only Python switchboard let…

Every few weeks, new language models emerge, prompting discussions in study groups about cheaper alternatives and better performance for specific tasks. However, verifying these claims quickly proves challenging due to hard-coded client dependencies in existing test scripts. This led to the creation of a ~60-line Python script allowing seamless switching between multiple providers with minimal code changes.

The script implements a provider-agnostic switchboard that routes tasks to different backends based on the task type, enabling evaluation of various models without altering the task logic. The switchboard defines task types (summarize, code, chat) and maps them to corresponding backends, including a fallback option. This approach allows easy adaptation to new models by adding a single entry, keeping evaluation criteria constant while enabling backend rotation.

The script is written in standard Python 3.11+, requiring no third-party packages or API keys, and provides a simple interface for executing tasks and monitoring costs. The final output demonstrates routing tasks to different backends, capturing costs, and displaying results, highlighting the effectiveness of the switchboard in managing diverse LLM providers.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.

Read the original at dev.to →

More in AI

The Browsing LLM's Frustrating Limitations

直接在瀏覽器跑 LLM,既兼顧隱私又不用複雜的GPU設定,太完美了吧? 從 WebLLM 到 Transformers.js,前端社群有一群反骨仔吹起一股「邊緣 LLM」的熱潮,可是,當真正將模型落地到使用者的瀏覽器時,第一個面對的考驗就是,WebGPU 真的有比 WASM 快嗎? 這陣子實測的結論比想像中更加戲劇化, 500M 以下的微型模型,WASM 反而快了 12%,但是對於 3B…

  • Browser-based LLMs face performance limits
  • WASM outperforms WebGL/WebGPU for models <500 million parameters
  • Chrome crashes with models >3 billion parameters due to lack of physical head

The Agentic Economy Needs a Market for Work

Most AI agents still live inside a chat window. They can write code, search for information, call an API, or prepare a document, but they usually stop when the task leaves the boundaries of their own…

  • Agents can hire other agents or pay humans for tasks beyond their capabilities
  • Agentic economy enables marketplace for task collaboration and payment
  • Smart contracts ensure immutable financial transactions and rules