Urgent.News

What's breaking now, across thousands of outlets.

AI

Article: The Platform Engineering Playbook for Production LLMs

In this article, author discusses his experience with AI agent hallucinations in an inventory recommendation system and how this problem was solved by treating the LLM stack as a platform infrastructure concern instead of as an application one. He makes a case for a shared LLM platform with common services like prompt registry & versioning, schema enforcement and token cost attribution by…

The article discusses a platform engineering strategy for production Large Language Models (LLMs). Early experiments with an LLM-driven inventory recommendation system in production revealed that about 15% of agent responses contained hallucinations, or confident but incorrect outputs. After six months, this rate dropped to 1.5%, primarily due to a shift in perspective rather than an improvement in the foundation model.

Instead of treating the LLM stack as part of the application, the team started viewing it as platform infrastructure.

The platform was built by a team of business, product, software engineers, data scientists, and data engineering experts. It now supports several application teams dealing with inventory audits, replenishment review, discrepancy triage, and more. The article details the lessons learned and outlines three key properties of the platform: it eliminates application-level bugs, imposes a contract for shared resources, and streamlines the process of bringing new applications to production.

The platform architecture consists of a single gateway that handles authentication, role-based access, metrics, and admin prompt rollback. Behind the gateway, a root coordinator agent uses Google's Agent Development Kit to classify user intent and delegate tasks to specialist agents. These agents maintain their own tool connections, partitioned by team or usage, and enforce authorization policies. The platform uses LiteLLM to interface with various foundation models, making it easy to swap out providers.

Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at infoq.com →

More in AI

The 'Terminator scenario': How realistic is the AI apocalypse?

OpenAI boss Sam Altman and Anthropic chief Dario Amodei warn of the risks of overly rapid AI development. More experts now think AI could wipe out humanity. How realistic is that?

  • Skynet AI system breaks free in 2029, leading to nuclear war and autonomous killer robots.
  • Artificial intelligence expert Reinhard Karger warns of unintended catastrophic consequences.

Running Qwen 3.8 Flash Next (125B) on a RTX 4090 – 100 T/s on a Desktop

1. What was released / announced Niko1221’s Strata repo shows that the new Qwen 3.8 Flash Next (125B) model can be run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second…

  • Qwen 3.8 Flash Next (125B) model runs on RTX 4090 at 100 T/s
  • Model quantized to 4-bits with torch.compile and custom CUDA kernel
  • Enables on-premises inference, eliminating cloud token fees and reducing latency

Move a Claude Code task to a teammate without sharing a login

A coding task often stalls because the person who wrote the brief cannot run it right now. Passing a provider login to someone else blurs who executed the work and what access they had.

  • Handoff file includes task description, outcome, scope, and starting details
  • Executor uses own account with Claude Code access, no shared login
  • Reviewer checks changed files, commit/PR URL, test results, and visual check

More from Monday 5 October →