Urgent.News

What's breaking now, across thousands of outlets.

Editions

AI

Nvidia finds that simple linear math can replace costly AI model handoffs

When an agentic AI system hands a task from a small model to a larger one — or back down again — it pays a steep tax: the receiving model has to recompute the entire conversation from scratch, driving up compute costs and latency. This is a major bottleneck for enterprises building long-horizon, multi-LLM workflows. To solve this challenge, researchers at Nvidia have introduced a cross-model KV…

Nvidia finds that simple linear math can replace costly AI model handoffs

Nvidia researchers have developed a technique to transfer the Key-Value (KV) cache from a smaller AI model to a larger one without re-running the entire conversation. This eliminates the need for costly deep learning model training and reduces compute costs and latency in multi-LLM workflows. The linear mapping process is 2.7 to 25 times faster than re-computing the conversation while maintaining up to 98% of the target model's accuracy.

The technique works by transforming the KV cache of one model into the expected format of another, enabling seamless handoffs between models in agentic AI systems.

Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at venturebeat.com →

More in AI

AI Capex Has Moved Into Credit's Jurisdiction

There are two honest ways to talk about the AI infrastructure boom. One starts with demand. Model usage is rising, enterprise budgets are moving from pilots to deployment, and cheaper inference can…

Sub‑50 ms On‑Device TTS: Instant Voice for Games & Streams

Ultra‑Low‑Latency TTS: How to Generate Voice in < 50 ms on‑device Introduction Imagine a game NPC that answers your question the instant you speak it, or a live‑streamer who adds a multilingual…

  • Sub-50 ms latency achieved for on-device TTS in games and streams.
  • Models like VITS-Lite, FastSpeech-2+, and Glow-TTS-Tiny enable sub-50 ms performance.

AI Capex Is Turning Into an Infrastructure Bill

The AI bubble argument got louder this week because it stopped being only about Nvidia's chart. Axios framed the U.S. as being in a capital squeeze, with federal debt, entitlement spending, defense…

AI Is Learning to Write Genetic Code

This sort of research is both exciting and terrifying: The two models in question were told to generate complete genomes for a viable bacteriophage—a type of virus able to infect and replicate itself…

More from Friday 21 August →