Urgent.News

What's breaking now, across thousands of outlets.

AI

Multiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story.

A 438-billion-parameter reasoning model isn’t an obvious choice when speed is a priority. Multiverse Computing is betting that compression can The post Multiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story. appeared first on The New Stack .

Multiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story.

Multiverse Computing has unveiled Quasar 438B, its first large-scale 438-billion-parameter reasoning model tailored for coding and enterprise agents. The model claims a score of 43 on Artificial Analysis' Intelligence Index and 69.3 on Terminal-Bench v2.1, outperforming competitors like Mistral Medium 3.5 and NVIDIA Nemotron 3 Ultra.

Quasar boasts a 1-million-token context window, offered in English and Spanish, accessible via the CompactifAI API. Compression technology reduces model size by 80-95% with minimal accuracy loss, yet Multiverse hasn't disclosed the specifics of Quasar's compression or starting point. The model's performance shows trade-offs, with a Terminal-Bench v2.1 score of 69.3 placing it ahead of Mistral Medium 3.5 but behind the best frontier systems, while Claude Opus 5 leads with a score of 89.1.

Multiverse claims Quasar starts responding in 1.1 seconds and produces a 500-token response in 15.3 seconds, but agent tooling and repeated model calls may affect overall latency. Quasar is a proprietary model available only through Multiverse's API, making it hard to assess its real-world performance. As European AI companies develop their own models and infrastructure, Multiverse's compression approach presents an intriguing middle ground, but further benchmarks are needed to validate its speed claims.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

OpenAI’s new reasoning technique alarms AI safety experts

OpenAI’s new Astra model will use “recurrent depth,” a technique that allows the model to operate outside of the sequential thinking that characterizes most reasoning models.

  • OpenAI's Astra model introduces opaque recurrence, a novel reasoning technique.
  • AI safety experts express concern over diminished ability to monitor model's thought process.
  • Opaque recurrence involves looping queries, reducing legible traces and chain-of-thought records.

More from Wednesday 2 September →