Urgent.News

What's breaking now, across thousands of outlets.

AI

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Writer , the enterprise AI agent platform used by Fortune 500 companies including Accenture, Uber, and Vanguard, released its new flagship model Palmyra X6 today, alongside a rebuilt agent orchestration "harness" and new governance tools designed to give IT leaders control over runaway token spending. The headline numbers are striking: Writer says its agent product now operates at an average 52%…

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Writer, the enterprise AI agent platform utilized by Fortune 500 companies such as Accenture, Uber, and Vanguard, introduced its latest flagship model, Palmyra X6, alongside a revamped agent orchestration system and new governance tools aimed at assisting IT leaders in managing excessive token expenses. The platform's headline figures reveal that the agent product now operates at an average 52% lower cost, with a 48% speed improvement and a 10% quality enhancement when utilized with the Palmyra X6 model.

However, the more significant aspect of this development might be the method by which Writer achieved these cost reductions and the implications for the future of enterprise AI.

It is essential to note that Palmyra X6 is not an entirely new model but rather a post-trained version of GLM-5.2, an open-weight mixture-of-experts model developed by Beijing-based Z.ai, formerly Zhipu AI. Writer openly discloses this information in its technical report, placing the company at the heart of a contentious debate within the industry about whether American enterprises should leverage Chinese open-source foundations for their AI models.

Despite not being directly connected to the original developers of GLM-5.2, Palmyra X6 is fully run on Writer's US infrastructure, as confirmed by Matan-Paul Shetrit, Writer's director of product management, in an exclusive interview with VentureBeat.

The timing of Writer's announcement is particularly noteworthy, as the enterprise AI market is increasingly grappling with the economic challenges associated with agentic AI. Unlike traditional chatbots, which typically generate a single response per user query, AI agents can initiate multiple rounds of planning, retrieval, tool calls, validation, and retries for a single request, with each loop consuming metered tokens.

This leads to rapidly escalating costs, as Goldman Sachs projects that token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month. The issue is exacerbated by the fact that falling per-token prices do not guarantee reduced total costs; a task requiring 20 times more tokens while unit prices drop 75% can still result in total charges increasing fivefold.

Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at venturebeat.com →

More in AI

DeepSeek Harness developer preview

https://github.com/deepseek-ai/deepseek-harness https://deepseek-harness.github.io/deepseek-harness/en/guide... Comments URL: https://news.ycombinator.com/item?id=49285244 Points: 426 # Comments: 197

More from Thursday 13 August →