Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache.

OpenAI released GPT-6 Sol and Luna on Tuesday, essentially more affordable versions of GPT-6 Astra that come closer to Astra The post OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache. appeared first on The New Stack .

OpenAI cut GPT-6 token prices in half. The bigger lever may be the cache.

On Tuesday, OpenAI unveiled GPT-6 Sol and Luna, more budget-friendly iterations of GPT-6 Astra, which closely mirror Astra's alignment while not quite matching the flagship model. One of the standout features of these new models is the substantial reduction in token prices, making them considerably more affordable to utilize. OpenAI explained that advancements in caching and inference enable them to render these models at reduced costs, with API prices for Sol and Luna falling 50% compared to their GPT-5.6 counterparts, and Luna output tokens down 58%.

The improvements primarily revolve around higher cache hit rates by default, preserving earlier context even when reasoning effort and tool availability change, and the introduction of new tools to monitor and diagnose caching performance. While prompt caching is not entirely new to GPT-6, the enhancements in Sol and Luna aim to retain more previously processed context as the agent advances on a task.

By default, GPT-6 models now boast a higher cache hit rate, enabling agents to reuse more context, respond more quickly, and enjoy discounts of up to 90% on cached input-token reads. This enhances cost reduction, as the model no longer needs to process identical context from the beginning for each call, thereby lowering latency and token costs.

Furthermore, GPT-6 offers developers more flexibility to optimize caching performance, allowing them to adjust reasoning effort and tool availability without disrupting cached context. This adaptability benefits both speed and cost.

OpenAI has also introduced a Prompt Caching Dashboard, providing developers with visibility into caching performance. The dashboard enables users to monitor how much context is being reused and how this amount fluctuates over time. In addition, the diagnostics tool flags missed caching opportunities to help developers identify areas for more efficient caching. By making cache performance more visible, OpenAI aims to encourage developers to actively measure and optimize cache reuse.

These caching enhancements and lower token prices are already proving beneficial. According to GitHub, these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models, spanning the past several months. While token prices play a significant role in making agents more affordable, OpenAI emphasizes that a combination of reduced costs for fresh processing and minimized reprocessing of the same context is key to managing agent expenses effectively.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

More from Wednesday 23 September →