Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

OpenAI on Tuesday released GPT-6 Sol and Luna, which will complement the flagship GPT-6 Astra model in OpenAI’s lineup. As The post OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half appeared first on The New Stack .

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

On Tuesday, OpenAI unveiled GPT-6 Sol and Luna, two new models set to join the company's lineup alongside the flagship GPT-6 Astra model. However, there is currently no GPT-6 Terra model. The most significant announcement is the halving of token prices for these models compared to their predecessors. GPT-6 Sol will cost $2/$10 per million input/output tokens (down from $4/$20 for GPT-5.6 Sol), while GPT-6 Luna will cost $0.10/$0.50 (down from $0.20/$1.20).

OpenAI attributes these price reductions to improvements in caching and inference, allowing them to serve the models at a lower cost and pass those savings directly to users and customers. Benchmarks show clear improvements in GPT-6 models over the GPT-5.6 predecessors, but they are not overwhelmingly dramatic. For instance, on Zapier's AutomationBench, which tests business workflow tasks, GPT-6 Luna improves by 5.4 percentage points over the previous version, while GPT-6 Sol essentially matches Anthropic's Fable model at only 20% of the cost.

Anthropic has recently released Opus 5.5, which has further shifted the comparison landscape, as its per-token pricing has been reduced to $4/$20 from $5/$25. However, Anthropic claims that Opus 5.5 also uses fewer tokens per task, resulting in 40% lower costs than Opus 5 on typical workloads. No head-to-head comparisons have been conducted between GPT-6 Sol and Opus 5.5 yet.

One notable change with GPT-6 is the improved prompt caching, which may matter more to developers building agents than the token prices themselves. OpenAI claims that the caching improvements deliver higher cache hit rates by default, with discounts of up to 90% on cached input tokens. Additionally, developers can now adjust reasoning effort and tool availability without invalidating the cache.

These caching improvements have reportedly cut the share of prompt tokens that require fresh processing by more than half across billions of requests to OpenAI models.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 3 other outlets

Read the original at thenewstack.io →

More in AI

More from Tuesday 22 September →