Urgent.News

What's breaking now, across thousands of outlets.

AI

'That is where the machine starts winning on cost': Expert pits AMD Radeon AI PRO R9700 rig against ChatGPT and gives surprising verdict

A dual-GPU AMD workstation beats expensive cloud AI subscriptions on cost only once monthly token volume climbs high enough.

'That is where the machine starts winning on cost': Expert pits AMD Radeon AI PRO R9700 rig against ChatGPT and gives surprising verdict

Two AMD Radeon AI PRO R9700 cards, each with 32 GB of memory, were installed in a $18,775 workstation for testing. The setup measured electricity consumption, token throughput, and amortized hardware costs, then compared these figures to various cloud subscription pricing tiers. The cloud providers charge per million tokens generated, ranging from $1.20 for GPT-5.6 Luna to $30 for GPT-5.6 Sol, with mid-tier models like Claude Sonnet 5 at $10 and Claude Opus 5 at $25.

Using both AMD cards together, the workstation produced 156.2 tokens per second while handling eight simultaneous users. However, with a technique called Multi-Token Prediction, the throughput increased to 320.2 tokens per second, nearly doubling the output quality while maintaining the same cost. At this speed, the hardware starts outperforming GPT-5.6 Sol within 3.5 hours of weekly use, surpassing Claude Opus 5 after 4.2 hours, and outperforming Gemini 3.1 Pro after 8.8 hours.

For teams generating over 20 million tokens per month against GPT-5.6 Sol pricing, this hardware setup results in significant savings, costing approximately $6,262 annually in electricity and amortized hardware, compared to roughly $18,000 annually in matching Sol fees. This $11,738 yearly difference illustrates the cost-effectiveness of the AMD rig for heavy monthly usage.

If a company opts for GPT-5.6 Luna, priced at $1.20 per million tokens, the AMD workstation would need 94.3 hours of weekly use to match the lower cloud subscription cost. Since a week contains only 168 hours, achieving this break-even point is challenging without constant, saturated usage. Electricity itself was a minor factor, as both cards together consumed between 310 and 510 watts during sustained load conditions.

The economics favors AMD's hardware for larger teams producing higher volumes of tokens. A smaller team generating five million tokens monthly would spend $6,262 annually to save $720 compared to cloud costs, resulting in a loss. The hardware only becomes financially advantageous when usage is high enough to close a substantial yearly cost gap.

A single R9700 card managed a 27 billion parameter AI model independently, processing 34.5 tokens per second without needing a second card. Running just one card also lowers the effective break-even point, as it consumes fewer watts under equivalent load conditions.

Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techradar.com →

More in AI

Grok Bot and the Rise of AI Teammates

In my last post, I talked about the potential risks AI could bring to humans in the not-too-distant future. Today, I want to look at another side of AI: what happens when we give AI more freedom to…

  • Grok Bot is an autonomous AI agent for workflows
  • Participates in 72-hour demo event with user-defined goals
  • Capabilities include web search, command execution, API calls

More from Sunday 13 September →