Urgent.News

What's breaking now, across thousands of outlets.

AI

What Does a Local LLM Actually Cost per Month? I Read the Meters.

What Does a Local LLM Actually Cost per Month? I Read the Meters. The Local LLM Lab — Part 5 One controlled experiment. One number. One verdict. The question nobody answers in the local-LLM hype is the boring one: what does the electricity bill say? Not "how many tokens per second." Not "how many GB of VRAM." The bill. The one that arrives on the first of the month and is the only number that…

The cost of running a local LLM inference stack on a single NVIDIA RTX 3090 GPU can be as low as €2.00 per month, according to a recent analysis of power consumption and electricity costs. The experiment involved running Whisper transcription as a permanent service, an embedding model for a Retrieval-Augmented Generation (RAG) pipeline, and a 27B chat model on a second machine. The machine was equipped with a power meter and a dual-rate electricity tariff (0.30 BGN/kWh during the day, 0.18 BGN/kWh at night).

The average power draw of the GPU while hosting the Whisper service was 22 W, with peaks of up to 120 W during transcription. The embedding model added only 23 cents to the monthly bill. The total power consumption across all services averaged 25 W over the 30-day period. The analysis concludes that the electricity cost of running a local inference stack is a rounding error compared to the hardware cost and time required to maintain the setup.

For home inference stacks, the main cost is the hardware itself, with electricity consumption being minimal.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Jev's Decision Model + NylonME: The 'Disassembly Era' of AI Is Here

A Model That Can't Write Just Broke the Internet On September 15, TypeSafe AI released a model called Jev. The company was founded by former OpenAI researcher Diogo Almeida — a core author of RLHF and…

  • Jev AI model delivers structured decisions with calibrated probabilities
  • 20-200 times faster and 40-400 times cheaper than LLMs for decision-making
  • Jev exemplifies shift from language generation to decision-making components

More from Sunday 20 September →