{
  "id": 9181713,
  "title": "tokens too cheap to meter",
  "url": "https://urgent.news/2026/09/22/tokens-too-cheap-to-meter",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-22T16:28:45.000Z",
  "source": {
    "name": "Lobsters",
    "slug": "lobsters",
    "url": "https://jyn.dev/tokens-too-cheap-to-meter/"
  },
  "original_language": "en",
  "account": "Machine learning intelligence is becoming increasingly affordable, with prices dropping by orders of magnitude each year. Large language models (LLMs) are expected to become integrated into every aspect of computing infrastructure within the next few years, as they are no longer solely used as standalone products. Local deployment of LLMs on commodity hardware is projected to become feasible in the next 3-6 years, making quality and accessibility the primary limiting factors for AI use rather than the sheer number of tokens.\n\nAI models can be proprietary, such as GPT-6 Astra, or open weight, like GLM-5.3-flash. Open weight models can be hosted by third-party providers like Z.ai or run locally on individual devices. Smaller local models, such as Muse Glimmer and Qwen3 Coder, are more common for local deployment, while larger models intended for hosting are typically more powerful.\n\nAdvancements in GPU efficiency are contributing significantly to the decreasing cost of using AI models. GPUs become exponentially more efficient with each generation, resulting in a logarithmic increase in efficiency, where a straight line on the graph represents an exponential improvement. This efficiency increase is unprecedented since the introduction of Moore's Law in the 1960s.\n\nThe cost of completing a task with a model is decreasing over time. Models are generally priced per token, with one token representing a fragment of a word, approximately 1.5 tokens per word. While smaller models may have lower costs per token, they often require more tokens to achieve the same results as larger models due to their need to think more or correct their initial drafts. The chart below shows the Pareto frontier of cost per task, illustrating the best tradeoff between cost and performance.\n\nAs LLMs continue to improve and become more cost-effective, the future of AI appears to be driven by the balance between quality, accessibility, and the cost of completing tasks.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}