Urgent.News

What's breaking now, across thousands of outlets.

AI

Claude Haiku 5.5 Pricing: The 100K-Token Rule for Agents

Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill. Anthropic released the model on 7 October 2026 on its own platform, AWS, Google Cloud and Microsoft Azure the same day. Output is $0.50 per million tokens at or below the threshold and $2.50 above it. The context window is 1M tokens with…

Claude Haiku 5.5 pricing introduces a 100K-token cliff, significantly increasing costs for agents handling prompts above that threshold. The model, launched on 7 October 2026, is the most affordable option available, costing $0.10 per million input tokens up to 100K tokens, and $0.50 per million above it. Output costs are $0.50 per million for prompts at or below the limit, and $2.50 above it.

The context window is 1M tokens, with 128K tokens available for output. For a high-volume agent loop with 8,000 input tokens and 500 output tokens per call, using Haiku 5.5 without caching would cost $1,050 per million calls. However, implementing prompt caching can reduce this to around $520 per million calls. The risk lies in exceeding the 100K token limit, which can happen quickly if context is not managed properly.

Key factors to consider include: counting tokens before sending requests to ensure they stay below 90,000 tokens, truncating tool output at the boundary, and routing large jobs to models built for long inputs. The Agent Run Cost Simulator can help model multi-step runs to identify potential costs. While the cheap tier offers savings, it does not eliminate all costs, such as retries and tool output, which can significantly impact overall expenses.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Running Local AI Speech-to-Text in Browser CPU with WebAssembly: Zero Server Uploads

Beaming confidential client interviews, confidential board meetings, or unreleased podcast recordings to remote cloud speech-to-text APIs is a massive privacy risk.

  • SolveMyMedia Transcribe processes audio locally in browser, no server uploads.
  • Uses Whisper Tiny English model via @xenova/transformers library.
  • Zero privacy risk, zero costs compared to cloud speech APIs.

Is a README for Humans or for LLMs?

I read the README of one of my own repositories and stopped at the tenth sentence. I had written it with my README skill, and the pre-publication check had said it was fine to publish.

  • Separate README sections for humans and LLMs improve clarity
  • Human-readable section contains concise, essential information
  • LLM-readable section includes additional facts for AI assistants

Kazakhstan seeks to build an AI-powered economy and manufacturing sector

Kazakhstan wants artificial intelligence to become more than a digital tool: the country is seeking to make it a new engine of economic growth, from universities and factories to energy and…

  • Kazakhstan aims to become an AI-powered economy and manufacturing sector.
  • AI should augment human capabilities and improve production efficiency.
  • Kazakhstan is developing all five layers of the AI economy.

More from Thursday 8 October →