Urgent.News

What's breaking now, across thousands of outlets.

AI

The ML you need to operate LLMs, not train them

You do not need to understand backpropagation to run a large language model well in production. You need a smaller, more practical thing: the operator's mental model. What a token actually costs you, what a context window actually bounds, what temperature and top_p actually do, and why evals catch what your own reading of the output will not. None of this is deep learning theory. All of it is the…

Running large language models in production does not require deep understanding of backpropagation. The key mental model is how much each token costs, the impact of context windows, the distinction between tokens and words, and the behavior of temperature and top_p parameters. Tokens differ from words; longer or uncommon words split into multiple tokens, while punctuation and non-tokenised strings can cost more.

Tokenisation counts vary across model families due to different tokenisers, which can cause discrepancies in cost estimation and context limits. Context windows represent a shared budget, not separate allowances for different components of a conversation. Model-specific constraints on temperature and top_p parameters must be understood, as AWS does not consistently document these details.

Treating temperature and top_p as interchangeable dials is incorrect, as only one can be specified per model on certain Bedrock configurations. Extensive thinking features are incompatible with temperature, top_p, or top_k modifications, requiring the removal of these parameters when thinking is enabled. Ultimately, a model's hallucinations result from its objective of predicting the most statistically plausible next token, not a malfunction in its operation.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Recursive Governance: When Agents Write the Rules They Execute

We almost lost forty modules to a file that never changed. The sync job ran every six hours, mirroring our shared working directory into a test workspace.

  • Agents create and enforce their own rules, leading to conflicts and inconsistencies.
  • Recursive governance system protects modules and enables rule auditing and validation.

Your AI Agent Will Do Something Terrible. Here's How to Survive It.

Here's a pattern I keep seeing. A team wires up an AI agent that can do real things — send emails, run commands, query and modify the database, call external APIs. The demo is magical.

  • Implement least privilege principle for AI agents
  • Require human approval for high-impact actions
  • Treat all input as untrusted to prevent prompt injection

Best LLM Gateways in 2026: A Production-Ready Comparison

TL;DR An LLM gateway is production-ready when it adds negligible latency under load, fails over across providers without application code, enforces budgets per team, governs MCP tool calls, and…

  • Bifrost is open-source Go-based LLM gateway with high performance
  • Favored for enterprises needing scalability and reliability
  • Not published details for Kong AI Gateway and Cloudflare AI Gateway

More from Tuesday 6 October →