{
  "id": 12336061,
  "title": "The ML you need to operate LLMs, not train them",
  "url": "https://urgent.news/2026/10/06/the-ml-you-need-to-operate-llms-not-train-them",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-06T08:29:02.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/rajmurugan/the-ml-you-need-to-operate-llms-not-train-them-2io9"
  },
  "original_language": "en",
  "account": "Running large language models in production does not require deep understanding of backpropagation. The key mental model is how much each token costs, the impact of context windows, the distinction between tokens and words, and the behavior of temperature and top_p parameters. Tokens differ from words; longer or uncommon words split into multiple tokens, while punctuation and non-tokenised strings can cost more. Tokenisation counts vary across model families due to different tokenisers, which can cause discrepancies in cost estimation and context limits. Context windows represent a shared budget, not separate allowances for different components of a conversation. Model-specific constraints on temperature and top_p parameters must be understood, as AWS does not consistently document these details. Treating temperature and top_p as interchangeable dials is incorrect, as only one can be specified per model on certain Bedrock configurations. Extensive thinking features are incompatible with temperature, top_p, or top_k modifications, requiring the removal of these parameters when thinking is enabled. Ultimately, a model's hallucinations result from its objective of predicting the most statistically plausible next token, not a malfunction in its operation.",
  "summary": "You do not need to understand backpropagation to run a large language model well in production. You need a smaller, more practical thing: the operator's mental model. What a token actually costs you, what a context window actually bounds, what temperature and top_p actually do, and why evals catch what your own reading of the output will not. None of this is deep learning theory. All of it is the…",
  "key_points": [
    "Tokens differ from words; longer or uncommon words split into multiple tokens.",
    "Temperature and topp parameters are model-specific constraints, not interchangeable dials."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}