{
  "id": 3299975,
  "title": "Tokenmaxxing is out. How to minimize AI spend without sacrificing security capability.",
  "url": "https://urgent.news/2026/08/25/tokenmaxxing-is-out-how-to-minimize-ai-spend-without-sacrificing",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-25T16:00:00.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/minimize-ai-security-spend/"
  },
  "original_language": "en",
  "account": "Recent findings reveal that premium AI models are financially burdensome for routine, high-volume security tasks, prompting a shift in strategy. A security operations team leader discovered that their detection costs were reduced to roughly $1 per day for trust and safety work, challenging the notion that enterprise AI is prohibitively expensive. The key to minimizing AI expenses lies in designing work correctly, focusing on the detection funnel. By narrowing and precisely targeting the cases fed to models, the cost per accurate outcome decreases dramatically. Rules-based pattern detection can capture a substantial portion of potential abuse before any model is engaged, followed by lightweight models handling the first pass. High-confidence outcomes are resolved automatically, while lower-confidence cases escalate to more capable models with broader context and stronger reasoning. The accuracy gap between lightweight and more advanced models is minimal, with only a 1-2% difference in accuracy for the specific use cases tested. Prompt engineering plays a crucial role, with well-crafted prompts requiring less complexity and fewer resources. For instance, a single prompt can address multiple scenarios and decision points, while simpler prompts may suffice for less complex cases. However, when dealing with ambiguous situations that require human judgment, such as differentiating legitimate security researchers from malicious actors or resolving interpretive disputes in security policies, human reviewers are indispensable. The objective is to allow humans to focus on the cases that genuinely require their expertise, rather than sifting through individual events. The strategy involves iterative prompt engineering, testing, and fine-tuning to adapt to evolving threats, much like traditional detection engineering. By treating AI spend as a design problem from the outset, teams can significantly reduce costs without compromising security capabilities.",
  "summary": "Security teams are discovering that the most capable AI models cost too much to run on routine, high-volume work, and The post Tokenmaxxing is out. How to minimize AI spend without sacrificing security capability. appeared first on The New Stack .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}