{
  "id": 13706699,
  "title": "I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.",
  "url": "https://urgent.news/2026/10/11/i-turned-the-reasoning-dial-to-high-on-4-models-it-fixed-one-thing",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T11:54:46.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/abeera_lodhi/i-turned-the-reasoning-dial-to-high-on-4-models-it-fixed-one-thing-and-billed-me-for-everything-41o5"
  },
  "original_language": "en",
  "account": "Reasoning Dial, a Kaggle benchmark, focuses on the reasoning_effort setting which can be adjusted to none, low, medium, or high. Developers typically choose this setting by intuition, often increasing it for challenging questions and decreasing it when in production. However, the benchmark reveals that this decision can significantly impact accuracy, output tokens, and cost per correct answer.\n\nThe benchmark tested four AI models under fixed conditions on three types of tasks: Deduce (logic puzzles), Arith (word problems), and Distract (object counting). Across 7 model-task combinations, high reasoning effort led to a notable accuracy boost for gpt-5.4-mini, from 15% to 97.5%. However, in 7 out of 12 model-task combinations, increasing the reasoning effort resulted in higher costs ranging from 1.5 to 3.4 times for each correct answer, while accuracy remained unchanged.\n\nFurthermore, the benchmark uncovered inconsistencies in model behavior. Some models required up to 14 times more tokens at high reasoning effort, while others would reject low reasoning effort with an HTTP 400 error. Additionally, one model continued to operate at none, suggesting the reasoning dial may not uniformly affect all models.\n\nThe benchmark's pre-registration documented its hypotheses, statistical methods, and locked data to ensure rigorous analysis. The total cost for the main run was $4.20. Overall, the findings highlight that the reasoning_effort setting can significantly influence AI model performance, token usage, and cost, despite initial intuition that it might only affect the billing.\n\nFINAL ANSWER: The reasoning_effort setting in AI models can drastically impact accuracy, output tokens, and cost, with high settings sometimes improving accuracy but often inflating costs and causing inconsistent model behavior.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, \"Who gives the talk on Friday?\" With reasoning effort set to none , it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave the same wrong answer, word for word, on the second repeat. At high it spent 1,333 tokens, cost…",
  "key_points": [
    "High reasoning effort boosts gpt-5.4-mini accuracy from 15% to 97.5%",
    "Increasing reasoning effort leads to 1.5 to 3.4 times higher costs for correct answers",
    "Model behavior varies significantly with reasoning effort across tasks and models"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}