{
  "id": 10768430,
  "title": "Local LLM vs Claude Code: 96% of My Requests Failed Locally. Half the Steps Didn't.",
  "url": "https://urgent.news/2026/09/29/local-llm-vs-claude-code-96-of-my-requests-failed-locally-half-the",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-29T19:53:48.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/kenimo49/local-llm-vs-claude-code-96-of-my-requests-failed-locally-half-the-steps-didnt-1600"
  },
  "original_language": "en",
  "account": "The article compares the performance of running Claude Code locally versus using an online service. The author ran 100 requests that had previously been sent to Claude Code and checked whether a local model could have completed each one independently. Out of the 100 requests, 96 failed to execute fully on the local model. However, when the author examined the requests step-by-step, 97 out of 200 steps could be performed locally. The key difference is that the local model excels at tasks like web page extraction, which are not considered planning. The author implemented a set of rules to determine whether a step should be handled locally or sent to the larger Claude Code model. Initially, the author used a probability model to decide routing, but found that a simple rule-based approach outperformed the complex model in this case. The routing logic is straightforward: check each rule from top to bottom, and apply the first match found. If a rule matches, return that route decision. If no rules match, fall back to the default route, which is the online Claude Code model. To enable local execution for specific tools like WebFetch, the author added matching rules related to those tools. The article concludes with details on how to configure Ollama to output probability scores for model decisions, allowing the router to make informed routing choices based on the model's confidence.",
  "summary": "The pitch for running Claude Code on a local model goes like this: point it at Ollama, stop paying per token, keep your code on your machine. I have an RTX 4070 with qwen3.5:4b on it, so I wanted that to be true. Before swapping anything, I counted. I took 100 requests I had actually sent to Claude Code and asked, for each one, whether a local model could have done it end to end. 96 could not. So…",
  "key_points": [
    "96% of 100 requests failed locally",
    "97 out of 200 steps could be performed locally",
    "Rule-based routing outperformed probability model"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}