{
  "id": 5451099,
  "title": "I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster",
  "url": "https://urgent.news/2026/09/04/i-ran-3-open-weight-llms-head-to-head-on-a-24gb-mac-one-was-3x-faster",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-04T00:07:41.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/pitambarmahato/i-ran-3-open-weight-llms-head-to-head-on-a-24gb-mac-one-was-3x-faster-4k2g"
  },
  "original_language": "en",
  "account": "Three open-weight large language models (LLMs) were benchmarked on a 24GB Apple M2 Mac, yielding surprising results. OpenAI's gpt-oss-20B, Alibaba's Qwen3-14B, and Mistral's Mistral-Small-24B were evaluated on four tasks: GSM8K math reasoning, HumanEval+ code generation, IFEval instruction following, and MMLU academic knowledge. The headline results showed gpt-oss-20B leading in accuracy (100% on GSM8K and HumanEval+, 72% on MMLU) and speed (3-4x faster than Qwen3-14B). Qwen3-14B, considered the strongest 14B model, trailed in math and code tasks but remained the most well-rounded option. Mistral-Small-24B, the largest model, timed out on the larger benchmarks due to memory constraints. The study highlights that model performance depends on both size and available memory, with gpt-oss-20B offering a compelling balance of speed and accuracy for resource-constrained setups.",
  "summary": "I ran three Apache 2.0 open-weight LLMs — OpenAI's gpt-oss-20B, Alibaba's Qwen3-14B, and Mistral's Mistral-Small-24B — on the same Apple M2 24GB machine, through the same Ollama 0.12.x runtime, on the same four benchmark tasks. The fastest model is the one I didn't expect to win. The most accurate model is the one I expected to win. And the biggest model lost on both axes. The headline numbers:…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}