{
  "id": 5678180,
  "title": "When Benchmark lies to you, SWE-Bench ProMax with the real score that the best model can do is only 41.2%.",
  "url": "https://urgent.news/2026/09/05/benchmark-swe-bench-promax-41-2",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-05T00:26:18.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/sarantoon/emuue-benchmark-okhkkhun-swe-bench-promax-kabkhaaenncchringthiiomedlekngsudthamaidaekh-412-1fl5"
  },
  "original_language": "th",
  "account": "A new research paper on arXiv has introduced a benchmark called SWE-Bench ProMax, which evaluates the performance of AI models in refactoring code across multiple files and languages. The results show that even the best models achieve a success rate of only 41.2%, contradicting the high scores of up to 90% reported by some AI companies. The study also found that open-weight models, such as GLM-5, can perform similarly to proprietary models at a significantly lower cost. The researchers highlight the challenges of long-horizon tasks, where models struggle to maintain context across multiple files.",
  "summary": "เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2% โดย Nokka (นก-กา), นักเขียนอิสระสายเทคโนโลยี ผู้เขียนบทความอธิบายเทคโนโลยีให้คนทั่วไปเข้าใจ 30+ บทความบน dev.to | 5 กันยายน 2026 บทความนี้เขียนโดย AI (glm-5.3 via ollama-cloud) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์, Nokka (นก-กา), อ้างอิงจาก paper วิจัย SWE-Bench ProMax บน arXiv ฉบับเต็ม เลข…",
  "key_points": [
    "SWE-Bench ProMax model achieves only 41.2% performance",
    "Benchmark data misleading, contrary to 2026 AI trends",
    "Nokka, tech writer, explains complex AI concepts in simple terms"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}