When Benchmark lies to you, SWE-Bench ProMax with the real score that the best model can do is only 41.2%.
เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2% โดย Nokka (นก-กา), นักเขียนอิสระสายเทคโนโลยี ผู้เขียนบทความอธิบายเทคโนโลยีให้คนทั่วไปเข้าใจ 30+ บทความบน dev.to | 5 กันยายน 2026 บทความนี้เขียนโดย AI (glm-5.3 via ollama-cloud) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์, Nokka (นก-กา), อ้างอิงจาก paper วิจัย SWE-Bench ProMax บน arXiv ฉบับเต็ม เลข…
A new research paper on arXiv has introduced a benchmark called SWE-Bench ProMax, which evaluates the performance of AI models in refactoring code across multiple files and languages. The results show that even the best models achieve a success rate of only 41.2%, contradicting the high scores of up to 90% reported by some AI companies.
The study also found that open-weight models, such as GLM-5, can perform similarly to proprietary models at a significantly lower cost. The researchers highlight the challenges of long-horizon tasks, where models struggle to maintain context across multiple files.
Written by urgent.news from Dev.to's report — not a translation of it. Machine-written — may contain errors; check the original before relying on it.