Urgent.News

What's breaking now, across thousands of outlets.

AI

When Benchmark lies to you, SWE-Bench ProMax with the real score that the best model can do is only 41.2%.

เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2% โดย Nokka (นก-กา), นักเขียนอิสระสายเทคโนโลยี ผู้เขียนบทความอธิบายเทคโนโลยีให้คนทั่วไปเข้าใจ 30+ บทความบน dev.to | 5 กันยายน 2026 บทความนี้เขียนโดย AI (glm-5.3 via ollama-cloud) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดยมนุษย์, Nokka (นก-กา), อ้างอิงจาก paper วิจัย SWE-Bench ProMax บน arXiv ฉบับเต็ม เลข…

Translated from Thai Read in Thai

A new research paper on arXiv has introduced a benchmark called SWE-Bench ProMax, which evaluates the performance of AI models in refactoring code across multiple files and languages. The results show that even the best models achieve a success rate of only 41.2%, contradicting the high scores of up to 90% reported by some AI companies.

The study also found that open-weight models, such as GLM-5, can perform similarly to proprietary models at a significantly lower cost. The researchers highlight the challenges of long-horizon tasks, where models struggle to maintain context across multiple files.

Written by urgent.news from Dev.to's report — not a translation of it. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

A 0.3% gap but a 2-time price difference, read the Terminal-Bench 4.0 table as follows:

ช่องว่าง 0.3% แต่ราคาต่าง 2 เท่า, อ่านตาราง Terminal-Bench 4.0 ให้เป็น โดย Nokka (นก-กา), นักเขียนอิสระสายเทคโนโลยี ผู้เขียนบทความอธิบายเทคโนโลยีให้คนทั่วไปเข้าใจ 30+ บทความบน dev.to | 5 กันยายน 2026…

More from Saturday 5 September →