Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemini 4 หลุด แต่ paper ที่ Google เพิ่งตีพิมพ์ตรวจสอบได้ทุกตัวเลข

Gemini 4 หลุด แต่ paper ที่ Google เพิ่งตีพิมพ์ตรวจสอบได้ทุกตัวเลข โดย Nokka (นก-กา) | 19 กันยายน 2026 บทความนี้เขียนโดย AI (โมเดล deepseek-v4.1-flash ของผู้ให้บริการ ollama-cloud) ผ่าน Hermes Agent จาก Nous Research ตรวจสอบและเรียบเรียงโดย Nokka ข้อความในเครื่องหมายคำพูดที่เป็นคำแปลเป็นคำแปลของผม ไม่ใช่สำเนาต้นฉบับ สัปดาห์นี้มีสองเรื่องที่เกี่ยวกับ Gemini 4 เกิดขึ้นพร้อมกัน เรื่องแรกคือตาราง…

In the week of September 19, 2026, two separate stories involving Gemini 4 emerged. The first story centered around a leaked benchmark table that circulated on various platforms. The second story involved a research paper published by Google DeepMind on September 14, which was verified through DOI and authorship details. While the leaked benchmark table lacked clear origin and verification, the published paper was more reliable and available for reading.

Google confirmed its involvement in the topic with direct statements, making two statements that lacked doubt. The first statement announced the launch of Gemini 3.6 Flash, while the second provided additional details during the earnings call. Sundar Pichai reiterated that Gemini 4 would be a more significant model compared to Gemini 3 Pro and mentioned Google's desire to compete at the frontier.

However, Pichai also mentioned that Google itself was uncertain about the extent of Gemini 4's capabilities. The leaked data included a table with unspecified authorship and no publication date, while the published paper came from Google DeepMind, featuring 17 authors and a DOI. The paper, titled "Dream-RSI: Recursive Self-Improvement through Evolving Worlds," was published on arXiv (2609.14858) on September 14, 2026, by Tong Zheng and 16 co-authors from the University of Maryland and Virginia.

The paper demonstrated that the team had found a way to test agent self-improvement within a simulator based on evolutionary worlds. The results showed that Gemini-3.1-Pro improved performance significantly, reducing the number of calls by 162 times, while still maintaining comparable performance in kernel benchmarks and layer normalization functions.

The most significant point highlighted in the paper was that the team did not fine-tune or update the model weights, but rather adjusted the search policy above the model. The gap between the paper's findings and the term "RSI" in general discourse remains significant.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 2 other outlets

Read the original at dev.to →

More in AI

More from Saturday 19 September →