Urgent.News

What's breaking now, across thousands of outlets.

AI

I Benchmarked 4 Frontier LLMs on Catching ML's "Silent Killers" — DeepSeek-R1 Missed the Most Basic Bug

This is a submission for the Kaggle Benchmarking Challenge Most public AI leaderboards test if a model can write code or pass a syntax check. But in real-world Machine Learning, the most dangerous code isn't syntactically broken—it's methodologically flawed. It passes unit tests, shows a green dashboard, and then dies silently in production. For the Kaggle Benchmarking Challenge, I built "The…

For the Kaggle Benchmarking Challenge, the author constructed a benchmark called "The Silent Killer" to test if large language models (LLMs) could identify methodological flaws in machine learning pipelines. This benchmark focused on four specific types of "silent killers": data leakage, using the wrong metric, and target leakage.

The author evaluated four frontier LLMs against this benchmark: Gemini 3.7 Flash, Claude Sonnet 4.5, Grok 4.20 Reasoning, and DeepSeek-R1. While Gemini 3.7 Flash, Claude Sonnet 4.5, and Grok 4.20 Reasoning all successfully detected the flaws, DeepSeek-R1 failed to identify the most basic bug, missing it entirely. This discrepancy highlights a critical limitation in DeepSeek-R1's ability to comprehensively audit ML pipelines.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

llama-server's sleep mode loses or crashes on a request that arrives just before it sleeps

TL;DR : llama-server --sleep-idle-seconds N unloads the model after N idle seconds and is documented to reload it for "any new incoming task".

  • Llama-server sleep mode fails or crashes on requests arriving just before sleep
  • 17,653-token prompt remains in queue without response after server falls asleep
  • One-token completion avoids both crash and hang in 123 out of 123 attempts

O Harness, o SDD e o Vibe-Coding: uma abordagem de engenharia

Infraestrutura como contrato: o Harness contendo a entropia da automação Mover um software além da fase “mashup” exige substituir a intuição por uma governança de engenharia rígida.

  • O Harness, SDD e Vibe-Coding address software engineering challenges
  • Savage Worlds Dice Roller tested approach, automated RPG mechanics
  • Avoid "vibe coding" with AI-generated code, enforce business rules

I made CodeRabbit's reviews a third less noisy with an open-source Claude Code skill

AI code review has a noise problem. On a public benchmark of 50 real pull requests, CodeRabbit raised 300 issues. 77 of them were real bugs on the benchmark's list.

  • Tanay Kulkarni created three Claude Code skills for pr-proof
  • pr-proof kept 72 of 77 real bugs (93.5%) while reducing noise issues by 34%
  • pr-proof repo is available at https://github.com/TanayK07/pr-proof

More from Saturday 3 October →