{
  "id": 13042523,
  "title": "LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls",
  "url": "https://urgent.news/2026/10/09/llms-pass-the-data-science-quiz-then-give-different-advice-a-kaggle",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T05:49:19.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/khushi886987/llms-pass-the-data-science-quiz-then-give-different-advice-a-kaggle-benchmark-of-36-measured-2i6b"
  },
  "original_language": "en",
  "account": "This report details a Kaggle Benchmark of 36 measured judgment calls involving large language models (LLMs). The benchmark involved 12 topics, each with three numeric variants, resulting in 36 cases. Each case presented a short competition situation where the correct call was established on Kaggle Playground S6E9 and S6E10, or it was a mathematical fact about the metric.\n\nThe benchmark measured two tasks for each model: ds-judgment (recognise) and ds-judgment-open (generate). In ds-judgment, models recognized the measured answer 94-100% of the time, but only gave the correct recommendation 56-81% of the time when asked without options. Every model lost 6-16 cases in the open-ended format. The benchmark also revealed that the biggest model in a family did not always perform best in the open format, and Haiku had the largest drop of any model despite a perfect multiple-choice score.\n\nThe failures clustered on four topics, with open-ended accuracy averaging 97% for every other topic. Topics T02 (Pearson vs Spearman), T03 (a 0 that means not applicable), T11 (picking final submissions), and T07 (more folds raised OOF, not the leaderboard) had lower accuracy. Three of these topics (T02, T07, T03) are genuine misses, where the textbook rule and measured answer disagree. The fourth topic, T11, turned out to be mostly the grading being too strict, which is explained later in the report.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked I spend a lot of time in Kaggle tabular competitions, and the decisions that cost me the most were never about model architecture. They were judgment calls: is this +0.0001 real? Should I append the original dataset? Which two submissions do I pick on the last day? (I once let the platform auto-pick, and an honest run…",
  "key_points": [
    "LLMs achieved 94-100% accuracy in recognizing measured answers",
    "Open-ended performance varied, with 56-81% correct recommendations",
    "Four topics showed lower accuracy, including T02, T03, T11, and T07"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}