LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls
This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked I spend a lot of time in Kaggle tabular competitions, and the decisions that cost me the most were never about model architecture. They were judgment calls: is this +0.0001 real? Should I append the original dataset? Which two submissions do I pick on the last day? (I once let the platform auto-pick, and an honest run…
This report details a Kaggle Benchmark of 36 measured judgment calls involving large language models (LLMs). The benchmark involved 12 topics, each with three numeric variants, resulting in 36 cases. Each case presented a short competition situation where the correct call was established on Kaggle Playground S6E9 and S6E10, or it was a mathematical fact about the metric.
The benchmark measured two tasks for each model: ds-judgment (recognise) and ds-judgment-open (generate). In ds-judgment, models recognized the measured answer 94-100% of the time, but only gave the correct recommendation 56-81% of the time when asked without options. Every model lost 6-16 cases in the open-ended format. The benchmark also revealed that the biggest model in a family did not always perform best in the open format, and Haiku had the largest drop of any model despite a perfect multiple-choice score.
The failures clustered on four topics, with open-ended accuracy averaging 97% for every other topic. Topics T02 (Pearson vs Spearman), T03 (a 0 that means not applicable), T11 (picking final submissions), and T07 (more folds raised OOF, not the leaderboard) had lower accuracy. Three of these topics (T02, T07, T03) are genuine misses, where the textbook rule and measured answer disagree. The fourth topic, T11, turned out to be mostly the grading being too strict, which is explained later in the report.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.