Your First Model Should Be Embarrassing
TL;DR: Before anything clever, fit a model that guesses the majority class and a plain logistic regression. Together they take seconds, and they're the only way to tell what the expensive model actually bought you. On six public datasets, guessing alone was 95.3% accurate on hypothyroid. Default boosted trees beat logistic regression by 2 to 15 AUC points. A 200-fit hyperparameter search added…
When starting a machine learning project, it is essential to build a baseline model that sets the standard for comparison. The simplest baseline models are those that ignore the features entirely, such as guessing the majority class or using a logistic regression model. By fitting these eminently basic models first, you can quickly gauge how much value the more complex models provide.
When applied to six public datasets from the Penn Machine Learning Benchmark (PMLB) collection, it was discovered that a dummy classifier that ignores features achieved up to 95.3% accuracy on the hypothyroid dataset, while logistic regression achieved 96.2% accuracy. On the other hand, gradient-boosted trees provided only marginal improvements of 2 to 15 AUC points across most datasets.
A hyperparameter search added only about 0.5 points of improvement, with gains varying between -0.3 and +0.2 points on average across the four random splits.
In conclusion, it is crucial not to overlook the importance of building a simple baseline model before diving into more complex ones. This approach ensures that the more sophisticated models you develop provide genuinely meaningful improvements and helps avoid false optimism in model performance.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.