96.5% Accurate Spam Filter, and the 58 Spam Messages It Let Through
My spam classifier scores 96.5% accuracy on the test set. That sounds great. The confusion matrix tells a more useful story. The setup The dataset is mail_data.csv , with 5,572 messages: 4,825 ham (86.6%) and 747 spam (13.4%), no missing values. I mapped labels to 0 (spam) and 1 (ham), split 70/30 with random_state=3 , fit a TfidfVectorizer (English stopwords removed, vocabulary of 6,896 terms)…
A spam filter achieved an impressive 96.5% accuracy on test data. However, a closer look reveals some concerning details. With 4,825 legitimate messages and 747 spam messages in the dataset, the filter allowed 58 spam messages to slip through. Although the model almost never flagged a genuine email as spam, it managed to miss about one in four spam messages.
This trade-off between precision and recall is crucial to understand. The model prioritizes precision, correctly identifying 99.4% of spam emails while only misclassifying 1.3% as legitimate. On the other hand, it maintains a nearly perfect 99.9% recall for legitimate messages. The confusion matrix shows that 174 of the actual spam messages were correctly flagged as spam, while 58 were falsely identified as ham.
This discrepancy highlights the limitations of relying solely on the accuracy score when evaluating a spam filter.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.