Urgent.News

What's breaking now, across thousands of outlets.

Tech

My fine-tuned model scored 100%... The benchmark was lying

I fine-tuned Mistral 7B on my laptop to detect personal data in log lines and support messages. On my first test set it scored 100%. Perfect. Every single line classified correctly. I did not publish that number, because the same test set gave few-shot prompting 94%, and a six-point gap over a prompt you can write in five minutes is not a reason to fine-tune anything. The honest conclusion looked…

The author fine-tuned Mistral 7B on their laptop to detect personal data in log lines and support messages. Initially, the model scored 100% accuracy on their test set, but when they rebuilt the test set using real public data, the accuracy dropped to 95%. The same model and training recipe resulted in a 29-point gap between the fine-tuned and prompting methods.

The author discovered that their benchmark was influencing their results, as the same model and code gave different scores depending on the test set used. They concluded that the benchmark was "lying" and that LoRA, a technique that trains only a small subset of the model's parameters, is a more reliable method for fine-tuning.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Rationale

The rationale is simple - I want my project repos to be self-sufficient in the sense that it includes both its source code as well as goals/issues/docs.

More from Tuesday 11 August →