My fine-tuned model scored 100%... The benchmark was lying
I fine-tuned Mistral 7B on my laptop to detect personal data in log lines and support messages. On my first test set it scored 100%. Perfect. Every single line classified correctly. I did not publish that number, because the same test set gave few-shot prompting 94%, and a six-point gap over a prompt you can write in five minutes is not a reason to fine-tune anything. The honest conclusion looked…
The author fine-tuned Mistral 7B on their laptop to detect personal data in log lines and support messages. Initially, the model scored 100% accuracy on their test set, but when they rebuilt the test set using real public data, the accuracy dropped to 95%. The same model and training recipe resulted in a 29-point gap between the fine-tuned and prompting methods.
The author discovered that their benchmark was influencing their results, as the same model and code gave different scores depending on the test set used. They concluded that the benchmark was "lying" and that LoRA, a technique that trains only a small subset of the model's parameters, is a more reliable method for fine-tuning.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.