My fine-tuned model scored 100%... The benchmark was lying
I fine-tuned Mistral 7B on my laptop to detect personal data in log lines and support messages. On my first test set it scored 100%. Perfect. Every single line classified correctly. I did not publish that number, because the same test set gave few-shot prompting 94%, and a six-point gap over a prompt you can write in five minutes is not a reason to fine-tune anything. The honest conclusion looked…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.