Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems;…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.