Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where uncertainty remains high, and stop where estimates are precise or stable enough. The framework builds on hierarchical…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.