Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.