CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.