Urgent.News

What's breaking now, across thousands of outlets.

AI

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

Using AI to thwart election misinformation

Election rumors, such as unverified accounts of voter fraud, have run rampant in the U.S. in recent years. Powered by the speed of social media and artificial intelligence (AI), it can be hard for…

More from Thursday 27 August →