SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.