The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.