Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.