Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Administrator 0 阅读

AI Digest - ArXiv AI

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. These algorithms differ in their statistical structure, compute accounting, and failure modes. Treating these procedures as interchangeable un


Source: ArXiv AI