TTPO: Test-Time Policy Optimization
Introduces a cutting-edge unsupervised training technique that advances reasoning capabilities beyond label-dependent methods.
AI Summary
Researchers propose TTPO, a novel test-time policy optimization method that enhances large language models' mathematical reasoning without ground-truth labels.
Excerpt
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts th
