SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Administrator 0 阅读

AI Digest - ArXiv AI

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. We show that trajectory centering generally does not center the token-weighted quantity consumed by the actor, and fix the mismatch by standardizing


Source: ArXiv AI