LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

Chronological Source Flow
Back

AI Fusion Summary

Reinforcement learning for multi-turn search reasoning often depends on terminal outcome rewards, failing to distinguish between useful, redundant, or harmful intermediate interactions. To address this, LAPO is introduced as a self-generated process-supervision method utilizing backward leave-one-turn attribution. By replacing specific search turns and retrieval observations with a [DELETE] placeholder, LAPO calculates the Answer-Likelihood Gain. This process estimates each turn's contribution to the gold answer while preserving downstream interactions and evaluating early evidence within the reasoning context.
Community Comments
Loading updates...
0