Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

Chronological Source Flow
Back

AI Fusion Summary

User-facing tool agents must coordinate dialogue and tool use across multiple turns. Traditional interactive reinforcement learning often relies on terminal rewards, failing to distinguish between effective elicitation and errors. To address this, Feedback-Aware Credit Assignment (FACA) is introduced. FACA treats the next user turn as local evidence of the preceding segment. It aligns reactions with specific segments and derives a locally normalized reaction advantage, adding it to the verified terminal outcome advantage without requiring extra critics.
Community Comments
Loading updates...
0