面向更好的多轮用户交互智能体:下一轮用户输入不止是上下文
Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context
浏览论文内容
中文总结 AI 辅助
该研究针对多轮用户交互智能体的信用分配问题,提出FACA方法,在多领域测试中提升了8B和14B规模模型的性能,验证了下一轮用户反应可提供局部信用。
中文摘要 AI 辅助
面向用户的工具智能体必须在用户目标随多轮交互展开时协调对话与工具使用。然而,交互式强化学习通常将每个回合简化为一个最终奖励,对有效引导、错误及后续修正分配相同的信用。下一轮用户输入不止是上下文:它还提供了关于前序用户交互片段的噪声、时间局部证据。我们提出反馈感知信用分配(Feedback-Aware Credit Assignment,简称FACA),该方法将每个反应与该片段对齐,推导局部归一化的反应优势,并将其添加到已验证的最终结果优势中,无需额外的评判器或回合。在模拟环境中与仅基于结果的交互式GRPO控制(匹配了可见对话、初始化、回合及优化设置)相比,FACA在8B和14B规模下,三个独立训练运行的九领域τ族平均值分别提升了5.91和10.22个百分点。增益集中在电信领域;在8B规模下,随机化反应极性会消除电信领域的增益。在Pare-Bench和Co-Gym上的零样本测试也呈现相同的排序。这些结果表明,下一轮用户反应为改进多轮用户交互智能体提供了可操作的局部信用。
英文摘要
User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repair. The next user turn is more than context: it also provides noisy, temporally local evidence about the preceding user-to-user segment. We introduce \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}), which aligns each reaction with that segment, derives a locally normalized reaction advantage, and adds it to verified terminal outcome advantage without an extra critic or rollout. Against an outcome-only Interactive GRPO control matched in simulator, visible dialogue, initialization, rollout, and optimization, \textsc{FACA} improves the nine-domain $τ$-family average across three independently trained runs by 5.91 and 10.22 percentage points at 8B and 14B, respectively. Gains concentrate in Telecom; at 8B, randomizing reaction polarity removes the Telecom gain. The same ordering holds zero-shot on Pare-Bench and Co-Gym. These results demonstrate that next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents.
发表机构
- Fudan University(复旦大学)
- Ant International, Ant Group(蚂蚁集团蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。