IAPO:面向多轮服务智能体信用分配的感知影响策略优化
IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents
浏览论文内容
中文总结 AI 辅助
本研究提出IAPO方法,将轨迹转化为影响依赖图以优化多轮服务智能体的信用分配,在三个服务智能体基准及BFCL-v4多轮任务上,性能优于多轮RL基线。
中文摘要 AI 辅助
大型语言模型(LLM)智能体越来越多地通过与用户及外部工具的多轮交互解决长周期任务。在这些场景中,相关任务信息往往随时间逐步展开,而非在初始提示中完全指定。服务智能体使这一挑战尤为具体:用户可能会澄清或修改其目标,而工具响应则提供后续决策所需的信息。因此,仅靠最终奖励无法表明哪些动作促成了任务解决。现有方法依赖其他轨迹或重采样延续的比较证据,或依赖单独构建的步骤级学习信号来优化信用分配。然而,完整的轨迹已记录了信息和错误如何在智能体动作间流动。我们提出感知影响策略优化(IAPO),它将每个轨迹表示为可训练智能体动作上的类型化影响依赖图,以用户和工具观测作为证据。IAPO 将支持使用和失败使用结构转化为路由权重,重新分配同一条轨迹级优势。对 Qwen3-4B 和 Qwen3-8B 的实验表明,在三个服务智能体基准——τ²-Bench、UserBench 和 AgentChangeBench 上,其性能优于多轮强化学习(RL)基线。BFCL-v4 多轮进一步显示,这些提升并未损害多轮函数调用性能。本研究推进了对多轮用户交互中信用分配的理解,并提供了一种从稀疏结果反馈中训练服务智能体的原则性方法。
英文摘要
Large Language Model (LLM) agents increasingly solve long-horizon tasks through multi-turn interactions with users and external tools. In these settings, relevant task information often unfolds over time rather than being fully specified at the initial prompt. Service agents make this challenge especially concrete: users may clarify or revise their goals, while tool responses provide information needed for subsequent decisions. Thus, a final reward alone cannot indicate which actions contributed to resolving the task. Recent methods rely on comparative evidence from other trajectories or resampled continuations, or on separately constructed step-level learning signals, to refine credit. However, a completed rollout already records how information and errors flow between agent actions. We introduce Influence-Aware Policy Optimization (IAPO), which represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations serving as evidence. IAPO converts support-use and failed-use structure into routing weights that redistribute the same trajectory-level advantage. Experiments with Qwen3-4B and Qwen3-8B demonstrate superior performance over multi-turn reinforcement learning (RL) baselines across three service-agent benchmarks: ${τ^2}$-Bench, UserBench, and AgentChangeBench. BFCL-v4 Multi-Turn further shows that these gains do not compromise multi-turn function-calling performance. This work advances the understanding of credit assignment in multi-turn user interactions and provides a principled approach to training service agents from sparse outcome feedback.
发表机构
- Fudan University(复旦大学)
- WeChat, Tencent Inc.(腾讯公司微信业务)
机构由 AI 辅助整理,请以论文原文为准。