发表机构
Ant International(蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放式真实世界交互中基于组的RL轨迹行为不可比导致的奖励公平问题,提出ARC训练方案,结合\textit{inter}范式与\textit{inter-86K}语料库,提升工具使用基准并缩短首次token时间。
AI 中文摘要
开放式真实世界交互允许多种有效行为:智能体可直接回答、请求澄清、提供进度更新,或在行动前确认。这种灵活性打破了基于组的强化学习(RL)背后的核心假设:组内比较的轨迹不再保证行为上的可比性。因此,奖励模型对交互风格的偏好会扭曲相对优势,引导优化走向奖励偏好的行为,而非符合上下文的行为。我们将此形式化为「奖励公平问题」,并提出ARC(Advantage Regularization via Conditioning,基于条件的优势正则化),这是一种训练方案,通过策略条件化的轨迹分组、混合奖励和熵正则化来恢复更公平的相对比较。我们在提出的\textit{inter}中研究ARC,这是一种响应式、可引导、感知执行的用户-智能体交互新范式,将用户可见的通信与潜在推理及工具使用解耦。\textit{inter}还提供了注释和蒸馏管道,用于构建\textit{inter-86K}——我们的策略注释训练语料库,用于监督学习和RL训练。实证结果表明,ARC显著增强了核心$\tau/\tau^2$工具使用基准,而\textit{inter}相较于思考式基线,将首次 token 时间从4.91秒降至1.27秒。这些结果共同表明,开放式交互学习的核心瓶颈不仅在于智能体如何获得奖励,还在于其行为是否从一开始就得到公平比较。ARC的实现和\textit{inter-86K}训练数据将被发布。
英文摘要
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a \textit{reward fairness problem} and propose \textbf{ARC} (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core $τ/τ^2$ tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.