发表机构
The University of Queensland; Griffith University(昆士兰大学; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究个性化智能体推理问题,提出基于强化微调框架ODYSSE,核心方法是逐情节GRPO,通过情节级奖励机制等解决长动作跨度等问题,经实验验证其在个性化GUI推理任务中优于其他模型,有效提升个性化智能体推理能力。
AI 中文摘要
智能体系统在与现实世界环境交互、利用外部工具和为用户提供服务的能力方面迅速发展。然而,与具有明确指令的自然世界任务不同,以人类为中心的场景具有模糊请求的特点,导致庞大的开放式解决方案空间。因此,解码用户的个性化偏好对于缩小候选解决方案空间至关重要。这引入了新挑战——个性化智能体推理,需要智能体与用户和环境共同交互以提供个性化服务。本文提出ODYSSE,一种用于个性化智能体推理的强化微调(RFT)框架。其核心是逐情节GRPO(ESPO),是组相对策略优化(GRPO)的新扩展,旨在解决个性化智能体推理中的长动作跨度和强跨步骤依赖性。ESPO引入情节级奖励机制和情节优势估计,而不是独立优化单个步骤,使上游证据能有效指导下游个性化决策,并允许智能体在多个交互步骤中逐步解决模糊的用户请求。还提出情节批量采样器,将同一情节中的动作分组为统一训练批次,促进ESPO下的连贯优化。在现实的长跨度个性化GUI推理任务上评估ODYSSE。实验结果表明ODYSSE始终优于专业和通用LVLMs,突出其在个性化智能体推理方面的有效性。
英文摘要
Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-centered scenarios are characterized by ambiguous requests that lead to large, open-ended solution spaces. Decoding users' personalized preferences is therefore essential for narrowing the candidate solution space. This introduces a new challenge, personalized agentic reasoning, which requires agents to jointly interact with both users and environments to deliver personalized services. In this paper, we present ODYSSE, a Reinforced Fine-Tuning (RFT) framework for personalized agentic reasoning. At its core, ODYSSE proposes Episode-wise GRPO (ESPO), a novel extension of Group Relative Policy Optimization (GRPO) designed to address long action horizons and strong cross-step dependencies in personalized agentic reasoning. Rather than optimizing individual steps independently, ESPO introduces an episode-level reward mechanism together with episodic advantage estimation, enabling upstream evidence to effectively guide downstream personalized decisions and allowing agents to progressively resolve ambiguous user requests across multiple interaction steps. We further propose an episodic batch sampler that groups actions from the same episode into unified training batches, facilitating coherent optimization under ESPO. We evaluate ODYSSE on realistic long-horizon personalized GUI reasoning tasks. Experimental results demonstrate that ODYSSE consistently outperforms both specialist and general-purpose LVLMs, highlighting its effectiveness for personalized agentic reasoning.