发表机构
UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PAIR和RePAIR,用于评估并提升移动GUI智能体在个性化界面中的跨用户可靠性,实验表明RePAIR显著提高任务成功率。
AI 中文摘要
移动图形用户界面智能体日益在受用户历史和偏好影响的界面上运行,但其在不同用户间的可靠性仍未得到充分探索。我们引入了PAIR(个性化应用状态实例化与渲染),一个用于构建用户条件化应用状态的流水线,使得能够在不同用户间对同一任务进行受控评估。我们进一步引入了RePAIR(具有个性化感知交互奖励的强化学习),一种训练方法,从子目标结果的跨用户差异中学习,以提高在用户条件化移动环境中的可靠性。在六个智能体上,我们发现不同用户间任务成功率存在显著差异,且在用户条件化用户界面情境中子目标达成率持续较低(6.98至15.4个百分点)。对于从每个用户自身内容中提取的个人目标,这一差距进一步扩大(8.77至22.0个百分点)。这些情境中的失败常常涉及选择另一个项目而非预期目标,尤其是在目标暴露之前。最后,RePAIR在未见过的用户上,相较于其监督微调基线,将用户条件化SAR提高了5.87个百分点,全成功提高了7.50个百分点,整体任务SR提高了9.42个百分点,提供了初步证据表明,明确从跨用户差异中学习可以提高图形用户界面智能体的可靠性。
英文摘要
Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation and Rendering), a pipeline for constructing user-conditioned application states that enables controlled evaluation of the same task across different users. We further introduce RePAIR (Reinforcement learning with Personalization-Aware Interaction Rewards), a training approach that learns from cross-user differences in subgoal outcomes to improve reliability across user-conditioned mobile environments. Across six agents, we find substantial variation in task success across users and consistently lower subgoal achievement in user-conditioned UI contexts (6.98 to 15.4 pp). This gap further increases for personal targets drawn from each user's own content (8.77 to 22.0 pp). Failures in these contexts frequently involve selecting another item instead of the intended target, particularly before target exposure. Finally, RePAIR improves user-conditioned SAR (+5.87 pp), all-success (+7.50 pp), and overall Task SR (+9.42 pp) over its supervised fine-tuning parent on unseen users, providing initial evidence that explicitly learning from cross-user variation can improve GUI-agent reliability.
Comments23 pages, 12 figures