arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24174cs.AI

面向GUI奖励建模的任务自适应评分规则

Task-Adaptive Rubrics for GUI Reward Modeling

Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有GUI奖励验证器任务自适应不足的问题,提出AdaptRubric框架,经实验验证其在离线奖励评估和在线强化学习中均优于现有方法,提升了F1值与任务成功率。

中文摘要 AI 辅助

近期关于GUI智能体的研究日益关注结果奖励建模,该方法通过判断执行的轨迹是否满足用户指令隐含的成功标准来分配结果奖励。然而,现有的GUI奖励验证器往往未明确说明如何为每个任务实例构建这些标准:无论是使用通用评分规则结构还是隐式模型推理,其判断标准都不够任务自适应,可能跨任务传递检查、忽略当前指令中的具体约束,或因强制未明确要求而变得过于严格。为解决这一局限,我们提出AdaptRubric,一种由粗到细的评分规则框架,通过类别级粗阶段和实例级细阶段构建任务自适应判断标准:首先将指令路由至GUI任务族并检索可复用的任务族标准,完成类别级粗评分规则检索;再进行实例级细评分规则生成,提取当前指令中具体值、范围和约束的紧凑线索。在离线奖励评估和在线强化学习优化中,AdaptRubric的表现始终优于现有奖励智能体:在匹配图像预算下,其F1值较基线平均提升3.6个百分点,任务成功率则获得4.23个百分点的增益。

英文摘要

Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks across tasks, overlook concrete constraints in the current instruction, or become overly strict by enforcing unstated requirements. To address this limitation, we propose AdaptRubric, a Coarse-to-Fine Rubrics Framework that constructs task-adaptive judging criteria through a category-level coarse stage and an instance-level fine stage. AdaptRubric performs category-level coarse rubric retrieval by routing the instruction to a GUI task family and retrieving reusable task-family criteria, then conducts instance-level fine rubric generation to surface compact cues for concrete values, scopes, and constraints in the current instruction. Across offline reward evaluation and online reinforcement learning optimization, AdaptRubric consistently outperforms prior reward agents, improving F1 by 3.6 points over the baseline average under a matched image budget and yielding a 4.23-point task-success gain.

发表机构

  • Zhejiang University(浙江大学)
  • MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)

机构由 AI 辅助整理,请以论文原文为准。

↑