发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对POR框架下GUI智能体反思器的决策依据薄弱问题,提出EFR两阶段反思器,解耦视觉差异提取与结果验证,在两个基准上提升了反思器准确率与端到端任务成功率。
AI 中文摘要
规划器-操作者-反思器(POR)框架被广泛应用于GUI智能体中,通过模块化协作在复杂任务中保持目标对齐。然而桌面GUI带来了一个关键挑战:大型密集界面常出现细微或分散的状态变化,这使得大部分负担落在反思器上,反思器必须比较动作前后的屏幕,而规划器和操作者仅基于单一状态推理。现有反思器将变化检测与结果验证合并为一步,导致证据不明确、决策依据薄弱。为解决此局限,我们提出证据优先反思(EFR),这是一种两阶段反思器,明确将动作诱导视觉差异提取与结果验证解耦。EFR通过标记集(Set-of-Marks)标注识别动作位置与候选变化区域,描述并过滤与动作相关的变化,再基于清理后的证据做出最终判断。这种证据推理解耦设计使反思更贴合屏幕转换,同时降低视觉搜索复杂度与推理负担。在OSWorld-Verified和WindowsAgentArena上的实验表明,EFR使反思器准确率提升7.11%,在两个基准上分别带来5.94%和4.95%的平均端到端任务成功率提升。
英文摘要
The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective alignment in complex tasks through modular collaboration. However, desktop GUIs introduce a key challenge: large, dense interfaces often exhibit subtle or scattered state changes, placing most of the burden on the reflector, which must compare pre- and post-action screens, while the planner and operator reason over a single state. Existing reflectors collapse change detection and outcome verification into one step, leaving evidence implicit and yielding weakly grounded decisions. To address this limitation, we propose Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification. EFR identifies the action location and candidate changed regions with Set-of-Marks annotations, describes and filters action-relevant changes, and makes the final judgment from the cleaned evidence. This evidence-reasoning decoupled design makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden. Experiments on OSWorld-Verified and WindowsAgentArena demonstrate that EFR improves reflector accuracy by 7.11%, yielding average end-to-end task success gains of 5.94% and 4.95% on the two benchmarks, respectively.