AI 中文总结
研究长期运行网络代理易偏离轨道问题,提出FCPAgent框架,将计划步骤表示为FCU,通过计划-测试-修复循环执行,经混合承诺测试模块和范围感知修复提高成功率,在WebArena上相比基线有显著提升。
AI 中文摘要
长期运行的网络代理在最终失败前常偏离轨道:即便当前状态、复用技能或计划假设不再支持用户指令,轨迹仍可能局部合理。现有代理能规划、反思或复用经验,但计划很少指明仍应信任当前步骤的证据。我们提出了FCPAgent,一种用于健壮的长期网络代理的可证伪承诺规划框架。FCPAgent将每个计划步骤表示为一个可证伪承诺单元(FCU),包括基于可复用技能的子目标以及确认证据、证伪证据和置信度得分。执行组织为计划-测试-修复循环。混合承诺测试模块在候选动作修改浏览器前及执行后检查观察结果,通过轻量级证据匹配与基于大语言模型的诊断验证提高效率。当证据证伪承诺时,范围感知修复将矛盾定位到执行、技能或规划层面并修正最小的适当部分。在WebArena上,FCPAgent相对于最强基线平均成功率有13.8%的相对提升,在长期任务上增益尤其大。
英文摘要
Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reuse experience, but their plans rarely specify the evidence under which an active step should still be trusted. We propose FCPAgent, a falsifiable commitment planning framework for robust long-horizon web agents. FCPAgent represents each plan step as a Falsifiable Commitment Unit (FCU): a subgoal grounded in a reusable skill, together with confirming evidence, falsifying evidence, and a confidence score. Execution is organized as a plan-test-repair loop. The hybrid commitment testing module checks candidate actions before they modify the browser and checks observations after execution; for efficiency, it combines lightweight evidence matching with LLM-based diagnostic verification. When evidence falsifies a commitment, scope-aware repair localizes the contradiction to the execution, skill, or planning level and revises the smallest adequate part. On WebArena, FCPAgent achieves a 13.8% relative improvement in average success over the strongest baseline, with especially large gains on long-horizon tasks.