发表机构
Salesforce AI Research(赛富时人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CLIFT利用保形自验证将裁判反馈转化为可复用的训练信号和无需裁判的测试时扩展信号,在多个网络智能体基准上达到最先进性能。
AI 中文摘要
开源网络智能体现在已足够强大,能够执行真实的浏览器任务,但使用强化学习训练它们仍依赖于较弱的监督信号:二元任务成功信号对于信用分配来说过于稀疏,而前沿语言模型裁判在每一步调用成本过高,且无法假设在部署时可用。我们提出了CLIFT,一种基于保形自验证的训练与测试时扩展方法。在训练阶段,智能体回答关于自身轨迹的自然语言验证问题;组合保形认证器仅保留其URL条件证据与训练时裁判一致的问询信号,通过极性感知提升分配带符号的信任权重,并将由此得到的验证器分数以不削弱裁判基线的方式融合进每步奖励中。在测试阶段,相同的认证库被冻结并复用为保形轨迹选择的结构化证据:智能体采样一条贪心轨迹及一条或多条多样化重试轨迹,自验证器对每条URL轨迹进行总结,保守的多数投票规则决定是否替换当前最优轨迹,全程无需调用任何外部裁判。这一统一机制支持三种场景。在WebArena Infinity上,CLIFT在开源网络智能体中达到了最先进性能。在VisualWebArena上,使用开放模型训练的库在测试时迁移至GPT-5.5,并在标准测试框架下达到最先进性能。在Online Mind2Web上,无需在该基准上训练智能体,仅翻译认证问询库即可在零样本评估中提升实时网络智能体的性能。综合这些结果,保形自验证可被视为将昂贵的裁判反馈转化为可复用训练信号及无需裁判的测试时扩展信号的有效途径。
英文摘要
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.