探针即行动:通过主动视觉探测提升浏览器使用智能体
Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- StepFun
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出主动探测框架P2A,通过将DOM符号渲染为像素证据并在决策时对齐,提升浏览器智能体任务成功率,在VisualWebArena上显著提升专有和开源模型性能。
AI中文摘要:
浏览器使用智能体需要在结构化网页元数据与视觉信息之间实现无缝对齐,同时在长时间交互中保持相关上下文。现有接口通常依赖截图级动作预测或静态的标记叠加(Set-of-Marks),使得模型在每次操作前必须解决密集的DOM-像素对齐问题。我们提出了探针即行动(Probe to Act, P2A),一种用于浏览器智能体循环的主动探测框架,将对齐过程转移到决策时刻。P2A通过将按需的符号化DOM结构渲染回像素,解决了符号化DOM假设与截图布局之间的不对称桥梁问题。在提交状态改变型浏览器操作之前,智能体可以发出轻量级探针,将DOM句柄转换为像素证据,将屏幕区域映射回DOM候选,注册仅视觉目标,并提交已验证的注释。这些交错过程自然产生了基于证据的记忆:只有被探测、被操作或被明确提交的观察结果才会跨步骤保留,从而在长视野上下文中仅保留决策关键证据。P2A可以作为标准DOM+SoM接口下专有模型的提示策略,并可通过冷启动合成和自我引导的SFT(监督微调)蒸馏到开放权重模型。在三个浏览器使用基准测试中,P2A在任务成功率上对专有模型和微调模型均表现出明显提升;例如,在VisualWebArena上,它将Gemini-3-Pro从54.1%提升至61.2%,将Qwen3-VL-8B从24.6%提升至32.9%,同时以仅约1.2倍的峰值保留输入上下文匹配了昂贵的完整观察历史(约3倍)的性能。
英文摘要:
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history ($\sim$3$\times$) at only $\sim$1.2$\times$ the peak retained input context of action-only history.