发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对计算机使用智能体操作层脆弱的问题,提出开源工具Tactile,将异构UI证据转换为基于动作的接口状态,通过特定循环操作。在相关任务中能提升Codex Success@100,证明可靠计算机使用需更强模型及可重用执行基础。
AI 中文摘要
计算机使用智能体正成为有能力的软件操作员,但其与桌面应用程序的接口往往是脆弱的操作层,将目标定位、动作执行和结果验证合并为一个模糊操作。我们提出了Tactile,一个开源工具层,为智能体提供更可靠的桌面使用“手脚”。Tactile将异构用户界面证据转换为基于动作的接口状态。智能体通过观察-定位-动作-验证循环操作,在macOSWorld风格任务中,添加Tactile可提高Codex Success@100。结果表明可靠的计算机使用不仅需要更强的模型,还需要可重用的执行基础。
英文摘要
Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor layer: they look at screenshots, predict coordinates, click, and hope that the visible state changed as intended. This collapses target grounding, action execution, and outcome verification into a single ambiguous operation. We present Tactile, an open-source tool layer that gives agents a more reliable "hands and feet" for desktop use. Tactile converts heterogeneous UI evidence--operating-system accessibility semantics, OCR-grounded text, and visual fallback regions--into action-grounded interface states: compact target candidates with source labels, roles or text, state, geometry, executable affordances, and verification cues. Agents operate through an observe-ground-act-verify loop that prefers native semantic actions when available, falls back to OCR-grounded coordinates when visible text is the best evidence, and keeps full provenance for replay and failure attribution. On macOSWorld-style tasks, adding Tactile improves Codex Success@100 from 41.1% to 50.0% overall and from 45.2% to 55.3% on accessibility-adapted tasks; a 96-task cross-agent subset shows consistent gains across Codex, Claude Code, OpenCode, and Goose. These results suggest that reliable computer use requires not only stronger models, but also a reusable execution substrate that exposes software actions as semantic, verifiable, and auditable objects rather than anonymous screen coordinates.