arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26991cs.AI

ASIL:用结构化状态与语义动作替代截图点击

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

发表机构上海交通大学 · 北京智源人工智能研究院
查看机构详情
  • Shanghai Jiao Tong University(上海交通大学)
  • BIGAI(北京智源人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Rui Xie, Lu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ASIL智能体-软件交互层,以结构化状态与语义动作替代低效的截图点击界面,在多应用基准任务中表现优于截图点击,还可提升Qwen系列模型的性能。

中文摘要 AI 辅助

强大的代码智能体可执行脚本、调用工具、管理文件,但许多重要应用仍主要通过图形用户界面(GUI)访问。本文认为,截图点击是软件操作智能体的低效界面:截图的状态信息不完整,GUI动作脆弱、语义薄弱,且与长程规划适配性差。我们提出ASIL(智能体-软件交互层),一种原生智能体界面,通过结构化JSON观测值与可执行代码的语义动作暴露软件,为每个应用提供最深可行的访问路径。我们在15个应用、300项单应用任务及80项多应用任务组成的基准上实例化ASIL。ASIL在闭源模型下准确率超80,且每项任务执行动作少于5个。在修复运行时环境、50步截图预算下,截图点击控制下的相同任务严格成功率分别为6.6%和26.6%,在更简单的OSWorld可比区间下升至15.0%和53.3%。在匹配任务上与应用原生接口对比,ASIL比LibreOffice的UNO API高28-38个严格点,仅与该链接的MCP内容契约相当。结构化模态也适配训练:小规模SFT将Qwen3.5-2B从58.0提升至72.1,Qwen3.5-9B从66.6提升至80.4,资源受限的在线策略强化学习进一步将其提升至74.4和82.2。

英文摘要

Powerful code agents can execute scripts, call tools, and manage files, yet many important applications remain accessible primarily through graphical user interfaces. We argue that screenshot-and-click is an inefficient interface for software-operating agents: screenshots are state-incomplete, and GUI actions are brittle, semantically weak, and poorly matched to long-horizon planning. We introduce ASIL (Agent-Software Interaction Layer), an agent-native interface that exposes software through structured JSON observations and code-executable semantic actions, realized through the deepest feasible access path for each application. We instantiate ASIL across 15 applications and a benchmark of 300 single-application and 80 multi-application tasks. ASIL reaches above 80 with closed models while executing fewer than five actions per task. Under a repaired runtime and a 50-step screenshot budget, the same tasks yield 6.6 and 26.6 strict success under screenshot-and-click control, rising to 15.0 and 53.3 on an easier OSWorld-comparable band. Against application-native interfaces on matched tasks, ASIL exceeds LibreOffice's UNO API by 28-38 strict points but only matches draw.io's MCP content contract. The structured modality also suits training: small-scale SFT raises Qwen3.5-2B from 58.0 to 72.1 and Qwen3.5-9B from 66.6 to 80.4, and resource-limited on-policy RL further raises them to 74.4 and 82.2.

补充信息

相关深度报道

↑