arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35789cs.SEcs.HC

Agent可调用功能覆盖率:衡量软件对AI Agent的就绪度

Agent-Callable Feature Coverage: Measuring Software Readiness for AI Agents

Zedong Peng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ACFC指标和ARC框架,量化软件对AI Agent的就绪度,通过30个系统的实证和10个系统的测试,发现可发现性显著影响Agent执行,且多数系统在文档和安全治理上滞后。

中文摘要 AI 辅助

AI Agent已经能够通过截图进行基于视觉的计算机操作来使用图形软件,因此紧迫的问题不是Agent能否操作软件,而是软件通过结构化、可控的渠道对Agent的支持程度。我们将这一需求形式化为GUI-API对等原则:人类用户通过图形界面可用的每一项能力,也应通过带有适当安全元数据的结构化、可调用接口提供给Agent。为使该原则可操作化,我们提出两项贡献。首先,Agent可调用功能覆盖率(ACFC),一个产品级就绪度指标,量化系统面向人类的功能中Agent能够访问和使用的比例。其次,Agent就绪一致性(ARC),一个框架,从三个维度对每项能力评分:可访问性(Agent能否调用它)、可发现性(Agent能否找到并理解它)和可控性(Agent能否安全使用它),每个维度0-3分,综合得出0-9分的就绪度评分。借鉴Richardson成熟度模型,ARC提供了比二元覆盖率更细粒度的评估。在对5个类别共30个软件系统的实证研究中,基于ARC的ACFC得分显著低于仅考虑可访问性的指标所显示的水平:大多数系统实现了中等程度的API覆盖率,但在面向Agent的文档和安全治理方面滞后。在一个包含10个已部署系统的受控测试平台中,仅使用API的Agent完成了58个任务中的56个,这些任务的目标能力通过结构化接口暴露;文档消融实验将完成数降至58个中的50个,表明即使在可访问的能力中,可发现性也会影响执行,而作为阴性对照的不可访问能力则未被解决。这些发现在三种Agent模型上保持一致,包括在本地运行的Google开放权重Gemma4。

英文摘要

AI agents already operate graphical software through screenshot-based computer use, so the pressing question is not whether agents can operate software, but how well software supports them through structured, controllable channels. We formalize this need as the GUI-API parity principle: every capability available to human users through a graphical interface should also be accessible to agents through a structured, callable interface with appropriate safety metadata. To operationalize this principle, we introduce two contributions. First, Agent-Callable Feature Coverage (ACFC), a product-level readiness metric quantifying how much of a system's human-facing functionality agents can access and use. Second, Agent Readiness Conformance (ARC), a framework that scores each capability on three axes: Accessibility (can an agent call it), Discoverability (can an agent find and understand it), and Controllability (can an agent use it safely), each on a 0-3 scale, yielding a composite 0-9 readiness score. Drawing on the Richardson Maturity Model, ARC provides finer-grained assessment than binary coverage. In an empirical study of 30 software systems across five categories, ARC-based ACFC scores are substantially lower than Accessibility-only metrics would suggest: most systems achieve moderate API coverage but lag on agent-oriented documentation and safety governance. In a controlled testbed of ten deployed systems, API-only agents completed 56 of 58 tasks whose target capabilities were exposed through structured interfaces; a documentation ablation reduced this to 50 of 58, showing that Discoverability affects execution even among accessible capabilities, while inaccessible capabilities, included as negative controls, were not solved. These findings hold across three agent models, including Google's open-weights Gemma4 run locally.

发表机构

  • University of Montana(蒙大拿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑