arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23610cs.SE

从可追溯性到可辩护性:智能体软件工程中的问责结构

From Traceability to Justifiability: Accountability Structures in Agentic Software Engineering

Rashid Azarang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过对47个平台的文档调查和30个仓库的工具分析,发现智能体软件工程领域的记录缺乏可辩护性,存在可验证性漏洞,仅少数采用证明工具的仓库实现了端到端绑定。

中文摘要 AI 辅助

AI系统的流水线会发布记录,声称所评估的内容就是已部署的内容,且证据支持该转变。我们仅从公开材料中测量这些记录能否表达该主张,以及该主张在声明处是否成立。首先,采用固定的三标签协议对47个交付平台(20个CI/CD平台、27个模型服务/智能体平台)进行两类文档调查,分两次评分(第二次为盲评),每个查阅页面均按内容哈希和日期固定。在188个双评分单元中,我们发现没有任何平台的默认记录会输出行为元组(模型版本、指令、工具定义、检索及运行时配置)的内容寻址标识;盲评对该列的默认项评分在47个平台中为0。与此同时,不可变的名义版本控制正作为智能体平台的默认方案出现(27个平台中的16个):可变指针后的版本整数,而制品供应链已发现该层存在不足。其次,开发了一种工具,仅从流水线发布的内容中计算已实现的保证深度,并将其与声明的深度进行比较。该工具应用于由30个公开仓库组成的冻结两层框架,从哈希归档中进行两次评分(第二次为盲评;单元级一致性分别为30个中的23个和19个,两次独立发现了相同的5个完全实现案例),最显著的结果是存在可验证性漏洞:在15个采用证明工具的仓库中,有7个仅发布了源代码版本,因此其工作流声明的绑定无法在声明处检查。在可检查的情况下,大部分都能通过:7个可测量的采用者中有5个实现了端到端绑定;两个不足均出现在标识绑定环节。综合结果表明,该领域的记录在结构上缺乏可辩护性,而可辩护性是能够拒绝转变的关键层级。该调查带有失效时钟;我们说明了可证伪每项发现的条件。

英文摘要

A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition. We measure, from public material only, whether those records can express that claim and whether it holds where declared. First, a two-class documentation survey of 47 delivery platforms (20 CI/CD, 27 model-serving/agent) under one fixed three-label protocol, graded twice (second pass blind), every consulted page pinned by content hash and date. Across 188 double-graded cells we found no platform whose default record emits a content-addressed identity of the behavioral tuple (model version, instructions, tool definitions, retrieval and runtime configuration); the blind pass grades that column default on zero of 47. Immutable nominal versioning is meanwhile arriving as the agent platforms' default answer (16 of 27): version integers behind mutable pointers, a layer the artifact supply chain already found insufficient. Second, an instrument computes realized assurance depth from a pipeline's published exhaust alone and compares it with the declared depth. Applied to a frozen two-stratum frame of 30 public repositories graded twice from a hashed archive (second pass blind; cell-level agreement 23 and 19 of 30, both passes independently finding the same five full realizations), the sharpest result is a verifiability hole: seven of the 15 repositories chosen for adopting attestation tooling publish source-only releases, so the binding their workflows declare cannot be checked where declared. Where checkable it mostly checks out: five of seven measurable adopters realize the binding end to end; both shortfalls fall at identity binding. Together the results locate the field's records structurally short of justifiability, the one rung that can refuse a transition. The survey carries an expiry clock; we state what would falsify each finding.

补充信息

↑