arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI代理能否检测并修复网络实验中的工件漂移?

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu

arXiv 2609.09849首次发表:更新:

发表机构

Nokia Bell Labs; Telecom Paris; Nanjing University; University of California, Santa Cruz; The University of Hong Kong; Tsinghua University; Augusta University(诺基亚贝尔实验室; 巴黎电信学院; 南京大学; 加利福尼亚大学圣克鲁兹分校; 香港大学; 清华大学; 奥古斯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出工件完整性概念及NetArtifactBench基准,测试AI代理修复网络实验记录不一致的能力,发现代理在跨工件隐含关系修复上表现不佳,主张将其作为AI代理的核心评估要求。

AI 中文摘要

近年来,AI代理已发展为能够在数字环境中执行多步骤任务的得力助手。网络系统社区正开始在操作和实验环境中探索这些能力。然而,在网络系统中运行的代理不应仅凭其是否完成当前任务来评判。它所修改的实验记录也必须保持可信。我们将这一属性称为工件完整性:记录中的声明必须始终得到现有证据的支持,限定在该证据确立的范围内,并可通过编码其支持的工件进行追溯。为使该属性可度量,我们引入了NetArtifactBench,该基准测试AI代理能否修复源自公共网络系统工件的不一致记录,同时保留仍受支持的声明。该基准包含52个实例,注入了从直接矛盾到跨多个工件未明确表述的关系等各类不一致。我们使用确定性评分,在三个通用AI代理运行时上评估了23种代理配置。在5,980个输出中,平均合同通过率为65.3%,但当修复需要恢复隐含关系并在工件间传播变更时,没有任何代理运行时超过30%。这些结果揭示了局部修正与完整记录级修复之间的明显界限。因此,我们主张工件完整性应成为在网络上运行的AI代理的一等设计和评估要求。

英文摘要

In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in operational and experimental settings. However, an agent operating in network systems should not be judged solely by whether it completes the immediate task. The experiment record it modifies must also remain trustworthy. We call this property artifact integrity: the record's claims must remain supported by the available evidence, confined to the scope established by that evidence, and traceable through the artifacts that encode their support. To make this property measurable, we introduce NetArtifactBench, which tests whether AI agents can repair inconsistent records derived from public network-system artifacts while preserving claims that remain supported. The benchmark contains 52 instances with injected inconsistencies ranging from direct contradictions to unstated relations spread across several artifacts. We evaluate 23 agent configurations across three general-purpose AI agent runtimes using deterministic scoring. The average contract pass rate is 65.3 % across 5,980 outputs, but no agent runtime exceeds 30 % when repair requires recovering implicit relations and propagating changes across artifacts. These results reveal a sharp boundary between local correction and complete record-level repair. Therefore, we argue that artifact integrity should become a first-class design and evaluation requirement for AI agents operating on network systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑