arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越文本:验证智能体撰写的论文是否得到其工件的支持

Beyond the Text: Verifying That Agent-Written Papers Are Backed by Their Artifacts

Qiuhong Shen, Benlong Wu, Hanjin Liu, Yuang Qi, Kejiang Chen

arXiv 2609.22111首次发表:更新:

AI 中文总结

针对智能体撰写论文与其代码工件间缺乏一致性验证的问题,提出自动化审计框架ReAgent,结合静态与动态分析识别不一致,实验证明其有效性。

AI 中文摘要

大型语言模型智能体越来越有能力自主进行研究,在生成研究文档的同时,还生成表面上支持这些文档的代码和实验。然而,所报告的研究结果是否始终得到相应实现和执行证据的支持,在很大程度上仍未得到探索:现有的评审实践主要评估文本质量,无法可靠地识别诸如硬编码指标、未实现的方法或不支持的实验结果等不一致之处。我们提出了ReAgent,一个用于评估智能体生成的研究文档与其相关代码库之间一致性的自动化审计框架。ReAgent从研究文档中构建科学主张的结构化表示,并利用这些表示来指导代码库分析和证据收集。静态审计检查所声称的方法、实现和实验配置是否在代码库中得到一致反映,而动态审计则执行相关实验并收集执行证据以评估实证结果。通过将静态分析与动态证据相结合,ReAgent能够识别出在任一视角单独下可能隐藏的不一致之处,例如那些复现了报告数字但偏离了所声称方法的实验。收集到的证据和审计决策被组织成结构化的代码库级审计报告,从而实现透明的证据可追溯性。我们在一个手工策划的智能体生成研究文档-代码库对基准上评估了ReAgent,并将其与代表性的静态基线和基于复现的基线进行了比较。实验结果表明,ReAgent能够有效识别所报告的研究发现与其支持性代码库证据之间的不一致之处。

英文摘要

Large language model agents are increasingly capable of conducting research autonomously, producing research documents alongside the code and experiments that ostensibly support them. Yet whether the reported findings are consistently supported by corresponding implementations and execution evidence remains largely unexplored: existing review practices primarily assess textual quality and cannot reliably identify inconsistencies such as hard-coded metrics, unimplemented methods, or unsupported experimental results. We present ReAgent, an automated auditing framework for assessing the consistency between agent-generated research documents and their associated repositories. ReAgent constructs structured representations of scientific claims from research documents and uses them to guide repository analysis and evidence collection. Static auditing examines whether claimed methodologies, implementations, and experimental configurations are consistently reflected in the repository, while dynamic auditing executes relevant experiments and collects execution evidence to assess empirical findings. By combining static analysis with dynamic evidence, ReAgent identifies inconsistencies that may remain hidden under either perspective alone, such as experiments that reproduce reported numbers while deviating from the claimed methodology. The collected evidence and audit decisions are organized into a structured repository-level audit report, enabling transparent evidence traceability. We evaluate ReAgent on a manually curated benchmark of agent-generated research document--repository pairs and compare it against representative static and reproduction-based baselines. Experimental results demonstrate that ReAgent effectively identifies inconsistencies between reported research findings and their supporting repository evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑