arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31587cs.SEcs.AIcs.CL

编码智能体的紧凑文档:基准、优化器及其不可迁移性

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

Md Shohel Arman, Igor Molybog

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建了往返基准和优化器,发现文档完整性决定保真度,但实验表明文档并不能提升智能体解决真实仓库问题的能力。

中文摘要 AI 辅助

我们研究了自然语言文档是否有助于编码智能体解决软件问题,并构建了用于构建和评估这些文档的工具。我们引入了一个往返基准,该基准通过从代码描述重新生成的代码是否通过原始测试来评分,并表明完整性而非长度决定描述的保真度。利用该基准作为优化信号,我们发现了一个能达到完全保真度并泛化到未见文件的描述编写提示。随后,我们检验了推动这项工作的假设:更好的文档有助于智能体解决真实仓库中的问题。跨越两个模型家族和十个仓库,并以一个确认我们的评估能检测到真实改进的阳性对照为参照,我们发现情况并非如此。当源代码存在时,无论是静态紧凑文档还是检索到的上下文,都不比仅凭问题本身更有效。我们报告了这一负面结果,连同基准和优化器,并刻画了文档发挥作用的边界。

英文摘要

We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.

发表机构

  • Daffodil International University(达福迪尔国际大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑