arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DoGBench:智能体能否达到面向用户文档的专家标准?

DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders

arXiv 2609.39909首次发表:更新:

发表机构

Promptless; Doc Detective; Helm; Mautic; PostHog(Promptless公司; Doc Detective公司; Helm公司; Mautic公司; PostHog公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DoGBench是首个面向用户文档生成基准,含292个开源项目条目,要求智能体判断文档是否需更新并生成可接受补丁,否则弃权;评估七个智能体,最高47.3分,识别出任务完成缺口等主要失败模式。

AI 中文摘要

我们提出了DoGBENCH(文档生成基准),据我们所知,这是首个用于生成和维护真实面向用户软件文档的基准。它考察智能体能否生成经验丰富的技术写作者在评审中会接受的文档。该基准包含来自开源项目的292个条目,包括Helm、PostHog和Mautic。每个条目为智能体提供一个变更前的代码仓库和一个触发器,例如代码拉取请求或报告的文档缺口。智能体必须首先决定文档是否需要更新。对于需要更新的条目,智能体必须一次性生成可接受的补丁。对于不需要更新的条目,智能体必须弃权(不执行)。任务特定的评分标准(经项目维护者验证)根据准确性、完整性、读者指导、位置和仓库约定对每个补丁进行评分。综合得分将补丁质量与正确弃权相结合,得分为100意味着智能体满足了任务的所有要求。得分不应被解释为专家能力的百分比。我们评估了七个智能体。得分最高的智能体在117个条目的保留测试集上达到47.3分(满分100)。在另一次对1,267个补丁的审计中,最常见的失败模式是任务完成缺口(45.5%)、技术不准确(36.6%)以及概念或参考覆盖不完整(32.5%)。对相应轨迹的分析确定了与这些失败相关的三个关键模式:(1)描述接口而不检查读者如何使用它们(36.0%),(2)缺少决定性证据并用看似合理的假设填补空白(33.1%),以及(3)在找到第一个看似合理的文档表面后停止,而让其他受影响的页面保持过时(30.1%)。

英文摘要

We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑