arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当智能体实现系统:缺陷、检测与评估严谨性的案例研究

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

Phanindra Reddy Madduru

arXiv 2609.01985首次发表:更新:

AI 中文总结

本研究通过案例研究分析LLM编码智能体实现多组件数据系统的缺陷,在HotpotQA基准上发现过滤检索比未过滤检索性能更优,验证了相关检索权衡的有效性。

AI 中文摘要

随着大语言模型(LLM)编码智能体越来越多地执行端到端工程工作,我们缺乏对它们在系统级需求上的行为的实证表征,这些需求包括模式设计、异步编排、配置正确性以及检索-过滤权衡。我们对一个此类智能体基于预先存在的详细规范实现多组件数据系统的过程开展案例研究。存储技术、模式、实体消歧算法和检索-过滤策略均预先固定;智能体的自主性体现在实现过程中、对自身引入缺陷的诊断与修复,以及未明确的交互设计选择上。在单个会话中,我们分类整理出5个此类缺陷,按违反的约束和检测方法划分。我们进一步在公开的HotpotQA基准上评估该架构中指定的一项检索权衡:在排序前将候选对象限制为图识别的实体集,与未过滤搜索的对比。由于我们无法访问LLM来运行实体识别阶段,因此我们用基准的黄金证据标签替代实体识别,并报告标准召回率而非基准自身的准确率指标。在1至10的检索预算、2994个段落组成的合并语料库的100个问题上,过滤后的召回率在预算为3时达到上限,这符合候选对象被限制为黄金段落本身的预期;而未过滤搜索即使在预算为10时,也仅能69%的时间恢复所有所需证据,在所有测试的预算下均存在这一差距,符号检验的p值小于0.0001。最后,我们讨论智能体自主性在哪些方面取得成功、哪些方面需要修正,包括一个声称的性能修复未在其针对的回归问题上重新测量的实例。

英文摘要

As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-offs. We present a case study of one such agent implementing a multi-component data system against a detailed pre-existing specification. Storage technologies, schema, entity-resolution algorithm, and retrieval-filtering strategy were fixed in advance; the agent autonomy was in the implementation, in diagnosing and fixing defects it introduced, and in interaction-design choices left open. Over a single session, we catalog five such defects, categorized by constraint violated and detection method. We further evaluate, on the public HotpotQA benchmark, the one retrieval trade-off specified in that architecture: restricting candidates to a graph-identified entity set before ranking versus unfiltered search. We substitute the benchmark gold evidence labels for entity identification, since we lacked LLM access to run that stage, and report standard recall rather than the benchmark own accuracy metrics. Across retrieval budgets from 1 to 10 and 100 questions against a pooled corpus of 2994 paragraphs, filtered recall reaches its ceiling by a budget of 3, expected once candidates are restricted to the gold paragraphs themselves, while unfiltered search recovers all required evidence only 69 percent of the time even at a budget of 10, a gap that holds at every budget tested, with sign test p less than 0.0001. We close with a discussion of where the agent autonomy succeeded versus required correction, including one instance where a claimed performance fix was never re-measured on the regression that motivated it.

Comments4 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑