上下文文件对编码智能体有帮助吗?针对真实仓库的双智能体消融研究
Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
浏览论文内容
中文总结 AI 辅助
本研究通过双智能体(Claude Code、Codex)的受控消融实验,发现上下文文件对编码智能体的正确性无显著提升,其失败源于实现技能而非知识缺失,还解释了过往研究矛盾的原因并公开了相关资源。
中文摘要 AI 辅助
持久上下文文件(this http URL)是指导AI编码智能体的标准做法,但其有效性证据相互矛盾。我们对两种前沿智能体(Claude Code和Codex)在3个仓库的17项真实任务(15项共享任务+2项Codex专属任务)上的上下文注入策略进行了受控消融研究,共完成288次带黄金测试评估的运行。等价检验显示,上下文策略对任一智能体的正确性均无显著提升(差异上限≤10-15个百分点)。失败模式分类揭示了原因:智能体的失败源于实现技能不足——包括功能设计、模式选择、精确衔接,而非上下文文件可提供的缺失仓库知识;操纵探针证实,真实this http URL从未在任一智能体上将接近成功的任务转为通过。我们还表明,边界任务难度具有智能体特异性(斯皮尔曼相关系数rho=0.75),这为先前的矛盾提供了候选解释:单智能体研究从不同智能体的信息范围内选取任务。我们发布了所有代码、数据和分析。
英文摘要
Persistent context files (AGENTS.md, CLAUDE.md) are standard practice for guiding AI coding agents, yet evidence for their effectiveness is contradictory. We present a controlled ablation of context-injection strategy across two frontier agents (Claude Code and Codex), 17 real tasks from 3 repositories (15 shared + 2 Codex-only), and 288 evaluated runs with gold-test evaluation. Context strategy does not measurably move correctness on either agent (bounded to <=10-15pp via equivalence testing). A failure-mode triage reveals why: agents fail on implementation skill---feature design, pattern selection, exact wiring---not missing repository knowledge that a context file could supply; a manipulation probe confirms the real AGENTS.md never converts a near-miss to a pass on either agent. We further show that borderline task difficulty is agent-specific (Spearman rho=0.75), offering a candidate explanation for prior contradictions: single-agent studies draw tasks from different agents' informative bands. We release all code, data, and analysis.