AI 中文总结
本研究针对AI编码智能体的可靠性问题,整合多源证据构建系统级评估与运行框架,区分模型与基础设施效应,提出可靠性记录目录及相关方法以提升智能体系统可靠性。
AI 中文摘要
AI编码智能体通常被作为模型进行评估,却作为系统部署,其可靠性不仅取决于模型能力,还依赖于测试框架(harness)、执行状态、检索机制、内存与状态管理、权限控制、审核接口以及资源分配。本专著研究了这些边界并开发了一套用于可靠评估和运行编码智能体的框架。研究通过结构化多视角综述、针对性更新审计、软件工程覆盖分析以及分布式系统证据综合,整合了164篇学术文献、100份从业者记录、29份基准记录和17份作者-系统案例记录。在所有证据中,许多表面上的模型故障实际源于系统的其他部分,而某一层的改进往往无法传递到端到端结果。评估和运行被视为一条依赖链,其中任务构建、执行环境、检索、状态管理、验证或可观测性方面的薄弱环节可能会使下游结论失效。本专著贡献了一个包含206条可靠性记录的版本化目录:193项门控实践(其中56项经过深入开发)以及13个研究方向;一个证据总账;一个用于分析智能体生命周期内依赖关系和修复不对称性的框架;来自已运行智能体系统的测量数据和故障案例;可运行的评估和可靠性协议;以及5项带有证据图谱的可复用智能体技能。这些内容共同提供了一套系统级方法,用于区分模型能力与基础设施效应、设计可辩护的评估方案,以及构建在组件故障时能安全恢复的系统。该综述具有结构性而非穷尽性,证据强度因主题而异,结果取决于工作负载和配置。该方法记录了已执行的搜索路径、未执行的路径以及证据分级主张的限制。
英文摘要
AI coding agents are commonly evaluated as models but deployed as systems. Their reliability depends not only on model capability, but on the harness, execution state, retrieval, memory and state management, permissions, review interfaces, and resource allocation. This monograph examines those boundaries and develops a framework for evaluating and operating coding agents reliably. It synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records, and 17 author-system case records through a structured multivocal review, targeted update audits, software-engineering coverage analysis, and distributed-systems evidence synthesis. Across this evidence, many apparent model failures originate elsewhere in the system, while improvements at one layer often fail to propagate to end-to-end outcomes. Evaluation and operation are treated as a dependency chain in which weaknesses in task construction, execution environments, retrieval, state management, verification, or observability can invalidate downstream conclusions. The monograph contributes a versioned catalog of 206 reliability records: 193 gated practices, including 56 developed in depth, plus 13 research leads; an evidence ledger; a framework for dependency and repair asymmetry across the agent lifecycle; measurements and failure cases from operated agent systems; runnable evaluation and reliability protocols; and five reusable agent skills with evidence maps. Together, these provide a system-level methodology for distinguishing model capability from infrastructure effects, designing defensible evaluations, and building systems that recover safely when components fail. The review is structured rather than exhaustive, evidence strength varies by topic, and results depend on workload and configuration. The methods record which search lanes were executed, which remain unexecuted, and limits on evidence-grading claims.
CommentsTechnical review and engineering monograph, 314 pages, 30 figures. Includes an evidence audit, a companion research artifact with 206 reliability records, and runnable protocols for evaluating and operating AI coding agents. August 2026. Source, companion, and reusable protocols: https://github.com/sjarmak/engineering-reliable-coding-agents