AI 中文总结
研究聚焦大型语言模型代理在模型-框架边界处的反应性错误,通过分析255份错误报告构建分类法,发现此类错误多为无明确测试预言的静默错误,检测和重现困难,还揭示了用户与开发者对错误归因及修复的不匹配,推动相关技术设计。
AI 中文摘要
大型语言模型(LLM)代理跨越命令行界面(如Codex)和代理框架(如LangChain),将后端LLM与解析模型输出、控制代理循环和管理上下文的框架代码集成。框架和LLM生成的响应共同塑造代理的执行。这种架构导致了一些仅检查单个组件无法轻易理解的错误,因为有些错误仅在特定LLM响应引发代理异常反应时出现。以往对代理错误的实证研究大多将故障归因于模型能力有限或框架端缺陷,而未对这些代理反应性(AR)错误进行特征描述。我们进行了第一项专注于代理反应性错误的实证研究。通过对来自Codex、Gemini-CLI、LangChain和CrewAI的255份错误报告进行人工分析,我们构建了一个涵盖可观察症状和触发这些症状的LLM行为的双轴分类法。我们的研究结果表明,许多AR错误表现为无明确测试预言的静默错误,这使得检测变得困难。LLM响应的随机性进一步使错误重现复杂化。我们还检查了用户提出并由开发者实施的修复。该分析揭示了一种不匹配:用户经常主张框架端的防护措施,而开发者可能将问题归因于LLM或对用户提出的修复反应缓慢。这些发现表明需要有助于用户和开发者理解AR错误根本原因和解决方案的机制。总体而言,该研究突出了LLM代理特有的挑战,并推动了针对AR错误的测试预言、重现支持和故障定位技术的设计。
英文摘要
LLM agents span command-line interfaces (e.g., Codex) and agent frameworks (e.g., LangChain), integrating backend LLMs with harness code that parses model outputs, controls agent loops, and manages context. Both the harness and LLM-generated responses jointly shape an agent's execution. This architecture gives rise to bugs that cannot be readily understood by inspecting either component alone, because some bugs occur only when a particular LLM response elicits an abnormal reaction from the agent. Prior empirical studies of agent bugs have largely attributed failures either to limited model capabilities or to harness-side defects, such as outdated APIs and configuration misalignment, without characterizing these AR bugs. We conduct the first empirical study focused on agent-reactive (AR) bugs. Through manual analysis of 255 bug reports from Codex, Gemini-CLI, LangChain, and CrewAI, we construct a two-axis taxonomy covering observable symptoms and the LLM behaviors that trigger them. Our findings show that many AR bugs manifest as silent errors without well-defined test oracles, which makes detection difficult. The stochasticity of LLM responses further complicates bug reproduction. We additionally examine fixes proposed by users and implemented by developers. This analysis exposes a mismatch: users frequently advocate harness-side guardrails, whereas developers may attribute the issue to the LLM or respond slowly to user-proposed fixes. These findings point to the need for mechanisms that help users and developers understand the root causes and resolutions of AR bugs. Overall, the study highlights challenges specific to LLM agents and motivates the design of test oracles, reproduction support, and fault-localization techniques for AR bugs.