arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14275cs.AIcs.MA

人工智能智能体并非独自失败:上下文首先失败

AI Agents Do Not Fail Alone:The Context Fails First

  • The University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Fouad Bousetouane

AI总结:

研究人工智能智能体可靠性,通过ProofAgent-Harness评估上下文七个标准,验证上下文工程质量是智能体可靠性的领先指标,经受控研究表明上下文质量标准可预测行为结果,确立其为预检信号及可审计层。

AI中文摘要:

上下文工程已成为构建可靠人工智能智能体的核心,但仍基本未被衡量。智能体并非孤立失败,其行为受上下文中积累的指令、工具、记忆、检索知识、护栏和不可信输入影响。当上下文薄弱时,智能体出现漂移、幻觉、滥用工具等问题。本文验证了上下文工程质量是智能体可靠性的独立领先指标。通过ProofAgent-Harness进行测量,该工具基于多陪审员、基于共识的评分评估上下文的七个标准。通过跨监管智能体领域的受控上下文质量研究表明,上下文质量标准能预测相应行为结果。这些发现确立了上下文测量作为智能体可靠性的有效预检信号,并将上下文工程定位为智能体评估和治理的可审计层。

英文摘要:

Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured. Agents do not fail in isolation: their behavior is shaped by the instructions, tools, memory, retrieved knowledge, guardrails, and untrusted inputs accumulated in their context. When this context is weak, agents drift, hallucinate, misuse tools, ignore constraints, become vulnerable to injection, and waste tokens. This paper validates context-engineering quality as an independent leading indicator of agent reliability. We implement the measurement in ProofAgent-Harness, an open-source infrastructure for AI agent evaluation that uses multi-juror, consensus-based scoring. The harness assesses context across seven criteria: role clarity, guardrail coverage, instruction consistency, tool schema quality, grounding sufficiency, injection hardening, and token efficiency. Crucially, the context score is isolated from behavioral metrics and release decisions, enabling a non-circular validation. Through a controlled context-quality study across regulated agent domains, holding frontier LLM agents fixed and varying only their operating context, we show that context-quality criteria consistently predict their corresponding behavioral outcomes. Grounding sufficiency predicts hallucination resistance, guardrail coverage predicts manipulation resistance, instruction consistency predicts instruction following, and tool-schema quality predicts tool use. These findings establish context measurement as a validated preflight signal for agent reliability and position context engineering as an auditable layer of agent evaluation and governance.

补充信息

↑