发表机构
Oracle Health AI; MBZUAI(甲骨文健康人工智能; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ContractEval通过将程序指令建模为查询活跃义务并与执行证据匹配,诊断LLM智能体执行中的结构性符合性失败,实现可审计的程序性符合性评估。
AI 中文摘要
随着LLM智能体从回答问题转向执行程序,失败可能是不合理的而非明显错误的:即使系统跳过了使答案合理的检查、分支、依赖或不变量,最终响应看起来仍然可接受。仅基于输出的评估看到答案,而基于轨迹的评判看到活动,但两者都无法识别哪些义务对于该查询是活跃的。我们引入了CONTRACTEVAL,一个诊断框架,用于使这些活跃义务明确化。它将程序性指令表示为查询活跃义务,并将其与响应或轨迹证据进行匹配,将遗漏、错误分支、顺序错误、多余动作、不变量违反和输出契约违反转化为不同的符合性失败。在一组经过审计的程序性契约受控套件上,仅基于输出和基于轨迹的LLM评判器遗漏了许多注入的结构性失败;在黄金预期和观察图下,ContractEval检测并定位了所有失败。基于LLM的提取保留了许多此类信号,但仍对校准敏感。因此,ContractEval不是合规保证;它使程序性符合性可审计,而非隐含在最终答案质量中。
英文摘要
As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions as query-active obligations and matches them against response or trace evidence, turning omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations into distinct conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges miss many injected structural failures; under gold expected and observed graphs, ContractEval detects and localizes all of them. LLM-backed extraction preserves much of this signal but remains calibration-sensitive. ContractEval is therefore not a compliance guarantee; it makes procedural conformance auditable rather than implicit in final-answer quality.