发表机构
University of Liverpool(利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于逻辑、信息片段和保证分类账的统一框架,用于组织语言模型智能体的轨迹证据,通过多种检查方法合并声明并识别残余风险,在τ²-bench基准上验证了有效性。
AI 中文摘要
评估语言模型智能体的方法包括对日志的规则检查器、技能覆盖与组合分析、支持检查器、前缀监控器、执行门控以及罕见事件估计器。每种方法观察运行的不同部分,并作出不同强度的声明,且没有统一的说明来解释这些声明如何组合或它们遗漏了什么。我们给出一个统一的说明,它基于逻辑、信息片段和保证分类账构建。需求是两层逻辑的公式:外层是记录事件上的有限迹时间逻辑,内层是智能体记录上下文在决策时支持的论证结构上的持续性逻辑。因此,规则违反和在撤回支持上采取的行动是同一语言的公式。评估者、智能体和执行门控可用的信息定义了该语言的片段;每种检查方法决定一个片段并返回类型化声明:精确的、在命名测试套件上的、风险界限的或描述性的。分类账合并证据并将每个义务分类为已建立、已处理但未建立或未处理。在发布的456条τ²-bench电信轨迹上,基准预言机标记了231次运行;添加规则检查器、支持替代和前缀监控基线将并集提高到391、401和407,每个都贡献了其他方法遗漏的标记,并且门控使一个禁令在门控部署上精确。49条未标记的运行和未满足或未处理的义务构成残余。发现使未知需求明确,新的或更强的检查器随后减少残余。
英文摘要
Methods for assessing language-model agents include rule checkers over logs, analyses of skill coverage and composition, support checkers, prefix monitors, execution gates, and rare-event estimators. Each observes a different part of a run and makes a claim of a different strength, and no common account says how these claims combine or what they leave unchecked. We give one, built from a logic, information fragments, and an assurance ledger. Requirements are formulas of a two-tier logic: an outer finite-trace temporal logic over recorded events, and an inner logic of standing over the argument structure that the agent's recorded context supports at a decision. A rule violation and an action taken on withdrawn support are thus formulas of one language. The information available to an assessor, to the agent, and to an execution gate defines fragments of that language; each checking method decides one fragment and returns a typed claim: exact, on a named test suite, a risk bound, or descriptive. The ledger merges the evidence and classifies every obligation as established, addressed but not established, or unaddressed. On 456 released $τ^2$-bench telecom trajectories, the benchmark oracle flags 231 runs; adding a rule checker, a support stand-in, and a prefix-monitor baseline raises the union to 391, 401, and 407, each contributing flags the others miss, and a gate makes one prohibition exact on a gated deployment. The 49 unflagged runs and the unmet or unaddressed obligations form the residual. Discovery makes unknown requirements explicit, and new or stronger checkers then reduce it.