arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18128cs.AIcs.LO

基于契约的LLM智能体符号时间监督

Symbolic Temporal Supervision of LLM Agents Using Contracts

  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Yifeng Xiao, Pierluigi Nuzzo

AI总结:

针对LLM智能体工具调用的安全监督问题,提出基于LTLf契约的ContrAgent框架,通过确定性自动机同时实现在线门控与离线评估,在四个基准上匹配现有方法且延迟大幅降低。

AI中文摘要:

由工具增强的大型语言模型(LLM)智能体能够通过工具调用在外部系统上执行操作,从而自动化复杂的多步骤任务,例如网页导航、代码生成和工作流编排。然而,LLM中的幻觉、分布不稳定性和对抗性操纵,以及某些工具调用不可逆的后果,可能导致有害结果。现有的安全措施要么事后用随机LLM评判器对记录的轨迹进行评分,要么一次一个调用地阻止不安全操作,且没有单一的确定性工件同时支持这两种角色。我们提出了ContrAgent,一个基于契约的框架,用于对LLM智能体进行符号时间监督。ContrAgent将智能体的行为捕获为一系列工具调用,并将其形式化为对一组固定可检查谓词的轨迹。然后,它使用有限轨迹上的线性时态逻辑(LTLf)中的假设-保证契约来指定所需行为。每个契约被编译成一个确定性有限自动机(DFA),该自动机扮演两个角色:在线门控智能体动作和离线评估记录的轨迹。契约库作为一个可复用的知识库,独立于智能体的模型进行维护,并可应用于同一任务领域内的不同智能体。我们在跨越这两个角色的四个基准上展示了我们方法的有效性,其中ContrAgent匹配了最先进的LLM评判器和基于规则的护栏基线,同时产生确定性、可复现的判定,并且在线模式下,每次调用的延迟降低了几个数量级。

英文摘要:

Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting on external systems through tool calls. However, hallucinations, distributional instability, and adversarial manipulations in LLMs, and the irreversible consequences of certain tool calls can lead to harmful outcomes. Existing safeguards either grade recorded trajectories post hoc with stochastic LLM judges or block unsafe actions one call at a time, and no single deterministic artifact supports both roles. We present ContrAgent, a contract-based framework for symbolic temporal supervision of LLM agents. ContrAgent captures an agent's behavior as a sequence of tool calls and formalizes it as a trace over a fixed set of checkable predicates. It then specifies required behaviors using assume-guarantee contracts in linear temporal logic over finite traces (LTLf). Each contract is compiled to a deterministic finite automaton (DFA) that serves two roles: gating agent actions online and evaluating recorded traces offline. A contract library, acting as a reusable knowledge base, is maintained independently of the agent's model and can be applied across different agents within the same task domain. We show the effectiveness of our approach on four benchmarks spanning both roles, where ContrAgent matches state-of-the-art LLM-judge and rule-based guardrail baselines while producing deterministic, reproducible verdicts and, in the online mode, orders-of-magnitude lower per-call latency.

↑