arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从自然语言策略到可执行义务:面向可靠车载大语言模型智能体的验证工具

From Natural Language Policies to Executable Obligations: A Verification Harness for Dependable In-Car LLM Agents

Radouane Bouchekir, Damir Safin, Tomas Bueno Momcilovic

arXiv 2608.23282首次发表:更新:

AI 中文总结

本研究提出AgentGuardUtil工具,将车载LLM智能体纳入验证-修正循环,通过运行时策略编译器和25个确定性门等机制,保障智能体满足车载操作策略,提升任务可靠性。

AI 中文摘要

部署在车辆中的大语言模型(LLM)智能体必须在每一轮都满足书面操作策略:单个幻觉标识符、遗漏的强制性副作用或过早的完成声明都会导致任务失败。我们展示了AgentGuardUtil——我们在CAR-bench Track 1的参赛作品,它将AI规划器(LLM)视为基于验证-修正循环中的易出错提议者。其核心创新是运行时策略编译器:每次对话附带的自然语言策略会针对每个策略编译一次,转换为类型化的机器可检查规则,其中一部分会被转换为可执行形式。确定性义务引擎会根据实时工具结果和草稿自身的模拟写入后状态来解释这些规则,发出带有计算参数的精确补救调用,而非自然语言提醒。围绕该引擎,25个确定性门(标识符来源、模式和枚举有效性、先收集后执行、确认和未来时间协议)以及一个LLM评论器会生成分层结果,驱动针对pass@k指标调优的有界修正循环。

英文摘要

Large Language Models (LLMs) agents deployed in vehicles must satisfy a written operating policy on every turn: a single hallucinated identifier, omitted mandatory side-effect, or premature completion claim fails the task. We present AgentGuardUtil, our entry to CAR-bench Track~1, which treats the AI planer (LLM) as a fallible proposer inside a grounded verify-and-revise loop. Its core novelty is a runtime policy compiler: the natural-language policy shipped with each conversation is compiled, once per policy, into typed machine-checkable rules, a subset of which receive an executable form. A deterministic obligation engine interprets these rules against live tool results and the simulated post-write state of the draft itself, emitting the exact remedial calls with computed arguments rather than natural-language reminders. Around this engine, 25 deterministic gates (identifier provenance, schema and enum validity, gather-before-act, confirmation and future-time protocols) and an LLM critic produce tiered findings that drive a bounded revision loop tuned for the pass k metric.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑