arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工具调用是否按预期执行?测量与修复LLM智能体中的意图-执行一致性

Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

Boyang Yang, Zhenhao Li, Ziyao Yang, Kanghui Jia, Xin Yin, Mingmou Liu, Haoye Tian

arXiv 2610.04375首次发表:更新:

发表机构

Yanshan University; Zhejiang University; Nanjing University; Aalto University(燕山大学; 浙江大学; 南京大学; 阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文定义意图-执行一致性(IEC),提出协议观察跳点接收内容并构建IntAct修复工具,在真实基准上证明路径改变调用导致大量失败,IntAct可恢复79.2%的失败。

AI 中文摘要

基于大型语言模型(LLM)的智能体通过工具调用来构建和运行软件。一次调用要经过多个跳点才能到达其程序,而任何跳点都可能在未通知的情况下改变该调用。当被改变的调用失败时,智能体会重试正确的调用,这会耗费用户的时间和金钱。基准测试和失败分析看不到这种改变,因为它们读取的是调用及其结果,而不是某个跳点实际接收到的内容。我们将意图-执行一致性(IEC)定义为:所执行的动作与发出的调用在工具契约下所表示的动作相匹配这一属性。我们的协议在不执行调用的情况下观察每个跳点接收到的内容,并通过接收方自身的解析器来命名第一个改变该调用的跳点。随后,IntAct以该跳点无法改变的形式传递调用,或者拒绝该调用。我们根据真实世界使用中观察到的改变构建了IEC-Bench,其中包含在4个广泛使用的执行框架的执行路径下的依赖调用链。在生产会话中的47,828次shell调用中,Claude Code的Bash工具改变了12.0%携带代码、转义序列或长文本的调用。对于反斜杠被改变的调用,80.7%的情况下错误动作在没有任何报错的情况下执行。所有10个被测量的执行框架都会改变调用。基于轨迹的判断将95.1%的生产失败归因于LLM,尽管路径导致了其中超过一半的失败。在IEC-Bench上,路径使每个通过任务的令牌成本提高了2.4倍(最高达12.3倍)。一个改变调用的跳点也会隐藏其之后的改变,因此一条路径上55.1%的失败仅在第一个跳点被修复后才显现。IntAct已部署在商业产品中,可恢复79.2%的带有被改变调用的失败。因此,执行框架应逐跳设计和测试,以确保正确的调用按预期执行或被拒绝。

英文摘要

Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not see the change, because they read the call and its result but not what a hop received. We define intent-execution correspondence (IEC) as the property that the executed action matches the action the emitted call denotes under the tool contract. Our protocol observes what each hop received without executing the call, and names the first hop that changed it by the receiver's own parser. IntAct then delivers the call in a form that this hop cannot alter, or refuses the call. We build IEC-Bench from the changes observed in real-world use, with chains of dependent calls under the execution paths of 4 widely-used harnesses. In 47,828 shell calls within production sessions, Claude Code's Bash tool changes 12.0% of the calls that carry code, escape sequences, or long text. For 80.7% of the calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change a call. Trajectory-based judgment attributes 95.1% of the production failures to the LLM, although the path caused more than half of them. On IEC-Bench, the path raises the token cost per passed task 2.4 times (up to 12.3 times). A hop that changes a call also hides the changes after it, so 55.1% of the failures on one path appear only after its first hop is repaired. IntAct, deployed in a commercial product, recovers 79.2% of the failures with a changed call. Harnesses should therefore be designed and tested hop-by-hop to ensure a correct call executes as intended or is refused.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑