arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经验证的工具调用可提升非原子故障下LLM智能体的可靠性

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

Isham Kalappurackal Mansoor, Abhishek Phadke, Pratip Rana

arXiv 2608.02645首次发表:更新:

发表机构

Old Dominion University; Christopher Newport University(奥多明尼昂大学; 克里斯托弗·纽波特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非原子故障导致LLM智能体可靠性不足的问题,提出带后置条件验证、重试前逻辑和幂等键的工具包装器,可减少重复动作并保持任务成功率,且无需修改底层LLM。

AI 中文摘要

大型语言模型(LLM)智能体依赖外部工具执行多阶段任务。现有智能体框架通常假设工具调用是原子的,仅返回成功或失败的二元信号,但现实系统存在非原子行为,如调度后超时、可见性延迟、部分状态更新等。这些不匹配会导致重复动作、任务失败、不必要的工具执行等可靠性问题。本文提出一种轻量级、感知验证的工具包装器,为工具调用添加后置条件验证、重试前验证逻辑和幂等键。该方法在注入非原子故障的受控模拟环境中,针对多个任务模板进行评估,结果表明,所提方法可显著减少重复动作,同时保持相当的任务成功率。总体而言,研究结果表明,强化工具交互语义是提升LLM智能体可靠性的有前景方向,且无需修改底层语言模型。

英文摘要

Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are atomic and return binary success or failure signals. However, real-world systems exhibit non-atomic behaviors such as timeouts after dispatch, delayed visibility, and partial state updates. These mismatches lead to reliability issues including duplicate actions, task success, and unnecessary tool executions. A lightweight, verification-aware tool wrapper is introduced that augments tool calls with postcondition verification, verify-before-retry logic, and idempotency keys. The approach is evaluated in a controlled simulated environment with injected non-atomic failures across multiple task templates. The results demonstrate that the proposed method significantly reduces duplicate actions, while maintaining comparable task success rates. Overall, the findings suggest that strengthening tool interaction semantics is a promising direction for improving LLM agent reliability without requiring modifications to the underlying language model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑