arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30147cs.CL

CAST:用于训练可靠长程工具调用智能体的批判感知监督

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

Amir Saeidi, Zehua Zhang, Rishitosh Singh, Naman Ahuja, Vivek Gupta, Ali Payani, Gaowen Liu, Jayanth Srinivasa, Chitta Baral

首次发表
浏览论文内容

中文总结 AI 辅助

提出批判感知训练框架CAST,通过分析智能体轨迹生成结构化批判,优化Qwen3模型,在零售、远程医疗等动态工具调用任务上提升可靠性,性能优于GPT-OSS-120B。

中文摘要 AI 辅助

大型语言模型(LLM)智能体正越来越多地部署在长程、交互式且有状态的环境中。在这些环境中,单个错误动作(如错误退款)可能导致不可逆的任务失败,必须在执行前拦截。此类失败并非每次运行都会出现,但会在多次试验中显现,因此跨步骤和跨试验的可靠性至关重要。然而,确保智能体可靠性颇具挑战性:即使是前沿LLM也难以解释动作为何可能错误,尤其是在受领域特定策略约束的冗长、交织轨迹中。近期许多工作依赖基于提示的批判智能体,而基于优化的方法缺乏生成丰富验证理由用于训练的系统方式。我们通过CAST解决这一缺口,这是一种批判感知训练框架,可将稀疏任务结果转化为动作级监督,用于批判学习和策略优化。CAST分析智能体轨迹,合成部分可观测性下解释动作有效性的结构化理由。生成的批判模型用于构建批判感知训练数据,以优化策略模型。在动态工具调用基准上微调Qwen3系列模型时,CAST提升了跨领域的可靠性,在零售任务上的通过率比GPT-OSS-120B高出10%以上,在跨域场景下的远程医疗任务上进一步提升9%。这些结果表明,批判感知训练可提升LLM智能体在现实动态环境中的鲁棒性。

英文摘要

Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the wrong purchase, can cause irreversible task failure and must be intercepted before execution. Such failures may not appear in every single run, but can emerge across repeated trials, making reliability across steps and trials critical. However, ensuring agentic reliability is challenging: even frontier LLMs struggle to explain why an action may be wrong, especially in long, intertwined trajectories governed by domain-specific policies. Much recent work relies on prompt-based critique agents, while optimization-based methods lack a systematic way to produce rich verification rationales for training. We address this gap with CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization. CAST analyzes agent trajectories to synthesize structured rationales explaining action validity under partial observability. The resulting critique model is used to construct critique-aware training data for optimizing the policy model. Fine-tuning Qwen3-family models on dynamic tool-calling benchmarks, CAST improves reliability across domains, outperforming GPT-OSS-120B by over 10% pass^4 on Retail tasks and yielding an additional 9% improvement on Telehealth in an out-of-domain setting. These results demonstrate that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.

发表机构

  • Arizona State University(亚利桑那州立大学)
  • Cisco Research(思科研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑