arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiDiffTIR:面向多轮工具集成推理的分层难度感知策略优化

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Yucan Guo, Xiaohan Wang, Miao Su, Saiping Guan, Zhongni Hou, Jiajun Chai, Wei Lin, Guojun Yin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng

arXiv 2608.21863首次发表:更新:

发表机构

State Key Laboratory of AI Safety; Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Meituan(人工智能安全国家重点实验室; 中国科学院计算技术研究所; 中国科学院大学; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HiDiffTIR是面向多轮工具集成推理的分层难度感知策略优化框架,通过在轨迹和轮次两级执行难度感知信用分配,无需额外监督即可提升工具集成推理性能。

AI 中文摘要

工具集成推理(Tool-Integrated Reasoning,TIR)是大型语言模型(LLM)智能体通过与外部工具迭代交互解决复杂任务的基础能力,强化学习(RL)已成为实现该能力的主流范式。然而现有方法通常为轨迹分配统一的优势值,并将所有正确工具调用同等对待,忽略了不同轨迹和推理步骤间的难度差异与学习价值,导致学习信号不精确,无法充分区分琐碎与具有挑战性的工具使用模式。为解决该局限,本文提出HiDiffTIR,一种面向多轮TIR的分层难度感知策略优化框架。HiDiffTIR在轨迹和轮次两级执行难度感知信用分配,使策略能够聚焦于更具信息量的轨迹和更难的推理步骤。值得注意的是,这种细粒度优化无需额外监督,仅依赖标准RL rollout得到的组级统计数据。在三个工具使用基准上开展的大量实验表明,HiDiffTIR相比强大的RL基线,在多轮TIR性能和工具调用准确率上均实现了持续提升,凸显了难度感知信用分配对工具集成LLM智能体有效策略优化的必要性。

英文摘要

Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, existing approaches typically assign uniform trajectory-level advantages and treat all correct tool calls equally, ignoring the varying difficulty and learning value across trajectories and reasoning steps. This can lead to imprecise learning signals that do not adequately distinguish between trivial and challenging tool-use patterns. To address this limitation, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR. HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels, enabling the policy to focus on more informative trajectories and harder reasoning steps. Notably, this fine-grained optimization is achieved without additional supervision, relying solely on group-level statistics derived from standard RL rollouts. Extensive experiments on three tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.

CommentsAccepted by EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑