arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04991cs.AI

CIPO:用于自适应工具粒度选择的反事实想象策略优化

CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection

Yu Li, Yunlu Wan, Zijian Zhu, Han Luo, Chao Ren, Long-Fei Li, Lei Feng

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体工具粒度选择问题,提出CIPO框架,通过反事实分支回滚训练粒度决策,提升任务成功率与决策效率。

中文摘要 AI 辅助

大型语言模型(LLM)智能体通过与外部工具的多步交互来解决复杂任务。这些交互通常包含重复出现的局部工具序列。将此类序列视为复合“技能”可以缩短工具使用轨迹,并减少重复的低层级决策。然而,当原子工具和复合技能共存时,技能使用就成为一个策略问题:智能体必须判断当前状态是需要原子级的精细控制还是技能级的抽象。在本文中,我们认为有效的技能使用应被研究为自适应工具粒度选择。解决这一问题最直接的训练信号是比较同一状态下可用的原子选择和技能选择所带来的后果。基于这一观点,我们提出了CIPO,一个用于自适应工具粒度的反事实想象策略优化框架。CIPO通过对成功的工具使用轨迹进行预算受限的挖掘来构建可执行的技能,并通过反事实分支回滚来训练粒度决策。对于每个基础回滚,CIPO在第一个符合条件的粒度决策点进行分支,并将所选动作替换为可行的原子或技能替代方案。配对的结果差异作为策略优化的补充奖励。在多个基准和模型骨干上的实验表明,CIPO在任务成功率和决策效率上优于基线。进一步的分析表明,CIPO通过根据当前状态改进原子工具与复合技能之间的选择来学习有效的技能使用,而不仅仅是提高技能使用频率。

英文摘要

Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fine control or skill-level abstraction. In this paper, we argue that effective skill use should be studied as adaptive tool granularity selection. The most direct training signal for this problem is to compare the consequences of atomic and skill choices available from the same state. Based on this view, we propose CIPO, a Counterfactual Imagination Policy Optimization framework for adaptive tool granularity. CIPO constructs executable skills through budget-constrained mining of successful tool-use trajectories and trains granularity decisions with counterfactual branch rollouts. For each base rollout, CIPO branches at the first eligible granularity decision and replaces the chosen action with a feasible atomic or skill alternative. The paired outcome difference serves as a supplementary reward for policy optimization. Experiments across multiple benchmarks and model backbones show that CIPO improves task success and decision efficiency over baselines. Further analyses show that CIPO learns effective skill use by improving the choice between atomic tools and composite skills based on the current state, without simply increasing skill frequency.

发表机构

  • Southeast University(东南大学)
  • KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑