arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DART-SD:面向多轮工具调用智能体自蒸馏的钻石拓扑感知检索与微调

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

Hangrui Xu, Jiarui Wang, Yang Yang, Chuanbo Zhu, Fangda Chen, Ziqi Wu, Jingming Cai, Yan Song

arXiv 2608.18524首次发表:更新:

发表机构

ByteDance; University of Science and Technology of China(字节跳动; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对多轮工具调用智能体训练中的拓扑崩溃问题,提出DART-SD框架,通过建模拓扑、检索恢复参考与渐进自蒸馏提升策略多样性,性能优于传统全轨迹基线。

AI 中文摘要

为大型语言模型(LLM)配备多轮工具调用能力,是构建自主智能体的核心需求。但现有进展受限于对全轨迹模仿的依赖:在涉及多个顺序无关子目标的任务中,最优解空间形成庞大的组合钻石格,将这种丰富拓扑强行纳入单一轨迹会引发严重的拓扑崩溃,无差别惩罚有效替代探索,大幅降低策略多样性。为解决该问题,本文提出DART-SD(Diamond-topology Aware Retrieval and Tuning for Self-Distillation,即面向多轮工具调用智能体自蒸馏的钻石拓扑感知检索与微调),该框架将范式从全局强制转向拓扑引导的局部修正。DART-SD首先将执行过程建模为收敛的交互状态转移图(ISTG),准确捕捉成功与失败探索路径的固有钻石拓扑;在自主回滚过程中,框架识别关键拓扑断点(CTB)并检索成功支撑的恢复参考;最后引入基于CTB引导局部监督的渐进自蒸馏范式,确保训练损失仅在生成的恢复步骤上计算,同时严格保护有效推理前缀免受破坏性梯度更新。在复杂多轮工具调用基准上的实验表明,DART-SD显著优于传统全轨迹基线。

英文摘要

Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks involving multiple order-independent sub-goals, the optimal solution space forms a vast combinatorial diamond lattice. Forcing this rich topology into monolithic trajectories causes a severe topological collapse, indiscriminately penalizing valid alternative explorations and severely degrading policy diversity. To address this, we propose DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction. DART-SD first models the execution process as a converging Interaction-State Transition Graph (ISTG), faithfully capturing the inherent diamond topology of successful and failed exploratory paths. During autonomous rollouts, the framework identifies the Critical Topological Breakpoint (CTB) and retrieves success-supported recovery references. Finally, we introduce a progressive self-distillation paradigm through CTB-guided localized supervision, ensuring that the training loss is calculated exclusively on the generated recovery steps while strictly protecting the valid reasoning prefix from destructive gradient updates. Experiments on complex multi-turn tool-calling benchmarks demonstrate that DART-SD significantly outperforms traditional full-trajectory baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑