arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

依赖感知的轨迹精炼用于高效多轮智能体微调

Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

Zhuo Chen, Zhen Zhang, Xinyu Wang, Kewei Tu

arXiv 2609.18417首次发表:更新:

发表机构

ShanghaiTech University; Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海科技大学; 上海智能视觉与成像工程技术研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多轮智能体轨迹中的冗余轮次,提出基于轮次级依赖DAG的轨迹精炼方法,在四个多模态问答基准上提升准确率并大幅降低推理成本。

AI 中文摘要

多轮智能体轨迹通常包含冗余轮次(失败的工具调用、并行的子查询、仅验证的步骤),这些冗余轮次增加了训练和推理成本。我们提出将每条轨迹视为一个轮次级依赖有向无环图(DAG),该图揭示了哪些轮次对最终答案具有全局关键作用,并在此DAG精炼后的轨迹上对智能体进行微调。给定由大语言模型(LLM)标注的DAG,这些编辑是确定性的且可解释的,并可选地进行改写。在这些精炼轨迹上训练的模型在较低推理成本下持续优于在原始轨迹上训练的模型。具体而言,在四个多模态问答基准上,我们的精炼方法相比普通监督微调(SFT)将下游准确率提升最多1.7个百分点(相比LLM删除基线提升5.7个百分点),同时将每个样本的推理消息数减少最多约40%,推理令牌数减少最多约48%,从而大幅节省计算和服务器成本。代码已公开。

英文摘要

Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a \emph{round-level dependency DAG} that exposes which rounds are globally load-bearing for the final answer, and fine-tune agents on trajectories refined through this DAG. Given an LLM-annotated DAG, these edits are deterministic and interpretable, with optional rephrasing. Models trained on these refined trajectories consistently outperform those trained on the original trajectories at lower inference cost. Specifically, across four multi-modal QA benchmarks, our refinements improve downstream accuracy by up to $1.7$\,pp over vanilla SFT (and $5.7$\,pp over an LLM-deletion baseline) while reducing per-sample inference messages by up to approximately $40\%$ and inference tokens by up to approximately $48\%$, translating to substantial savings in compute and serving cost. Code is available.

CommentsAACL 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑