arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为什么反馈增强的自蒸馏不能改进检索交织搜索智能体?

Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?

Fan Yang, Rui Meng, Yuxin Wen

arXiv 2607.17558首次发表:更新:

发表机构

Chapman University; Lawrence Berkeley National Laboratory(查普曼大学; 劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究反馈增强的自蒸馏在智能体搜索中失效的原因,发现模型存在解码崩溃现象,因监督信号不一致致学习不稳定,将其分解为模型和提示不一致,引入EMA教师减轻不一致,最终提高模型性能。

AI 中文摘要

在线策略自蒸馏(OPSD)为训练大语言模型提供了一种有前景的方法,且无需依赖单独的教师模型。然而,其在复杂智能体任务上的有效性很大程度上未被探索。在这项工作中,我们实例化了反馈增强自蒸馏(FA-SD),一种用于智能体搜索的自蒸馏算法,它利用成功的示范作为特权信息。我们发现模型可能依赖重复的推理和搜索输出模板,产生看似多样但对输入问题基本无感知的轨迹,使得基于KL的自蒸馏信号无信息。我们将此现象称为解码崩溃,这是现有评估指标可能遗漏的一种失败模式。为理解其根本原因,我们表明尽管自教师取得更强性能,但由于监督信号不一致,学习本质上仍不稳定。我们进一步将这种不一致分解为模型不一致和提示不一致,并表明后者会显著降低监督信号质量,限制自教师学习的有效性。为减轻这种不一致,我们引入指数移动平均(EMA)教师来稳定自教师并提供更一致的监督信号。尽管EMA教师需要一个热身阶段,期间性能可能暂时退步,但它最终通过提供更稳定的监督提高了模型性能。

英文摘要

On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored. In this work, we instantiate Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information. We identify that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal uninformative. We term this phenomenon decoding collapse, a failure mode that can be missed by existing evaluation metrics. To understand its underlying cause, we show that although the self-teacher achieves stronger performance, learning remains inherently unstable due to inconsistent supervision signals. We further decompose this inconsistency into model inconsistency and prompt inconsistency, and show that the latter can significantly degrade the quality of the supervision signal, limiting the effectiveness of self-teacher learning. To mitigate this inconsistency, we introduce an exponential moving average (EMA) teacher to stabilize the self-teacher and provide more consistent supervision signals. Although the EMA teacher requires a warm-up phase during which performance may temporarily regress, it ultimately improves model performance by providing more stable supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑