arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02101cs.CL

面向通用搜索智能体的跨域混合在线策略蒸馏

Cross-Domain Hybrid OPD for Generalizable Search Agents

Hongzhan Chen, Xiaoyu Liu, Dengming Zhang, Minzhou Huang, Dongliang Xu, Jingcheng Xie, Dongxiang Fang, Bowen Qin, Minsheng Hao, Yaozong Shen, Xiaojun Quan, Mona… 展开作者

Hongzhan Chen, Xiaoyu Liu, Dengming Zhang, Minzhou Huang, Dongliang Xu, Jingcheng Xie, Dongxiang Fang, Bowen Qin, Minsheng Hao, Yaozong Shen, Xiaojun Quan, Mona Zhou, Haosheng Zou, Jeff Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于混元3架构的元宝搜索智能体训练框架,通过跨域混合在线策略蒸馏联合优化搜索专业化与通用能力,缓解对齐损耗,在现实搜索场景中实现两者的良好平衡。

中文摘要 AI 辅助

强化学习(RL)的近期进展大幅提升了自主搜索智能体的能力,使其能够在动态信息源上进行复杂规划与迭代检索。然而,针对特定搜索行为优化语言模型常产生对齐损耗,即搜索性能的提升以通用能力为代价,限制了其作为通用助手的有效性。本技术报告介绍了元宝搜索智能体背后的训练框架,旨在实现搜索专业化的同时不牺牲通用智能。该框架基于混元3(Hunyuan3)架构构建,将用于自主搜索的智能体强化学习与跨域专家在线策略蒸馏(OPD)流水线相结合,将专注于互补通用领域的专家蒸馏为搜索专业化的学生模型,恢复并进一步增强其广泛能力。我们的混合训练策略未将专业化与通用能力视为相互竞争的目标,而是联合优化两者,有效缓解了对齐损耗。大量实验表明,所得模型在实现有竞争力的搜索性能的同时,持续提升了通用能力,在现实搜索场景中实现了专业化执行与广泛泛化之间的良好平衡。

英文摘要

Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and iterative retrieval over dynamic information sources. However, optimizing language models for specialized search behaviors often incurs an alignment tax, where gains in search performance come at the expense of general-purpose capabilities, limiting their effectiveness as universal assistants. In this technical report, we present the training framework behind the Yuanbao search agent, designed to achieve search specialization without sacrificing general intelligence. Built upon the Hunyuan3 architecture, our framework combines agentic reinforcement learning for autonomous search with a cross-domain expert On-Policy Distillation (OPD) pipeline. Experts specializing in complementary general-purpose domains are distilled into the search-specialized student, restoring and further enhancing its broad capabilities. Rather than treating specialization and general capability as competing objectives, our hybrid training strategy jointly optimizes both, effectively mitigating the alignment tax. Extensive experiments demonstrate that the resulting model achieves competitive search performance while consistently improving its general-purpose capabilities, providing a favorable balance between specialized execution and broad generalization in real-world search scenarios.

发表机构

  • Yuanbao Team, Tencent Shanghai Innovation Institute(腾讯上海研究院元宝团队)

机构由 AI 辅助整理,请以论文原文为准。

↑