arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoCo-ACE:用于保留感知知识注入的基于展开的在线蒸馏

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

Yan Hong, Wei Li, Kedong Xiu, Jun Lan, Shuheng Zhou, Zhongcai Lyu, Huijia Zhu, Weiqiang Wang, Jianfu Zhang

arXiv 2607.24771首次发表:更新:

发表机构

Ant Group; Zhejiang University; Shanghai Jiao Tong University(蚂蚁集团; 浙江大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对知识注入中拟合权威答案致行为偏差问题,提出RoCo-ACE方法,通过同展开无参考/基于参考条件的似然对比及参考侧锚定校正,在多知识注入设置等下,实现最佳注入知识准确性且保持保留率接近基础模型。

AI 中文摘要

知识注入通过新的事实或特定领域知识更新预训练的语言模型,但拟合完整的权威答案可能导致未更新行为出现偏差。在线蒸馏通过对模型生成的展开进行训练来减轻这种偏差,然而统一的基于参考条件的蒸馏提供的监督较为粗糙:它可能会低估参考支持的展开令牌,并只能间接监督遗漏的事实。我们引入了RoCo-ACE,一种用于知识注入的基于展开的在线蒸馏目标。RoCo使用同展开无参考/基于参考条件的似然对比,将额外的蒸馏权重重新分配给参考支持的展开令牌,而ACE则为未在展开中出现的权威锚点添加稀疏的参考侧锚定校正,无需完整答案模仿。在三种知识注入设置、六个保留基准、多个基线和多个基础模型上,RoCo-ACE在比较方法中实现了最佳的注入知识准确性,同时保持评估的保留率接近基础模型。

英文摘要

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly. We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation. Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑