arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20927cs.CLcs.AI

MentorPulse:为长文本生成更新跨模型隐层引导

MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

Ziwu Liu, Guozhong Li, Chen Qiu, Weiyang Kong, Panos Kalnis

首次发表
浏览论文内容

中文总结 AI 辅助

MentorPulse通过每16令牌刷新的插槽记忆更新技术,在13个数据集上缩小52.2%的导师-学生性能差距,在长文本生成任务中优于C2C、T2T等方法。

中文摘要 AI 辅助

跨模型隐层引导允许一个冻结的大型“导师”模型对输入进行一次编码,而一个冻结的小型“学生”模型则基于该生成信号进行生成。现有方法会固定该信号,假设其在输出增长时仍保持有用,但我们证明这在长文本生成中不成立。在多轮指令遵循任务中,静态引导使4B参数学生模型的约束满足度比无引导基线低2.5个百分点;每生成16个令牌进行一次无需训练的刷新,仅更改记忆内容,即可恢复出比该基线高2.0个百分点的增益。我们提出MentorPulse,以实用成本保持引导的新鲜度:它将导师模型的状态压缩为有限容量的插槽记忆,增量处理新生成的令牌,并通过门控交叉注意力更新学生模型读取的记忆,无需重置学生模型的键值(KV)缓存。窗口化刷新训练揭示了前缀条件记忆的关联。在13个数据集上,MentorPulse在宏平均水平上缩小了52.2%的导师-学生性能差距,优于C2C、T2T及同等预算的LoRA方法,在长输出上的增益最大。它在来自三个模型家族的全部11对导师-学生模型中均表现最佳,且随着能力差距增大,性能增益的幅度会缩小,此外,一种轻量型读取模式检查可在部署前预测增益。测量得到的成本确定了在长输出中主导文本引导的刷新间隔。

英文摘要

Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. On multi-turn instruction following, static guidance pushes a 4B student's constraint satisfaction 2.5 points below its no-guidance baseline; a training-free refresh every 16 tokens changes only the memory content and restores a 2.0-point gain over that baseline. We propose MentorPulse to keep guidance fresh at practical cost: it compresses mentor states into a capped slot memory, incrementally processes newly generated tokens, and updates the memory that the student reads through gated cross-attention without resetting the student's KV cache. Windowed Refresh Training exposes the bridge to prefix-conditioned memory. Across thirteen datasets, MentorPulse closes 52.2% of the mentor-student gap on macro average, outperforming C2C, T2T, and equal-budget LoRA, with the largest gains on long outputs. It performs best on all eleven mentor-student pairs from three model families, with margins that narrow as the capability gap grows, and a lightweight read-pattern check predicts the gain before deployment. Measured costs identify refresh intervals that dominate text guidance on long outputs.

发表机构

  • King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑