arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34754cs.CL

Draft-KV:在语言模型之间学习有用的潜在通信

Draft-KV: Learning Useful Latent Communication Between Language Models

Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang, Peng Zhong, Fengming Zhu, Xi Peng, Linqi Song, Jacky Keung, Jingyu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出Draft-KV,通过发送共享者起草答案时的键值状态作为潜在通信,使冻结的接收器在MMLU-Redux上达到78.04%,验证了消息内容的关键作用。

中文摘要 AI 辅助

潜在通信在语言模型之间传递内部状态而非解码文本,但接收器准确率的提高并不能证明接收器使用了消息内容。在五个方法-数据集组合中,将每条消息替换为来自无关问题的消息,准确率变化最多不超过0.60个百分点,即使通信相比单独使用接收器带来了15.44个百分点的提升。因此,接口可以提供增益,同时使共享者变得可有可无。Draft-KV则改为发送共享者在起草当前问题答案时形成的键值状态。线性投影将这些状态置于通过门控注意力分支读取的侧记忆体中,渐进式训练从消息重建过渡到在错误消息危害防护下的答案监督。两个模型均保持冻结,接口仅训练105万参数,比C2C少348倍。使用Qwen3-8B共享者时,冻结的Qwen2.5-0.5B-Instruct接收器在MMLU-Redux上达到78.04%的准确率,而单独使用时为37.45%,使用重新分配的消息时为36.40%。在固定接口大小下,将共享者从0.6B扩展到8B可将准确率从46.11%提升至78.04%;通信还能迁移到保留任务上,并且当两个模型各自持有不同证据时,可以超越两个模型各自的表现。

英文摘要

Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.

发表机构

  • City University of Hong Kong(香港城市大学)
  • University of Science and Technology of China(中国科学技术大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • AIPD, Tencent(腾讯AIPD)
  • Theory Lab, Huawei(华为理论实验室)
  • Hong Kong Metropolitan University(香港都会大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑