arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合语言模型中注意力机制的作用与循环控制的功能

What Attention Recalls and Recurrence Controls in Hybrid Language Models

Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov

arXiv 2609.04434首次发表:更新:

发表机构

YSDA; AIRI; HSE University(YSDA; AIRI; 高等经济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过两种缓存级干预方法,揭示混合语言模型中注意力与循环通道的功能划分,证实注意力负责精确检索、循环控制输出风格的作用。

AI 中文摘要

混合语言模型将注意力机制与固定大小的循环状态相结合,但各通道的作用仍不明确。我们引入两种缓存级干预方法:Split-prefill仅保留预填充上下文中的KV缓存或循环状态,然后生成答案;State-swap在单次前向传播中将一个上下文的KV缓存与另一个上下文的循环状态配对。在Qwen3.5和Falcon-H1上,两个通道的功能存在明显划分:精确检索仅通过注意力机制实现(达到完整准确率的64%-98%),而通过循环机制则降至零;输出语言和角色则呈现相反模式,两者均通过循环机制保留(准确率分别为70%-80%和3-5倍),而仅使用KV缓存时语言准确率降至约1%。State-swap从因果角度证实:答案的取值来自KV侧,语言风格来自循环侧;仅使用循环机制生成时,模型会接受上下文中从未出现但与已见内容含义或部分相似的词语。注意力机制提供对已说内容的查找,循环状态则塑造模型后续的输出方式。

英文摘要

Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions. Split-prefill keeps only the KV cache or only the recurrent state from a prefilled context, then generates an answer. State-swap pairs the KV cache from one context with the recurrent state from another in a single forward pass. On Qwen3.5 and Falcon-H1, the two channels split sharply by function. Exact retrieval survives only through attention (64-98% of full accuracy) and collapses to zero through recurrence. Output language and persona reverse the pattern: both survive recurrence (70-80% and 3-5x) while KV-only drops to ~1% language accuracy. State-swap confirms this causally: the answer takes its value from the KV side and its language from the recurrent side. Recurrent-only generation also accepts words that were never in the context but share meaning or parts with seen items. Attention provides a lookup over what was said; the recurrent state shapes how the model says it next.

CommentsAccepted to Findings of EMNLP 2026. 13 pages, 3 figures, 8 tables. Code: https://github.com/kirillTerra/split-prefill

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑