arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

K/V缓存干预在仅解码器语言模型中解耦表示对齐与人格表达

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models

Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han

arXiv 2609.11020首次发表:更新:

AI 中文总结

本研究通过K/V缓存干预在Llama-3.1-8B上发现表示对齐与人格表达解耦,中层替换兼顾表达与多样性,位置扰动抑制表达,揭示轨迹级移植机制。

AI 中文摘要

我们研究了K/V缓存干预——将目标条件化的K/V轨迹移植到源人格生成中——作为仅解码器语言模型中人格控制的结构化表面。针对固定的源到目标人格对,在应用于Llama-3.1-8B的13种干预配置中,我们报告了表示层对齐与行为表达之间的两种一致解耦,以及在位置扰动下的常见失败。首先,所有层带的K/V替换(早期、中期、晚期)都实现了强局部V空间对齐(V-gap分别为0.91、0.89、0.84),但只有中层替换(第9-20层)结合了显著的目标标记表达与保留的词汇多样性。其次,全层和中层替换诱导了可比较的对齐(V-gap分别为0.94和0.89),却产生了不同的词汇多样性特征(TTR分别为0.65和0.77)。第三,位置扰动(滞后和洗牌)应用了不同的操作,却一致地抑制了目标人格表达——这是一种常见的行为失败,而非严格的解耦。因此,在我们研究的机制中,仅靠表示层相似性指标并不能充分预测下游人格表达;K/V缓存成为一个可控但结构受限的干预表面。由于移植的轨迹携带了目标自身生成的令牌历史,我们将该干预表征为轨迹级移植,而非孤立的人格表示注入;一个相同令牌序列对照,在源与目标条件下解码相同的令牌序列,重现了L28表示转移的符号和层定位,表明该转移并非仅由导入的令牌历史解释。这些发现表征了高信号设置中的表示-行为解耦,而非建立跨模型或人格对的普遍性。

英文摘要

We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixed source-to-target persona pair, we report two consistent dissociations between representation-level alignment and behavioral expression, plus a common failure under position perturbations. First, all layer-band K/V replacements (early, mid, late) achieve strong local V-space alignment (V-gap 0.91, 0.89, 0.84), but only mid-layer replacement (layers 9-20) combines substantial target-marker expression with preserved lexical diversity. Second, full and mid-layer replacement induce comparable alignment (V-gap 0.94 vs. 0.89) yet produce different lexical-diversity profiles (TTR 0.65 vs. 0.77). Third, position perturbations (lag and shuffle) apply distinct operations yet uniformly suppress target-persona expression -- a common behavioral failure rather than a strict dissociation. Representation-level similarity metrics alone are thus not sufficient predictors of downstream persona expression in the regimes we study; the K/V cache emerges as a controllable but structurally constrained intervention surface. Because the transplanted trajectory carries the target's own generated token history, we characterize the intervention as trajectory-level transplantation rather than isolated persona-representation injection; a same-token-sequence control, decoding an identical token sequence under source vs. target conditioning, reproduces the sign and layer localization of the L28 representational shift, indicating the shift is not explained solely by imported token history. These findings characterize representation-behavior dissociation in a high-signal setting rather than establishing universality across models or persona pairs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑