arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAME:面向涌现性错位的令牌归因与掩蔽

TAME: Token Attribution and Masking for Emergent misalignment

Md Rayhanul Masud, Md Rizwan Parvez

arXiv 2609.16754首次发表:更新:

发表机构

University of California, Riverside; Qatar Computing Research Institute (QCRI)(加州大学河滨分校; 卡塔尔计算研究所(QCRI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对涌现性错位,提出TAME框架,通过令牌归因定位高影响令牌并掩蔽,使Llama和Qwen的EM分别降低23倍和36倍,揭示信号在于表达自信度而非领域词汇。

AI 中文摘要

在狭窄且有缺陷的数据上对对齐语言模型进行微调,可能会在训练领域之外诱发有害行为,这被称为涌现性错位(EM)。先前的工作已将EM定位在模型权重、激活和训练文档中,但尚不清楚哪些训练令牌携带相关的微调信号。我们引入了TAME(面向涌现性错位的令牌归因与掩蔽),一个三阶段框架:令牌归因通过前向传播经过发布的LoRA适配器,计算微调更新提高每个响应令牌似然度的强度;信号表征发现高归因令牌中的模式;因果验证通过归因引导的损失掩蔽来测试这些模式。在发布的EM生物体和一个包含6,849个示例的医疗建议子集上,归因是集中的(前5%的令牌持有32%的归因质量),并且在Llama中,医疗词汇的归因减少,但无根据确定性的语域却富集,即使在控制令牌稀有性之后也是如此。在重新微调期间掩蔽高归因令牌,使Llama中的EM减少23倍,Qwen中的EM减少36倍,困惑度成本集中在目标语域而非医疗内容上;等量的随机掩蔽则使EM保持不变。在Llama中,归因模式表明,EM相关信号更多在于缺陷内容表达的自信程度,而非其领域词汇;因果掩蔽效应本身在两个模型家族中均成立。

英文摘要

Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weights, activations, and training documents, but it remains unclear which training tokens carry the relevant fine-tuning signal. We introduce TAME (Token Attribution and Masking for Emergent Misalignment), a three-stage framework: token attribution scores how strongly the fine-tuning update raises each response token's likelihood, using forward passes through a released LoRA adapter; signal characterization finds patterns among high-attribution tokens; and causal validation tests them by attribution-guided loss masking. On released EM organisms and a 6,849-example medical-advice split, attribution is concentrated (the top 5% of tokens hold 32% of the mass) and, in Llama, depleted for medical vocabulary but enriched for a register of unwarranted certainty, even after controlling for token rarity. Masking high-attribution tokens during fresh fine-tuning cuts EM by 23x in Llama and 36x in Qwen, with the perplexity cost concentrated on the targeted register rather than on medical content; an equal random mask leaves EM unchanged. In Llama, the attribution pattern suggests that EM-relevant signal lies more in how confidently flawed content is expressed than in its domain vocabulary; the causal masking effect itself holds across both model families.

CommentsAccepted at EMNLP UncertaiNLP Workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑