arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于减轻大语言模型微调中隐藏行为的推理时共识

Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning

Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern

arXiv 2607.23394首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型微调中隐藏行为问题,通过从不同源收集多数据集并学习共性,在解码时聚合参考模型下一个token分布,引入两种共识解码器,有效抑制特定源不当行为,保留共享期望行为。

AI 中文摘要

近期研究表明,即便在少量中毒数据上微调语言模型,也会植入针对性不当行为,看似良性的数据也可能传播广泛的隐藏偏好。标准防御措施虽能减弱但无法消除这些影响。本文通过冗余追求鲁棒性,从不同来源收集多个数据集,仅学习它们之间的共性。为此在每个源数据集上微调一个单独的参考模型,并在解码时聚合其下一个token的分布。引入两种共识解码器:token-wise最小值,将每个token限制在任何源分配的最低概率;基于基线的变体,在源方向相反的任何token上恢复到基线概率。在受控中毒任务、潜意识学习和新兴错位中,共识解码抑制了特定源的不当行为,同时保留了共享的期望行为。

英文摘要

Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, mixing in harmless data, and regularization, attenuate these effects but do not eliminate them. We instead pursue robustness through redundancy: collecting multiple datasets from different sources and only learning what is common between them. Thus, if only a subset of sources are malicious, the misbehavior will be blocked. In order to implement this defense strategy, we fine-tune a separate reference model on each source's dataset and aggregate their next-token distributions at decoding time. We introduce two consensus decoders: a token-wise minimum, which caps each token at the lowest probability any source assigns, and a base-relative variant, which reverts to the base probability on any token the sources move in opposing directions. We further relax exact agreement to tolerate partial support across sources and different surface expressions of the same intention. Across controlled poisoning tasks, subliminal learning, and emergent misalignment, consensus decoding suppresses source-specific misbehavior while preserving shared desirable behavior, including cases where union training and weight averaging retain the unwanted behavior.

Comments21 pages, 4 figures. Code: https://github.com/AdhyyanNarang/consensus-aggregation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑