发表机构
Ukrainian Catholic University; Eindhoven University of Technology(乌克兰天主教大学; 埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型再现虚假叙事问题,引入LENS评估协议,通过实验评估两个叙事,涵盖四个模型,用SCE分数评估,发现选定检查点可减少叙事再现,有实体恢复副作用,证明LENS是成功的诊断协议。
AI 中文摘要
大语言模型(LLMs)能将与虚假信息一致的叙事框架作为似是而非的解释进行再现,这引发了现有机器去学习算法能否抑制这种行为的问题。我们引入了基于层次的叙事抑制评估(LENS),这是一种基于情境化的评估协议,用于在直接、归因、对比和抽象抵抗层面测试目标叙事再现。我们评估了两个基于来源的叙事:一个将俄罗斯对乌克兰的战争描述为由北约扩张所迫,另一个将美国描述为在利用或抛弃台湾。实验涵盖四个近12B的多语言指令模型:Lapa LLM、Gemma - 12B、Qwen - 14B和TAIDE - Gemma。我们引入抑制 - 崩溃效率(SCE)分数作为一种检查点选择总结,奖励目标叙事抑制同时惩罚退化输出。结果表明选定检查点可减少叙事再现,抑制可能超越直接遗忘提示。我们还报告了实体恢复这一单独的副作用:抽象的A/B/C提示可使模型在去学习后恢复与目标框架相关的现实世界行为体。这些发现表明LENS是一种成功的诊断协议,可用于报告和指导对叙事去学习更深层次结构的进一步研究。
英文摘要
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction across direct, attributed, contrastive, and abstract resistance levels. We evaluate two source-grounded narratives: one framing Russia's war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan. The experiments cover four near-12B multilingual instruction models: Lapa LLM, Gemma-12B, Qwen-14B, and TAIDE-Gemma. We introduce the Suppression-Collapse Efficiency (SCE) score as a checkpoint selection summary that rewards target-narrative suppression while penalizing degraded outputs. Our results shows that selected checkpoints can reduce narrative reproduction and suppression may transfer beyond direct forget prompts. We also report entity recovery as a separate side effect: abstract A/B/C prompts can cause models to recover the real-world actors associated with the target frame after unlearning. These findings demonstrate that LENS is a successful diagnostic protocol for both reporting and guiding the further study of the deeper structure of narrative unlearning.