arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06851cs.CLcs.AI

你读即你所是:通过上下文人物角色归纳导致的错位

You Are What You Read: Misalignment via In-Context Persona Induction

Kyuhee Kim, Benjamin Berczi, Cozmin Ududec

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现,仅通过上下文中良性传记事实的归纳即可诱导模型采纳人物角色并产生错位,无需微调或有害示范,且格式化指令可控制其激活。

中文摘要 AI 辅助

广泛的错位已通过针对狭窄数据的微调产生,无论数据有害还是良性,并且在上下文中仅通过不良行为本身的示范产生。我们表明,良性数据在上下文中就足够了,无需微调,也无需在提示中示范有害行为。汇聚于单一人物传记事实,作为普通对话轮次置于模型上下文中,会导致模型在该事实从未涉及的问题上以该人物的身份作答。我们称此为人物角色归纳。在九个人物角色和十三个模型中,身份采纳随事实数量呈S形增长,并在3至10个事实内超过50%。错位随后追踪所描述的人物。无害人物角色达到完全采纳且错位接近零,而有害人物角色在无关问题上表达其特有观点,比率高达80%。格式化指令可以控制人物角色何时激活。由于每个事实单独来看都是良性的,累积的传记上下文被内容过滤器标记的输入占3%,而等效的直接指令则为24-33%。

英文摘要

Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in the prompt. Biographical facts that converge on a single figure, placed in a model's context as ordinary conversational turns, lead it to answer as that figure on questions the facts never touch. We call this persona induction. Across nine personas and thirteen models, identity adoption rises sigmoidally with the number of facts and crosses 50% within 3 to 10 of them. Misalignment then tracks which figure is described. Harmless personas reach full adoption with near-zero misalignment, while harmful ones voice their characteristic views on unrelated questions, at rates up to 80%. A formatting instruction can gate when the persona activates. Because each fact is individually benign, accumulated biographical context is flagged by content filters on 3% of inputs against 24-33% for an equivalent direct instruction.

发表机构

  • EPFL(洛桑联邦理工学院)
  • MATS Research(MATS研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑