arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过迭代零空间投影提取人格子空间用于调制

Extracting Persona Subspaces Through Iterative Nullspace Projection For Modulation

Ananya Malik, Mai ElSherief

arXiv 2610.04676首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PaSS,通过迭代零空间投影提取人格子空间,在推理时调制LLM行为,无需重训练,实现更强、更大的调制并保持内容保真度。

AI 中文摘要

大型语言模型(LLMs)能够采用不同的人格来调整其语义、专业知识和视角,以适应不同的用户和任务。对这些特质的精确控制对于确保模型行为的安全性和可靠性至关重要。现有的方法如激活引导和基于提示的人格归纳将人格简化为单一主导方向,忽略了那些只有在主导信号被分解后才会出现的更细微、嵌套的特质。我们引入调制作为一种设置,其中人格上下文已经嵌入到被操作的内容中,要求控制方法放大或抑制已经存在的特质,而不是从头注入。PaSS是一种推理时的控制范式,它将人格建模为模型潜在空间中的多维子空间,无需监督对比示例。人格子空间通过迭代概念擦除提取,并应用于调制人格引导的生成,无需重新训练。为了提取该子空间,我们使用迭代零空间投影(INLP)来线性且迭代地隔离人格特定方向。我们针对MATH-500、TinyAlpaca、GSM8K和IFEval等多样任务对六种人格进行了因果评估,表明判别性、迭代的子空间提取能够捕捉给定人格下的多样特质,实现比单方向加法方法更强、更大的调制,同时保持内容保真度。我们进一步研究每个子空间内的单独剥离方向,以揭示它们编码的人格行为的不同方面。总体而言,我们表明人格子空间为调制LLM行为提供了一个可控、可解释且可泛化的框架,而不会牺牲任务性能。

英文摘要

Large Language Models (LLMs) can adopt distinct personas to tune their semantics, expertise, and perspective to different users and tasks. Precise control over these traits is critical to ensure safety and reliability in model behavior. Existing methods like activation steering and prompt-based persona induction reduce a persona to a single dominant direction, missing the finer, nested traits that emerge only once that dominant signal is factored out. We introduce modulation as a setting where the persona context is already embedded in the content being manipulated, requiring control methods to amplify or suppress a trait already present rather than inject it from scratch. PaSS is an inference-time control paradigm that models personas as multi-dimensional subspaces in a model's latent space without supervised contrastive examples. The persona subspaces are extracted via iterative concept erasure and applied to modulate persona-guided generation without retraining. To extract this subspace, we use Iterative Nullspace Projections (INLP) to linearly and iteratively isolate persona-specific directions. We causally evaluate six personas against diverse tasks like MATH-500, TinyAlpaca, GSM8K, and IFEval, showing that discriminative, iterative subspace extraction captures diverse traits underlying a given persona, enabling stronger and larger modulation than single-direction additive methods, while maintaining content fidelity. We further study individual peeled directions within each subspace to uncover the distinct aspects of persona behavior they encode. Overall, we show that persona subspaces offer a controllable, interpretable, and generalizable framework for modulating LLM behavior without sacrificing task performance.

Comments29 pages, 12 tables, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑