arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36388cs.AI

人格剂量:用于分级特质控制的校准激活引导

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出PersonaDose方法,通过校准激活引导实现对语言模型人格特质的精细控制,在多个模型上显著提升特质表达,并区分了控制器的行为范围与请求准确性。

中文摘要 AI 辅助

激活引导系数设定干预强度,但要求特定程度的人格表达需要行为量表。我们研究人格剂量:通过特质描述和请求的平均强度来控制语言模型。PersonaDose将共享的、描述条件的FLAS控制器专门化于人格响应,然后根据测量的特质表达校准其流动时间。训练响应不与请求的目标强度配对。在Llama-3.1-8B、Qwen3-8B和Gemma-3-4B上,PersonaDose在Persona Vectors一致性下限75处将核心特质表达分别提高了33.2、18.3和17.8个百分点,相对于对比激活加法。校准选择的设置在保留问题上保持表达优势,尽管一致性下限并非对所有特质都成立。在七个训练特质中,校准请求在每模型28个目标中14-22个可校准目标上产生4.7-6.2个点的平均目标误差。这些结果将控制器学习到的行为范围与该范围内请求的准确性区分开来。

英文摘要

An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • Zhejiang University(浙江大学)
  • University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑