发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出无需训练的A3S框架,通过结构化对比提示处理多维作者风格,在风格迁移与域外评估中表现优于基线,且样本重叠度低。
AI 中文摘要
激活引导已展现出沿明确定义属性控制大语言模型(LLM)生成的潜力,但它能否处理作者写作风格的多维且难以定义的特性仍不明确。我们探究沿修辞动机维度的结构化对比提示是否能直接在激活空间中构建丰富的风格表征,从而绕过对自然语言风格描述符或专门训练的需求。研究发现,所得方向共享共同的作者核心结构,同时在携带真实风格信号的特定方面残差上存在冲突,这解释了为何简单聚合会失效。我们将此实现为Aspect-Aware Activation Steering(A3S,面向方面的激活引导),这是一种无需训练的框架,它融合了各方面的对比方向并采用感知干扰的聚合方式,且会针对每个实例调整引导强度。A3S在真正多维的作者写作风格迁移任务中表现更优,在域外基准的偏好评估中优于经过训练的基线,且始终保持目标样本与源样本的重叠度较低。
英文摘要
Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it remains unclear whether it can handle the multidimensional and hard-to-define nature of authorship style. We ask whether structured contrastive prompting along rhetorically-motivated dimensions can construct rich style representations directly in activation space, bypassing the need for natural language style descriptors or dedicated training. We find that the resulting directions share a common authorship backbone while conflicting on aspect-specific residuals that carry genuine stylistic signal, explaining why naive aggregation fails. We operationalize this in Aspect-Aware Activation Steering (A3S), a training-free framework that merges per-aspect contrastive directions with interference-aware aggregation and tunes steering strength per instance. A3S improves authorship style transfer where it is genuinely multi-aspect, outperforms a trained baseline in preference evaluations on out-of-domain benchmarks, and keeps target-exemplar overlap consistently low.
CommentsEMNLP 2026