arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

采样揭示风格:无监督、免训练发现LLM激活中的提示条件风格轴

Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

Ajit Mallavarapu, Ziwei Gu

arXiv 2609.19150首次发表:更新:

发表机构

Cornell University; Harvard University(康奈尔大学; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种免训练方法,通过对单提示高温采样补全并对其隐藏激活做PCA,自动发现LLM中的风格轴,经人类标注验证有效,并揭示跨模型差异。

AI 中文摘要

大型语言模型(LLMs)在其隐藏激活中编码了丰富的风格结构,但发现哪些风格维度对给定提示是显著的通常需要监督对比数据。我们提出了一种免训练、提示条件相关的替代方法:我们反复对单个提示在高温下采样补全,对汇集后的隐藏激活应用主成分分析(PCA),并自动从极点生成中标记所得轴。我们在一个两阶段研究中,针对245个人类引发的风格标注验证了所发现的轴。在我们最强的模型(Qwen-3.5-4B-Instruct)上,前两个轴以72.8%的精确率和43.6%的宏召回率匹配自发请求的人类维度,75.6%的有效性评分判断轴的极点生成与其标签准确相符,相邻注释者间一致性为90.9%。可发现性强烈依赖于模型:Qwen模型和Llama-3.2-3B都暴露了人类显著的轴,而DeepSeek-7B-Chat的精确率降至35.3%,其主要成分由结构而非风格方差主导。因此,对模型自身解码方差进行简单PCA是探测LLM表示中风格结构的有效、低成本方法,也暴露了该结构组织方式上的显著跨模型差异。

英文摘要

Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions with 72.8% precision and 43.6% macro-recall, and 75.6% of validity ratings judge the axes' polar generations accurate to their labels, with 90.9% adjacent inter-annotator agreement. Discoverability is strongly model-dependent: both Qwen models and Llama-3.2-3B expose human-salient axes, while DeepSeek-7B-Chat drops to 35.3% precision, its leading components dominated by structural rather than stylistic variance. Simple PCA over a model's own decoding variance is thus an effective, low-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross-model differences in how that structure is organized.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑