在基础开放权重模型中复现情绪表征的几何结构
Replicating the Geometry of Emotion Representations in a Base Open-Weights Model
浏览论文内容
中文总结 AI 辅助
本研究在gemma-2-27b基础模型上复现Claude Sonnet 4.5情绪表征几何,验证情感环形结构及效价轴对齐,并揭示其源于预训练表征。
中文摘要 AI 辅助
Sofroniew等人(2026)报告称,Claude Sonnet 4.5中的情绪概念以向量形式表征,其几何结构映射了人类情感心理学。我们在基础预训练模型google/gemma-2-27b上复现了该研究的表征核心,继承了所有公开参数,通过公开规则解决未明确的步骤,仅更换了主体模型。从205,200个新生成的、符合原始语料库设计的Claude Sonnet 4.5故事中,我们提取了171个情绪向量并恢复了核心结果。主要主成分构成情感环形结构(PC1承载26.7%的方差,原始研究约为27%;PC2承载13.4%,原始研究约为14%),情绪聚类为相似的直观类别,且该几何结构在较宽的晚中段保持稳定。效价轴与人类规范对齐(r = 0.72,原始研究为0.81),并在不同尺度和深度下保持稳定。唤醒度在r = 0.67(原始研究为0.66)处对齐,但仅在完整的171情绪尺度和较晚深度下成立,因此我们未将其归类为已复现。扩展原始分析,一次46层扫描定位到L22-26处的尖锐接缝,在该处几何结构巩固且词汇读出变得清晰。嵌入层基线发现,大部分几何结构已存在于静态词元嵌入中,唤醒度除外。在最高激活的保留文本上,该几何结构以r = 0.907预测词元级共激活。至少52%的向量在结构上非概念的词元上达到峰值,这是最大激活混杂因素的测量下限。即使文档包含向量的情绪词,峰值落在该词上的概率仅为6.1%。由于主体是基础模型且刺激为Claude生成的虚构文本,恢复的结构是Claude渲染情绪预训练表征的属性。原始的因果分析和面向助手的分析不在范围内。代码和数据已发布。
英文摘要
Sofroniew et al. (2026) report that emotion concepts in Claude Sonnet 4.5 are represented as vectors whose geometry mirrors human affect psychology. We replicate the representational core of that study on the base pretrained model google/gemma-2-27b, inheriting every disclosed parameter, resolving unspecified steps by disclosed rules, and changing only the subject model. From 205,200 newly generated Claude Sonnet 4.5 stories matching the original corpus design, we extract 171 emotion vectors and recover the core results. The leading principal components form an affective circumplex (PC1 carries 26.7% of variance against the original study's ~27%, PC2 13.4% against ~14%), emotions cluster into similar intuitive families, and the geometry holds across a broad late-middle band. The valence axis aligns with human norms (r = 0.72, against 0.81) and is stable across scales and depth. Arousal aligns at r = 0.67 (against 0.66) but only at the full 171-emotion scale and late depth, so we do not classify it as replicated. Extending the original analysis, a 46-layer sweep locates a sharp seam at L22-26, where the geometry consolidates and the vocabulary readout becomes legible. An embedding-layer baseline finds much of the geometry already present in the static token embeddings, with arousal as the exception. On top-activating held-out text, the geometry predicts token-level co-activation at r = 0.907. At least 52% of vectors peak on structurally non-conceptual tokens, a measured floor for max-activation confounds. Even when a document contains the vector's emotion word, the peak lands on that word only 6.1% of the time. Because the subject is a base model and the stimuli are Claude-generated fiction, the recovered structure is a property of the pretrained representation of Claude-rendered emotion. The original's causal and assistant-facing analyses are out of scope. Code and data are released.
发表机构
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。