发表机构
Fudan University; Shanghai Innovation Institute(复旦大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PersonaManifold框架,将LLM人格表征建模为弯曲黎曼流形,利用测地线引导替代欧几里得运算,并引入BST基准,在三个开源模型上验证了更准确的行为相似性预测与更连贯的中间人格。
AI 中文摘要
在推理时控制大型语言模型(LLM)的人格对于角色扮演、个性化对话和社会模拟至关重要。近期方法从模型的激活空间中提取人格向量,并在线性表征假设下应用欧几里得运算——加法、缩放和线性插值。然而,这些方法自身报告了系统性失败:非正交的特质维度、不对称的天花板效应和阻力效应,以及多特质组合中的显著偏差,表明线性各向同性假设不成立。我们提出PersonaManifold,一个将人格表征建模为激活空间中弯曲的低维黎曼子流形上的点的框架。我们估计流形的内在几何——局部度量张量、测地距离和Ollivier-Ricci曲率——并引入测地线引导,该方法沿流形测地线而非欧几里得直线在人格之间进行插值。我们还提出了行为相似性三元组(BST)基准,该基准基于六个已建立的心理构念自动生成情境问题,并通过行为反应而非自我报告问卷来定义人格相似性。在三个开源LLM上的实验表明,人格激活形成具有异质曲率的流形,测地距离比欧几里得替代方法更准确地预测行为相似性,且各向异性和曲率具有独立贡献,测地线引导在我们的BST基准和外部评估中均产生更连贯的中间人格,其优势集中在流形偏离平坦度最大的高偏差区域。
英文摘要
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry---local metric tensors, geodesic distances, and Ollivier-Ricci curvature---and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.
CommentsAccepted to NeurIPS 2026