arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16589cs.AI

大型语言模型有价值观吗?大语言模型中价值观的定量分析与对齐框架

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出先验-环境-认知(PEC)框架,定量分析106个LLM的价值观,并建立自适应对齐处方,以最小干预实现高效、精确的价值观引导。

中文摘要 AI 辅助

随着大型语言模型(LLMs)日益处理复杂的主观任务,将其意图和行为与人类价值观对齐已成为一个关键的科学挑战。然而,当前的努力受到一个显著的行为悖论的困扰:它们在微小的措辞变化下会不可预测地波动(“摇摆”),却又顽固地忽视纠正根深蒂固偏见的明确指令(“僵化”)。解决这一双重性对于可靠的AI对齐至关重要。为了系统地理解并安全地引导这些潜在的主观偏好,我们的研究围绕三个基本问题展开。首先,LLMs是否拥有内在的价值体系?通过将106个LLM(每个模型15万次查询)和95,000份人类调查档案的响应投射到一个共享的社会学空间中,我们实证确认它们确实拥有。然而,它们并未反映人类的多样性,而是凝聚成一个高度集中、理想化的价值核心。其次,这些价值观如何被量化?我们提出了先验-环境-认知(PEC)框架。该模型在数学上将价值表达定义为固有倾向(如参数权重,即先验)、外部情境(如用户提示,即环境)和内部推理过程(如思维链,即认知)的联合结果。最后,如何将LLM的价值观对齐到期望的目标?利用PEC诊断,我们建立了一个自适应的“对齐处方”。该方法并非盲目应用资源密集型的训练,而是识别每个维度所需的最小有效干预,从零成本提示到针对性的参数更新。广泛的经验验证证实,我们的方法成功验证了LLM价值观的存在,准确量化了它们的转变,并实现了比传统盲目训练更高效、更精确的引导,且不降低通用能力。

英文摘要

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs' values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive "Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

发表机构

  • State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
  • School of Industry-education Integration, University of Chinese Academy of Sciences(中国科学院大学产教融合学院)
  • Beihang University(北京航空航天大学)
  • Beijing Jiaotong University(北京交通大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑