arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过价值冲突探究LLM价值表达的结构与动态

Probing the Structure and Dynamics of LLM Value Expression through Value Conflicts

Kaicheng Zhang, Jingyi Xiao, Renjun Hu, Xiaoling Liu, Yunshi Lan, Xuan Zhou

arXiv 2609.07296首次发表:更新:

发表机构

East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出冲突驱动的价值探测框架,通过干预价值冲突探究十个LLM的价值表达,发现表达二元性、功能可引导性和有限可塑性三种模式,揭示了价值表达的结构与动态边界。

AI 中文摘要

大型语言模型(LLM)的伦理评估通常将模型价值描述为静态且单一的。相反,我们认为LLM的价值表达更应被理解为一种结构化但动态的现象。为探究这一点,我们引入了冲突驱动的价值探测(Conflict-driven Value Probing),这是一个受控框架,它将LLM置于价值冲突中,并实施四种类型的干预来扰动这些冲突,以探测LLM的价值表达。将该框架应用于十个LLM,我们识别出三种反复出现的模式。(1)表达二元性:模型在抽象评估中从广泛的理想主义取向转向具体冲突中更务实的优先级。(2)功能可引导性:模型能轻易地将其表达的价值概况重新配置为任务定义的价值目标。(3)有限可塑性:这种重新配置并非没有约束,即压力引发安全和目标导向的优先级转变,而负面框架则区分了受保护的价值与那些更易被重定向的价值。综合来看,这些发现刻画了LLM价值表达的结构与动态:上下文灵活地重新配置表达的优先级,但仍在行为边界之内。这种行为描述为理解LLM的可控性、对齐和安全性提供了基础。代码和数据可在以下网址获取:此https URL。

英文摘要

Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic. In contrast, we argue that LLM value expression is better understood as a structured yet dynamic phenomenon. To investigate this, we introduce Conflict-driven Value Probing, a controlled framework that places LLMs in value conflicts and implements four types of interventions that perturb these conflicts to probe LLM value expression. Applying this framework to ten LLMs, we identify three recurring patterns. (1) Expression duality: models shift from broad idealistic orientations in abstract assessment toward more pragmatic priorities in concrete conflicts. (2) Functional steerability: models readily reconfigure their expressed value profiles toward task-defined value objectives. (3) Bounded plasticity: such reconfiguration is not without constraints, i.e. pressure induces a security- and goal-oriented priority shift while negative framing distinguishes protected values from those more amenable to redirection. Together, these findings characterize both the structure and dynamics of LLM value expression: context flexibly reconfigures expressed priorities, yet within behavioral boundaries. This behavioral account provides a foundation for understanding controllability, alignment, and safety in LLMs. Code and data are available at https://github.com/ZeroGen-Lab/CFProbe.

CommentsCode and data are available at https://github.com/ZeroGen-Lab/CFProbe

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑