arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.03160cs.AIcs.CL

VALUEFLOW:迈向大语言模型中多元化和可引导的基于价值的对齐

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

  • Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电气与计算机工程系)
  • Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔国立大学人工智能交叉学科项目)

机构由 AI 辅助整理,请以论文原文为准。

Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do

更新

AI总结:

提出VALUEFLOW框架,通过分层价值嵌入、强度标注数据库和锚定评估器,实现大语言模型在价值强度上的可控对齐,解决现有方法在提取、评估和引导方面的不足。

AI中文摘要:

将大语言模型(LLMs)与人类价值的多元光谱对齐仍然是一个核心挑战:基于偏好的方法通常无法捕捉更深层次的动机原则。基于价值的方法提供了更原则性的路径,但仍存在三个差距:提取常常忽略层次结构,评估检测存在但未校准强度,并且LLMs在受控强度下的可引导性仍未得到充分理解。为解决这些限制,我们引入了VALUEFLOW,这是第一个统一框架,涵盖提取、评估和引导,并具有校准的强度控制。该框架整合了三个组件:(i) HIVES,一个层次化价值嵌入空间,捕捉理论和跨理论的价值结构;(ii) 价值强度数据库(VIDB),一个大规模资源,包含基于排序聚合得出的强度估计的价值标注文本;(iii) 一个基于锚点的评估器,通过将模型输出与VIDB面板进行排序,产生一致的强度分数。使用VALUEFLOW,我们在十个模型和四个价值理论上进行了全面的大规模研究,识别了可引导性的不对称性和多价值控制的组合规律。本文建立了一个可扩展的基础设施,用于评估和控制价值强度,推进了LLMs的多元化对齐。

英文摘要:

Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and the steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, the first unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HIVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.

补充信息

↑