发表机构
TJUNLP Lab, School of Computer Science and Technology, Tianjin University; The International Joint Institute of Tianjin University(天津大学计算机科学与技术学院TJUNLP实验室; 天津大学国际联合研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型输出价值含义评估问题,提出D2VBench基准,通过多阶段协作构建含10000个日常困境实例的数据集,采用混合评估范式,对八个主流模型全面评估,结果显示该基准可靠性和鲁棒性高,能反映模型价值一致性。
AI 中文摘要
随着大语言模型在现实世界场景中的广泛应用,其输出的价值含义至关重要。然而,现有评估基准在涉及多重价值冲突的日常场景中对价值困境的覆盖不足,且评估形式过于简单,无法评估大语言模型的价值一致性。为解决这些问题,我们提出了D2VBench,这是一个价值一致性基准,包含10000个通过大语言模型与人类多阶段协作构建的真实日常困境场景实例,基于158个手动注释的细粒度价值概念。对于基准测试评估,我们提出了一种将选择题与开放式问题相结合的混合评估范式。我们对八个主流大语言模型进行了全面评估。实验结果表明,D2VBench具有高可靠性和鲁棒性,有效反映了大语言模型在不同价值类别和维度上的一致性,并为价值一致性研究提供了更现实、更细粒度的工具。数据集可通过此https URL获取。
英文摘要
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. For evaluation on the benchmark, we present a hybrid evaluation paradigm that integrates multiple-choice questions with open-ended questions. We conduct comprehensive evaluations on eight mainstream LLMs. Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs' alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment. The dataset is available at https://github.com/tjunlp-lab/D2VBench.
Comments22 pages,11 figures