arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21279cs.CL

用于指令微调的统一道德价值数据集

A Unified Moral-Value Dataset for Instruction Tuning

Zhaohui Zeng, Florian Mai

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型与人类价值观对齐问题,构建统一道德价值数据集用于指令微调,通过合并现有数据集并转换格式而成,实验表明其结合通用任务数据集训练可保持性能,为对齐研究提供资源。

中文摘要 AI 辅助

大语言模型发展迅速,成为日常生活中的重要工具。但如何使大语言模型与特定人类价值观对齐仍是开放问题。近期研究表明指令微调在零样本任务中有很大潜力。尽管已有许多指令微调数据集,但缺乏针对道德场景和行为的设计。我们构建了可直接用于指令微调的统一道德价值数据集,通过合并现有道德价值数据集并转换为指令-响应格式而成。实验表明结合通用任务数据集与我们的数据集进行混合训练可保持通用任务性能,还报告了混合比例对价值导向任务性能的影响。我们的工作为指令微调提供了道德价值数据集,为进一步的对齐研究提供了有用资源。数据集可通过此https链接获取。

英文摘要

Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open problem. Recent studies show that instruction tuning has strong potential for zero-shot tasks and may serve as an effective approach to addressing value alignment. Nevertheless, although many datasets for instruction tuning already exist, they are not specifically designed around moral scenarios and behaviors. We construct a unified moral-value dataset that can be directly used for instruction tuning. This dataset is built upon existing moral-value datasets by merging them into a unified corpus and converting them into an instruction-response format. We show that training on a mixed dataset combining general task datasets with our dataset preserves general-task performance, and we report preliminary observations on how the mixing ratio affects value-oriented task performance. Our work provides a moral-value dataset for instruction tuning and offers a useful resource for further alignment research. The dataset is available at https://huggingface.co/datasets/teohzzh/value-for-instruction-tuning.

发表机构

  • RWTH Aachen(亚琛工业大学)
  • University of Bonn(波恩大学)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉玛尔机器学习与人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑