用于提升大语言模型指令遵循能力的跨关系偏好学习
Cross-Relational Preference Learning for Better LLM Instruction Following
浏览论文内容
中文总结 AI 辅助
针对LLMs遵循复杂指令时忽略不同指令允许响应空间关系的问题,提出CRPL框架,通过两项关键技术构建高质量偏好数据,在多方法、主干及基准上实现显著改进与强泛化。
中文摘要 AI 辅助
大语言模型(LLMs)在遵循复杂指令方面的能力仍有限。现有方法常依赖偏好学习来提升该能力,但通常会忽略不同指令的允许响应空间之间的关系,这限制了模型对细微且多样的约束变化进行对齐。为解决此问题,我们提出跨关系偏好学习(CRPL),这是一种用于构建偏好数据的新框架,通过两项关键技术显式建模指令间关系:跨关系扰动和跨区域对采样。这能够生成更多样的偏好数据,以捕捉广泛的约束变化范围。此外,我们引入了一种基于原子约束的验证机制,以严格评估响应的满意度,确保构建高质量的偏好对。在多种偏好学习方法(如DPO、KTO)、LLM主干及四个指令遵循基准上进行的大量实验表明,我们的方法相较于现有基线取得了显著改进,并展现出强泛化性。
英文摘要
Large Language Models (LLMs) still exhibit limited capability in following complex instructions. While existing approaches often rely on preference learning to enhance this ability, they typically overlook the relationships between the permissible response spaces of different instructions, which restricts a model to align with subtle and diverse constraint variations. To address this, we propose Cross-Relational Preference Learning (CRPL), a novel framework for constructing preference data that explicitly models inter-instruction relationships through two key techniques: Cross-Relationship Perturbation and Cross-Region Pair Sampling. This enables the generation of more diverse preference data that captures a wide spectrum of constraint variations. Additionally, we introduce an atomic constraint-based verification mechanism to rigorously assess response satisfaction, ensuring high-quality preference pair construction. Extensive experiments across multiple preference learning methods (e.g., DPO, KTO), LLM backbones and four instruction-following benchmarks demonstrate that our approach achieves substantial improvements over prior baselines and exhibits strong generalization.
发表机构
- School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
- Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(西安交通大学陕西省大数据知识工程重点实验室)
- School of Distance Education, Xi’an Jiaotong University(西安交通大学远程教育学院)
机构由 AI 辅助整理,请以论文原文为准。