arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STAIF:一种用于复杂指令跟随的逐阶段优化方法

STAIF: A Stage-wise Optimization for Complex Instruction Following

Jian Hong, Chen Cheng, Quan Liu, Yuhao Chen, Enhong Chen

arXiv 2607.22649首次发表:更新:

发表机构

University of Science and Technology of China; iFLYTEK Research Group(中国科学技术大学; 科大讯飞研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型遵循复杂指令的挑战,提出STAIF逐阶段优化框架,先通过偏好优化提高软约束敏感度,再用可验证奖励强化学习确保硬约束遵守,构建相关数据集,验证了该方法的设计并展现出领先性能和泛化能力。

AI 中文摘要

对于大语言模型而言,遵循带有多个明确约束的复杂指令仍是一项基本挑战。现有的对齐方法,如DPO,优化的整体奖励信号往往会忽视对单个约束的严格满足,尤其是在分布外或多约束设置下。本文提出了STAIF,一种逐阶段优化框架,将主观(软)约束的对齐与客观可验证(硬)约束的优化解耦。第一阶段使用多个负样本进行偏好优化,以提高对软约束的敏感度,而第二阶段使用可验证奖励的强化学习来确保严格遵守硬约束。为支持该方法,构建了STAINSTRUCT,一个约31000条复杂多约束指令的高质量双语(英语、中文)数据集。大量分析验证了STAIF的设计,并在具有代表性的基准测试中,相对于强大的基线展示了领先的性能以及真正的泛化能力。

英文摘要

Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction of individual constraints, particularly under out-of-distribution or multi-constraint settings. In this paper, we propose STAIF, a stage-wise optimization framework that decouples the alignment of subjective (soft) constraints from the optimization of objectively verifiable (hard) constraints. Stage 1 applies preference optimization with multiple negative samples to sharpen sensitivity to soft constraints, while Stage 2 applies Reinforcement Learning with Verifiable Rewards (RLVR) to enforce strict compliance with hard constraints. To support this method, we construct STAINSTRUCT, a high-quality bilingual (English, Chinese) dataset of approximately 31,000 complex multi-constraint instructions. Extensive analyses validate the design of STAIF and show state-of-the-art performance on representative benchmarks against strong baselines, as well as genuine generalization.

Comments16 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑