Reasoning Up the Instruction Ladder for Controllable Language Models
在指令阶梯上推理以实现可控语言模型
机构 * Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) ; Microsoft Research(微软研究院) ; Allen Institute for AI(艾伦人工智能研究所)
AI总结 通过将指令层级解析重构为推理任务,并构建VerIH数据集进行轻量级强化学习,使模型能优先处理高层级指令,在冲突场景下提升约20%的指令遵循准确率,并增强对越狱和提示注入的鲁棒性。
Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 39332-39354