Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
反事实推理用于可操控的多元价值观对齐大语言模型
机构 * Renmin University of China(中国人民大学) ; Microsoft Research Asia(微软亚洲研究院) ; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心,教育部)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG
AI总结 COUPLE通过反事实推理框架实现多元价值观对齐,解决现有方法在处理细粒度价值目标时的依赖性和优先级控制问题。
Comments NeurIPS 2025. 41 pages, 7 figures
Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems. (NeurIPS 2025)