QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
QA-LIGN:通过宪法分解的问答对对齐大语言模型
AI总结 QA-LIGN通过分解奖励信号提升LLM对齐效果,降低攻击成功率并保持低拒绝率,实现安全与帮助性的帕累托最优。
Comments Findings of the Association for Computational Linguistics: EMNLP 2025, pages 20619-20642, Suzhou, China
Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 20619-20642, Suzhou, China, 2025