arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SIGMA:从模型规范中实现自我改进的对齐泛化

SIGMA: Self-Improving Alignment Generalization from a Model Spec

Jingyu Zhang, Shruti Palaskar, Daniel Khashabi, Benjamin Van Durme, Leon A. Gatys, Joseph Yitan Cheng

arXiv 2610.07935首次发表:更新:

发表机构

Apple; Johns Hopkins University(苹果公司; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SIGMA流水线,利用模型自身推理能力,基于模型规范生成任务并自我评判训练,实现安全对齐自我改进,在多轮智能体环境中显著降低有害性并保持通用能力。

AI 中文摘要

LLM智能体日益能够执行复杂任务,并在易于验证的目标(如软件工程和数学)上递归地自我改进。由于对齐更难验证,这造成了能力增长而缺乏适当安全对齐的风险日益增加,尤其是当能力扩展到自动研究和网络安全领域时。现有方法侧重于利用可验证反馈进行能力自我改进,或利用更强模型的监督或精选数据进行对齐训练,这形成了对齐的外部监督瓶颈。我们探究当前模型能否改进自身的安全对齐,并提出了SIGMA,一个数据生成和训练流水线,能够实现对齐的自我改进,并泛化到分布外场景。仅给定一个陈述模型期望行为的“模型规范”,SIGMA利用模型的推理能力来强化其自身的安全推理。SIGMA首先执行规范引导的任务合成,将候选模型作为任务设计智能体,生成多样的对齐困境场景,并将其转化为训练任务,以压力测试其对模型规范的理解。接下来,SIGMA通过监督微调和基于评分标准的强化学习进行自我判断的对齐训练,其中模型自身作为奖励模型。尽管仅在单轮对话数据上训练,SIGMA在多轮智能体环境(AgentHarm有害性从22.6降至14.8;智能体失对齐从79.1降至3.8)中提升了安全对齐,优于审慎对齐和宪法AI基线,并保留了通用能力。分析表明,平衡无害性和有用性的模型规范、用于安全审慎的测试时推理,以及来自SIGMA任务设计智能体的高质量评分标准,对于有效的自我改进至关重要。

英文摘要

LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a "Model Spec" stating the model's desired behavior, SIGMA leverages a model's reasoning capabilities to strengthen its own safety reasoning. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into training tasks that stress-test its understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA's task designer agent are crucial for effective self-improvement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑