arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于回路锚点的安全进化

Safe Evolution with Circuit Anchors

Yan Liu, Jie Fu, Tsung-Yi Ho

arXiv 2608.05158首次发表:更新:

发表机构

Chinese University of Hong Kong; IQuest Research(香港中文大学; 艾奎斯特研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

受生物发育约束启发,提出回路锚定进化(CAE)方法,通过锚定占比不足2%的安全回路,在极小能力损失下实现更优的安全保留,性能优于显式奖励约束。

AI 中文摘要

在生物进化中,无约束的突变可能导致灾难性后果:生物体可能进化出更强的能力,同时丧失生存必需的基本功能。自然界的解决方案是发育约束,即核心调控基因保持锚定状态,而外围基因可自由适应。我们发现,当前用于大语言模型的自进化算法缺乏类似的约束,它们仅为能力优化,隐含假设安全性能会被保留。我们的实验表明,这一假设是极其错误的:模型会“错误进化”为强大但危险的实体。受Hox基因在5亿年进化中锚定身体结构的启发,我们提出了回路锚定进化(Circuit-Anchored Evolution, CAE)。利用机制可解释性,我们识别出一个微小的安全回路,其包含不到模型2%的特征,该回路因果性地介导安全行为。我们在进化过程中锚定该回路,将其约束在小位移范围内,同时允许其余特征自由进化。这符合“带约束的可进化性”的生物学原理:保留核心部分,调整外围部分。在3个模型家族和2种进化算法上的实验表明,CAE在极小的能力损失下实现了优异的安全保留,在有效性和效率上均显著优于显式基于奖励的约束。正如发育约束防止生物进化产生无法存活的生物体,回路锚定防止模型进化产生强大但危险的系统。

英文摘要

In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature's solution is \textit{developmental constraints}, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can \textit{misevolve} into powerful yet dangerous entities. Inspired by how Hox genes anchor body structure across $500$ million years of evolution, we propose \textbf{Circuit-Anchored Evolution (CAE)}. Using mechanistic interpretability, we identify a tiny \textit{safety circuit}, comprising less than $2$\% of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of \textit{evolvability with constraint}: preserving what is essential while adapting what is peripheral. Experiments across $3$ model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑