发表机构
Chinese University of Hong Kong; IQuest Research(香港中文大学; 艾奎斯特研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
受生物发育约束启发,提出回路锚定进化(CAE)方法,通过锚定占比不足2%的安全回路,在极小能力损失下实现更优的安全保留,性能优于显式奖励约束。
AI 中文摘要
在生物进化中,无约束的突变可能导致灾难性后果:生物体可能进化出更强的能力,同时丧失生存必需的基本功能。自然界的解决方案是发育约束,即核心调控基因保持锚定状态,而外围基因可自由适应。我们发现,当前用于大语言模型的自进化算法缺乏类似的约束,它们仅为能力优化,隐含假设安全性能会被保留。我们的实验表明,这一假设是极其错误的:模型会“错误进化”为强大但危险的实体。受Hox基因在5亿年进化中锚定身体结构的启发,我们提出了回路锚定进化(Circuit-Anchored Evolution, CAE)。利用机制可解释性,我们识别出一个微小的安全回路,其包含不到模型2%的特征,该回路因果性地介导安全行为。我们在进化过程中锚定该回路,将其约束在小位移范围内,同时允许其余特征自由进化。这符合“带约束的可进化性”的生物学原理:保留核心部分,调整外围部分。在3个模型家族和2种进化算法上的实验表明,CAE在极小的能力损失下实现了优异的安全保留,在有效性和效率上均显著优于显式基于奖励的约束。正如发育约束防止生物进化产生无法存活的生物体,回路锚定防止模型进化产生强大但危险的系统。
英文摘要
In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature's solution is \textit{developmental constraints}, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can \textit{misevolve} into powerful yet dangerous entities. Inspired by how Hox genes anchor body structure across $500$ million years of evolution, we propose \textbf{Circuit-Anchored Evolution (CAE)}. Using mechanistic interpretability, we identify a tiny \textit{safety circuit}, comprising less than $2$\% of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of \textit{evolvability with constraint}: preserving what is essential while adapting what is peripheral. Experiments across $3$ model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.