arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Safin-1:通过内存原生状态演化实现内在安全

Safin-1: Safety from Within through Memory-Native State Evolution

Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu

arXiv 2609.00092首次发表:更新:

AI 中文总结

Safin-1 是遵循「内在安全」理念的基础模型系列,基于 MARCH 架构实现内存路由与状态演化,可提升下游安全任务性能,为安全能力的原生自适应维持提供路径。

AI 中文摘要

长 horizon 复杂任务要求基础模型积累信息、维持内部状态并在长期交互中自适应。安全应是模型本身的固有属性,而非仅依赖外部保障或事后对齐(如监督微调)的行为约束。这催生了「内在安全」理念,即与安全相关的能力通过模型的原生计算进行表征和调用。我们提出 Safin-1,这一基础模型系列通过内存路由和状态演化实现该理念。Safin-1 基于跨上下文历史的内存锚点路由(MARCH)构建,这是一种维持结构化内存状态并通过内容条件路由选择性检索相关历史信息的网络架构。它支持持久能力状态的测试时自适应,无需重复修改主干,可在共享基础上实现可控的专业化。我们通过「安全状态」在下游安全任务上研究该接口,展示了基于状态的自适应可带来显著的安全提升。更广泛而言,路由状态接口将上下文内存与持久能力自适应统一在模型的原生计算中,将内存从先前上下文的被动记录重构为维持和演化模型行为的主动载体。在通用能力、长上下文理解、检索及效率方面的评估进一步验证了 Safin-1。这些发现为将安全作为原生状态且可自适应维持的能力提供了路径。本工作仅是对「内在安全」的初步架构探索,要实现更广泛愿景仍需大量后续工作。

英文摘要

Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑