arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环状态可以安全遗忘什么?

What Can a Recurrent State Safely Forget?

Linzhe Zhang, Changming Xu

arXiv 2609.23366首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过预测商形式化循环模型的安全遗忘边界,提出可审计有限未来框架,并给出最优探针复杂度,实现有界语义失真的安全修正。

AI 中文摘要

循环模型必须保留改变未来行为的信息,同时抑制隐藏状态误差。这两个目标相互冲突:收缩提高了稳定性,但沿未来区分方向的收缩会破坏记忆。我们通过循环状态空间的预测商来形式化这一边界。当两个隐藏状态诱导相同的条件未来时,它们是等价的;它们的等价类形成预测纤维。每个精确保持语义的修正器在此商上充当恒等映射。在隐藏维度为d、预测维度为k的正则点上,它最多可以消除d - k个独立方向。这建立了一个离散-连续边界:有限预测状态允许正半径的精确修正盆地,而在有限维欧几里得空间中,不可数无限的未来可区分状态在任意正半径扰动后无法被解码。为了操作化这一原理,我们开发了一个可审计的有限未来框架。一个紧凑的部署库W在声明的修正域上,针对独立的审计库A(W ⊆ A)进行评估。在生成式探针访问和审计度量覆盖下,有限随机轨迹为分离边际Ω_{W|A}(δ)提供了高概率证书。在此认证边际内保留学习的W预测,保证了有界的审计语义失真。对于内在审计维度k,所需的探针结果规模为O(M * Ω^{-(k+2)}),其中M = |A|;匹配的极小极大下界证明了该指数是最优的。通过显式的完备性模量,将保证扩展到连续未来。受控实验在安全优先的评估范式下验证了认证边际、缩放规律和自动化探针细化。

英文摘要

Recurrent models must preserve information that changes future behavior while suppressing hidden-state error. These objectives conflict: contraction improves stability, but contraction along a future-distinguishing direction destroys memory. We formalize this boundary through the predictive quotient of a recurrent state space. Two hidden states are equivalent when they induce the same conditional future; their equivalence classes form predictive fibers. Every exact semantics-preserving corrector acts as the identity on this quotient. At a regular point with hidden dimension d and predictive dimension k, it can eliminate at most d - k independent directions. This establishes a discrete-continuous boundary: finite predictive states admit positive-radius exact correction basins, whereas an uncountable continuum of future-distinguishable states cannot be decoded after arbitrary positive-radius perturbations in finite-dimensional Euclidean space. To operationalize this principle, we develop an auditable finite-future framework. A compact deployment bank W is evaluated against an independent audit bank A (W subseteq A) on a declared correction domain. Under generative probe access and audit-metric coverage, finite stochastic rollouts furnish a high-probability certificate for the separation margin Omega_{W|A}(delta). Preserving learned W-predictions within this certified margin guarantees bounded audit-semantic distortion. For intrinsic audit dimension k, the required probe outcomes scale as O(M * Omega^{-(k+2)}), where M = |A|; a matching minimax lower bound proves this exponent is optimal. Extending guarantees to continuous futures is achieved via an explicit completeness modulus. Controlled experiments validate the certified margins, scaling laws, and automated probe refinement under a safety-first evaluation paradigm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑