arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

闭环知识动力学:饱和与逃逸的操作框架

Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape

Xuening Wu, Shan Yu, Shenqin Yin

arXiv 2607.14185首次发表:更新:

发表机构

Pfizer; Institute of Humanities and Social Science Data, Fudan University(辉瑞公司; 复旦大学人文社会科学数据研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究闭环知识系统饱和及逃逸问题,引入三级操作框架,利用李雅普诺夫漂移条件等刻画系统动态,通过案例展示反馈影响,建立了稳定性工具、干预效果与跨域诊断间的操作联系。

AI 中文摘要

反馈驱动循环支持大语言模型、强化学习和自主发现中的迭代改进,但在反复内部反馈下其收益往往会减少。我们研究闭环知识系统为何饱和以及何种外部信息能使其超越当前吸引子。引入了一个三级操作框架,其中知识状态$x_t$通过由结构参数$\theta$索引的转移核$K_{\theta}$演化。定义了控制结构、吸引子和盆地等概念,通过结构干预改变$\theta$并产生可检测的核差异。利用李雅普诺夫漂移条件表明稳定内部动力学趋近有界稳定区域,通过度量条件和KL下界刻画逃逸。案例研究展示了反馈强度和对齐如何影响质量改进逃逸。我们的贡献是在稳定性工具、可测量干预效果和跨域诊断之间建立了操作联系。

英文摘要

Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback. We study why closed-loop knowledge systems saturate and what external information can move them beyond their current attractors. We introduce a three-level operational framework in which knowledge states $x_t$ evolve through transition kernels $K_θ$ indexed by a structural parameter $θ$. The governing structure is defined as the observational equivalence class of $θ$ induced by these kernels, while attractors and basins are properties of the fixed-$θ$ dynamics. A structural intervention changes $θ$ and produces a detectable kernel discrepancy on pre-specified probe states, making structural change falsifiable. Using a Lyapunov drift condition, we show that stable internal dynamics approach bounded stability regions with exponentially attenuated transients and a noise-controlled residual floor. We characterize escape through a metric condition on intervention-induced attractor displacement and a baseline-relative KL lower bound for increasing escape probability. This analysis also explains why conditional mutual information alone cannot certify escape: it measures variation among intervention-conditioned updates rather than departure from the no-intervention law. Case studies in LLM code repair, sparse-reward reinforcement learning, and Bayesian optimization use matched continuation controls to illustrate how feedback strength and alignment affect quality-improving escape. Our contribution is an operational connection among stability tools, measurable intervention effects, and cross-domain diagnostics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑