arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27805nlin.CDcond-mat.dis-nn

理性中的混沌:思维链大语言模型(LLMs)如何探寻答案

Chaos in reason: How chain-of-thought LLMs can look for an answer

Gregorio Jaca, Kristóf Benedek, János Török

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将非线性动力学与混沌理论工具应用于思维链LLMs,证实其存在混沌标志性特征,明确各Transformer子模块对扰动的作用,揭示其与典型混沌系统的结构相似性及分形特性,指出注意力机制诱导的非线性耦合是混沌行为的关键驱动因素。

中文摘要 AI 辅助

大语言模型(LLMs)在广泛任务中取得了显著性能,但其内部动态仍知之甚少。本研究将非线性动力学与混沌理论工具应用于LLMs,通过分析文本与隐状态轨迹,证明LLMs表现出混沌的标志性特征:对初始条件的强敏感性,体现为近邻轨迹的间歇性跳跃式发散与有界演化,且在不同距离度量下结果一致。对Transformer子模块的精确雅可比分析显示,自注意力机制与前馈网络会放大并传播扰动,而归一化与残差连接则抵消这种放大并提升稳定性。递归图表明LLMs与洛伦兹吸引子等典型混沌系统存在结构相似性,维度分析揭示隐状态空间存在分形结构,尤其在最后几层表现显著。本研究提出,注意力机制诱导的非线性耦合是驱动这种混沌行为的关键因素。

英文摘要

Large Language Models (LLMs) have achieved remarkable performance across a wide range of tasks, yet their internal dynamics remain poorly understood. In this work, we apply the tools of nonlinear dynamics and chaos theory to LLMs. By analyzing both text and hidden state trajectories, we demonstrate that LLMs exhibit hallmark signatures of chaos, including strong sensitivity to initial conditions, manifested as intermittent, jump-like divergence of nearby trajectories combined with bounded evolution, with consistent results across different distance metrics. An exact Jacobian analysis of the Transformer's sub-blocks shows that self-attention and the feed-forward network expand and propagate perturbations, while normalization and residual connections counteract this expansion and promote stability. Recurrence plots show structural similarities between LLMs and canonical chaotic systems such as the Lorenz attractor, while dimension analysis reveals fractal structures in the hidden state space, particularly pronounced in the last layers. We propose that the nonlinear coupling induced by attention mechanisms plays a key role in driving this chaotic behavior.

补充信息

↑