arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12304cs.AI

探究大语言模型(LLM)内部:小世界连通性作为推理性能的标志

Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance

Zheng Huang, Sansheng Cao, Enpei Zhang, Weikang Qiu, Elynn Chen, Xiang Zhang, Yaoqing Yang, Rex Ying, Dawei Zhou, Yujun Yan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究以小世界连通性为LLM推理的结构标志,提出基于核心、桥分数的SWA剪枝方法,可更好保留小世界组织与性能,降低WikiText困惑度达20%。

中文摘要 AI 辅助

理解大语言模型(LLM)的推理能力,需要超越行为表现,考察推理能力如何在内部组织中体现。受神经科学研究发现的启发——更高的智力与功能脑网络中更强的小世界组织相关,本研究将小世界连通性作为LLM推理的结构标志进行探究。我们从注意力头激活相似性构建功能图,发现更高的小世界指数(SWI,捕捉局部聚类和短全局路径)在不同模型和训练检查点中,始终与更好的流体推理性能相关。由于局部聚类是小世界组织的核心,我们进一步考察对模型性能重要的注意力头在社区内部及跨社区的连接情况,发现这些注意力头在自身社区内的连接权重占比更大(核心分数高),且跨社区的权重分布更集中(桥分数低)。这些观察结果促使我们提出假设:高核心分数和低桥分数是注意力头对推理能力重要性的结构指标。我们通过剪枝验证该假设,提出小世界分配(SWA)方法,这是一种基于上述分数引导的分层稀疏分配方法。在六个LLM上的实验显示,SWA相比其他竞争分配策略能更好地保留小世界组织和模型性能,将WikiText困惑度降低了多达20%。综上,这些发现确定了小世界功能连通性是可测量的LLM推理性能标志,提供了补充行为评估的结构视角。

英文摘要

Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning. We construct functional graphs from attention-head activation similarities and find that a higher small-world index (SWI), capturing local clustering and short global paths, consistently correlates with better fluid reasoning performance across models and training checkpoints. Since local clustering is central to small-world organization, we further examine how heads important for model performance connect within and across communities. We find that these heads tend to have a larger share of connection weight within their own communities (high core scores) and a more concentrated weight distribution across communities (low bridge scores). These observations motivate the hypothesis that high core and low bridge scores serve as structural indicators of head importance for reasoning capability. We validate this hypothesis through pruning, introducing Small-World Allocation (SWA), a hierarchical sparsity allocation method guided by these scores. Across six LLMs, SWA better preserves small-world organization and model performance than competing allocation strategies, reducing WikiText perplexity by up to 20%. Together, these findings identify small-world functional connectivity as a measurable signature of LLM reasoning performance, offering a structural perspective that complements behavioral evaluation.

发表机构

  • Dartmouth College(达特茅斯学院)
  • Yale University(耶鲁大学)
  • New York University(纽约大学)
  • UNC Charlotte(北卡罗来纳大学夏洛特分校)
  • Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑