arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言编码的网络拓扑结构使大型语言模型能够对复杂网络进行推理

Language-encoded network topology enables large language models to reason about complex networks

Ucchwas Talukder Utsha, Sakib Mostafa, James Zou, Md Tauhidul Islam

arXiv 2609.03229首次发表:更新:

发表机构

Stanford University; Stanford University School of Medicine(斯坦福大学; 斯坦福大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BioGlyph将网络拓扑编码为结构角色语言,提升开源LLMs的网络结构推理能力,在20个跨领域网络上准确率较现有方法最高提升26个百分点。

AI 中文摘要

网络描述了生物学及其他领域的系统,从蛋白质相互作用、社会关系到电网和引文记录。对这类系统进行推理需要理解其结构:哪些元素是核心,哪些连接桥接了不同的社区,以及当元素被移除时结构会如何变化。尽管大型语言模型(LLMs)在自然语言处理方面表现出色,但当网络以边列表、句子或测量表形式呈现时,它们难以回答这类问题,因为必须推断其结构含义。本文提出BioGlyph,它将网络拓扑结构编译为可解释且可迁移的结构角色语言。BioGlyph结合图划分和结构测量来识别核心节点、社区核心、跨社区连接器等角色,并通过固定规则将其转换为通用词汇。该表示通过结构角色、支持证据和语义后果来描述每个元素,且不改变网络本身和大型语言模型。在涵盖五个领域的20个网络上,BioGlyph显著提升了开源大型语言模型回答结构推理问题的能力,在系统准确率上比基于边、数值和学习到的表示方法高出多达26个百分点。 ablation实验表明,该增益源于以语义可解释的术语显式编码结构角色。这种增益在密集的社区结构网络中更为显著,而在拓扑结构更容易从文本中推断的稀疏网络中则会减弱。在出芽酵母的蛋白质相互作用网络中,BioGlyph揭示了生物组织:跨社区连接器富集了必需基因,而外围蛋白则呈缺失状态。因此,BioGlyph为语言模型和科学家提供了一种可解释的表示,用于对网络结构进行推理。

英文摘要

Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central, which connections bridge separate communities, and how it changes when elements are removed. Although large language models (LLMs) excel at natural language, they struggle with such questions when networks are given as edge lists, sentences or measurement tables, because their structural meaning must be inferred. Here we introduce BioGlyph, which compiles network topology into an interpretable and transferable language of structural roles. BioGlyph combines graph partitioning and structural measurements to identify roles such as hubs, community cores and cross-community connectors, and fixed rules to translate them into a universal vocabulary. The representation describes each element through its structural role, supporting evidence and semantic consequences, leaving both the network and the LLM unchanged. Across twenty networks spanning five domains, BioGlyph substantially improves open LLMs' ability to answer structural reasoning questions, outperforming edge-based, numerical and learned representations by up to 26 percentage points in system accuracy. Ablations show that the gain comes from explicitly encoding structural roles in semantically interpretable terms. The gain is more prominent in dense, community-structured networks and diminishes in sparse networks whose topology is more readily inferred from text. In a budding-yeast protein-interaction network, BioGlyph exposes biological organization: cross-community connectors are enriched for essential genes, whereas peripheral proteins are depleted. BioGlyph thus provides an interpretable representation for both language models and scientists to reason about network structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑