arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19158cs.CLcs.IRcs.LG

VisKG-LM:将知识图谱编译为视觉记忆用于多项选择问答

VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

Yixin Peng, Er Jin, Shiwei Luo, Diego Collarana, Stefan Decker

首次发表
浏览论文内容

中文总结 AI 辅助

VisKG-LM将知识图谱子图离线编译为视觉记忆图像,供语言模型在推理时查阅,在三个问答基准上超越现有方法,证明了编译视觉记忆作为在线图传播的可行替代方案。

中文摘要 AI 辅助

知识图谱通常通过图神经网络对检索到的子图进行编码,并在在线推理路径中与语言模型融合,从而集成到问答系统中。因此,即使知识图谱从未改变,每次对一对(问题,候选答案)进行评分时,相同的子图都会从头开始重新编码,这跨越了训练轮次、随机种子和评估运行。我们提出疑问:检索到的知识图谱是否可以改为离线编译一次,然后作为只读记忆进行访问?VisKG-LM证明了这是可行的,它将图编码与语言推理解耦。它将每个检索到的候选特定子图序列化为关系标记路径,并将结果渲染为图像,其二维布局保留了路径的分支结构。每个图像离线编码一次,并缓存以供重用。在推理时,语言模型仅从文本中理解问题和候选答案,只有其最后一层会查阅缓存的视觉记忆,同时读取其全局布局和局部关系细节。因此,图信息仅在文本被理解之后才进入。在CommonsenseQA、OpenBookQA和MedQA-USMLE的测试集上,VisKG-LM相比GreaseLM分别提高了1.2、0.8和4.3个百分点,同时匹配或超越了拥有70亿参数的视觉语言模型GraphVis,而VisKG-LM仅使用约4亿在线参数。与接收相同关系标记路径的匹配文本仅控制组相比,它在三个基准上分别获得了4.2、6.5和5.1个百分点的提升。这些提升表明,完整的视觉记忆接口在路径文本化之外增加了价值,并支持将编译的视觉记忆作为在线图传播的替代方案。

英文摘要

Knowledge graphs are usually integrated into question answering by encoding a retrieved subgraph with a graph neural network and fusing it with the language model in the online inference path. The same subgraph is therefore re-encoded from scratch every time a pair is scored, across training epochs, seeds, and evaluation runs, even though the knowledge graph never changes. We ask whether the retrieved knowledge graphs can instead be compiled once, offline, and then accessed as read-only memory. VisKG-LM shows that it can, by decoupling graph encoding from language reasoning. It serializes each retrieved candidate-specific subgraph as Relation-Labeled Paths and renders the result as an image whose two-dimensional layout preserves the branching structure of the paths. Each image is encoded once, offline, and cached for reuse. At inference, the language model contextualizes the question and candidate from text alone, and only its final layer consults the cached visual memory, reading both its global layout and its local relational detail. The graph information thus enters only after the text has been understood. On the test sets of CommonsenseQA, OpenBookQA, and MedQA-USMLE, VisKG-LMimproves over GreaseLM by $1.2$, $0.8$, and $4.3$ points, respectively, while matching or surpassing GraphVis, a $7$B vision-language model, with only about $400$M online parameters. Against a matched text-only control that receives the identical Relation-Labeled Paths, it gains $4.2$, $6.5$, and $5.1$ points across the three benchmarks. These gains show that the complete visual-memory interface adds value beyond path textualization alone and support compiled visual memory as an alternative to online graph propagation.

↑