arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越线性上下文:基于图引导的证据导航用于长篇小说推理的本地9B语言模型

Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model

Wenji Fu

arXiv 2609.22939首次发表:更新:

发表机构

Research Institute of Economics and Management; Southwestern University of Finance and Economics(经济与管理研究院; 西南财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究测试冻结知识图谱能否赋予本地9B模型长篇小说推理能力,实验显示图路径在234题上达53.85%准确率,虽未通过统计校正,但发现证据集中于图核心及构建流程差异显著。

AI 中文摘要

长上下文模型阅读小说时,就像一个人阅读打印输出:按叙事顺序逐词处理,整个历史争夺固定的注意力预算。而侦探并非如此工作。他们会整理事件发生的时间顺序,并保留一张人物关系图,这样第一章的线索可以与书末提出的问题相遇。我们测试一个冻结的知识图谱能否赋予小型本地模型同样的自由。三十部侦探小说和234道多项选择题由一个固定的qwen3.5:9b阅读器在九种条件下回答:五种图路径、近期窗口基线、全书压缩、普通向量检索以及仅问题对照。最强的图路径达到53.85%(126/234),而近期窗口为46.15%,压缩为51.28%,向量检索为51.71%,仅问题为40.17%。在无书则任何模型都无法回答的子集上,图路径达到42.86%。十五个图-基线对比中没有一个通过Holm校正,因此我们将结果呈现为关于设计的探索性证据。两个结构性发现比头条数字更经得起审视:标注证据集中在这些图的拓扑核心(合并富集度2.35倍),且两条图构建流程在标注覆盖率上差异巨大(线索段落的16%对73%),因此仅合并准确率会掩盖正在测量的瓶颈。

英文摘要

Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective does not work that way. They sort what happened when, and they keep a map of who relates to whom, so a clue from chapter one can meet a question asked at the end of the book. We test whether a frozen knowledge graph can give a small local model that same freedom. Thirty detective novels and 234 multiple-choice questions are answered by one fixed qwen3.5:9b reader under nine conditions: five graph routes, a recent-window baseline, whole-book compression, ordinary vector retrieval, and a question-only control. The strongest graph route reaches 53.85% (126/234) against 46.15% for the recent window, 51.28% for compression, 51.71% for vector retrieval and 40.17% for question-only. On the subset that no model can answer without the book, the graph route reaches 42.86%. None of the fifteen graph-baseline contrasts survives Holm correction, so we present the result as exploratory evidence about a design. Two structural findings survive scrutiny better than the headline number: annotated evidence concentrates in the topological core of these graphs (2.35x enrichment, pooled), and the two graph-building pipelines differ so much in annotation coverage (16% versus 73% of clue paragraphs) that pooled accuracy alone would hide which bottleneck is being measured.

Comments10 pages, 11 figures, 2 tables. Code, graph data and an interactive demo: https://github.com/fuxiaoji/novel-graph-lab

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑