arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27669cs.CLcs.AI

路径至关重要:在小语言模型的KGQA中超越答案准确率进行评估

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini

AI总结:

本研究通过THESEUS框架隔离小语言模型在知识图谱问答中的导航能力,发现答案准确率与路径保真度(PED)常不一致,且受提示影响,主张评估应超越端点准确率。

AI中文摘要:

小语言模型(SLMs)日益与知识图谱(KGs)结合使用,然而端到端的知识图谱问答将图访问、搜索、导航、推理和答案生成混为一体。这种耦合使得难以确定SLM能否忠实地执行问题所隐含的推理路径,也难以将失败归因于导航而非流程的其他阶段。我们通过采用THESEUS导航与可追溯性框架,并使用冻结的、现成的SLM作为局部动作策略来隔离这一能力。在每一步中,环境暴露合法的出向图动作,模型选择一个可执行的图动作并决定是否停止,无需特定任务的参数更新、模型控制的束搜索或自由形式的答案生成。这种受控设置使我们能够使用Hits@1评估终端答案准确率,同时使用路径编辑距离(PED)作为主要轨迹指标评估路径保真度。在Kinship和MQuAKE-ST KGQA中,相似规模的局部模型在答案准确率和路径保真度上存在显著差异,这两个指标有时会偏向不同的模型。这种依赖模型的行为也扩展到提示(prompting),因为单个演示轨迹可以改善或恶化导航,具体取决于模型。这些结果促使我们超越端点准确率来评估SLM的图推理能力。

英文摘要:

Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a question and to attribute failures to navigation rather than to other stages of the pipeline. We isolate this capability by employing the THESEUS navigation and traceability framework and using frozen, off-the-shelf SLMs as local action policies. At each hop, the environment exposes the legal outgoing graph actions, and the model selects one executable graph action and decides whether to stop, without task-specific parameter updates, model-controlled beam search, or free-form answer generation. This controlled setting allows us to evaluate terminal-answer accuracy with Hits@1 together with path fidelity, using Path Edit Distance (PED) as the primary trajectory metric. Across the Kinship and MQuAKE-ST KGQAs, similarly sized local models differ substantially in answer accuracy and path fidelity, with the two metrics sometimes favoring different models. This model-dependent behavior also extends to prompting, as a single demonstrated trajectory can improve or degrade navigation depending on the model. These results motivate evaluating SLM graph reasoning beyond endpoint accuracy alone.

补充信息

↑